
本文详细介绍如何用地平线的地瓜S100P芯片部署yolov5检测模型时值十一前夕在此国庆之日。遥想2023在RV1126上部署yolo模型已经是3年之久彼时对深度学习理解非常浅薄甚至没有用pytorch 框架训练过模型当时用RK的toolkit 转换yolov3,转完之后是可以跑的但是转完v5之后死活不行折腾了2周才在长沙宇文工的帮助下跑通待返回青岛时我负责的业务已被砍掉。最近在地瓜S100P上转yolo模型此时对yolo的结构已经十分熟悉花费了2天的时间完成模型转换由于官方的demo还是有点儿坑自己也是在不确定中跑通的过程亦非完全顺利故予以作文以记之。1 .安装docker这一步是要配置一个和官方完全一样的模型转换、评估环境。基本没啥坑。S100P install docker我的电脑是CPU环境运行如下命令进入docker.bashrun_docker.sh data/ cpucd 到这个目录下PC上/open_explorer/samples/vision/ultralytics_yolo/conversion把地瓜S100P上的模型复制到本地PC上运行下面的命令我们可以看一看模型的信息hb_model_info yolov5x_672x672_nv12.hbm拉到最后可以看到是这样的其中一个输出尺寸是【18484255】其中84672/8762是图像的输入尺寸。2026-09-15 09:42:08,007 INFO data_y input[1,672,672,1]UINT82026-09-15 09:42:08,007 INFO data_uv input[1,336,336,2]UINT82026-09-15 09:42:08,007 INFO output output[1,84,84,255]INT322026-09-15 09:42:08,007 INFO1310output[1,42,42,255]INT322026-09-15 09:42:08,007 INFO1312output[1,21,21,255]INT322 导出onnx 模型我下载的是原版yolov5模型这个需要简单改一下前向传播的处形状。2.1 修改的代码ultralytics8.4.0在yolo.pydef forward(self, x): z[]# inference outputforiinrange(self.nl): x[i]self.m[i](x[i])# convbs, _, ny, nxx[i].shape# [bs, na*no, ny, nx] - [bs, na, no, ny, nx] - [bs, na, ny, nx, no]x[i]x[i].view(bs, self.na, self.no, ny, nx).permute(0,1,3,4,2).contiguous()# [bs, na, ny, nx, no] - [bs, ny, nx, na, no] - [bs, ny, nx, na*no]x[i]x[i].permute(0,2,3,1,4).reshape(bs, ny, nx, self.na * self.no).contiguous()z.append(x[i])returnxifself.trainingelsez# 训练时保持原样推理/导出返回 3 个张量在export.pyoutput_names[output0,output1,output2]afer this change2.2 exportpython export.py--weightsyolov5s.pt--includeonnx--imgsz672672--opset11--simplify输入如下可简单测试一下导出是否OK1973python3 -EOF1974importonnx1975monnx.load(yolov5s.onnx)1976print(inputs:)1977foriinm.graph.input:1978print( , i.name,[d.dim_valuefordini.type.tensor_type.shape.dim])1979print(outputs:, len(m.graph.output))1980foroinm.graph.output:1981print( , o.name,[d.dim_valuefordino.type.tensor_type.shape.dim])1982EOF3 onnx 转 hbm在 /open_explorer/samples/vision/ultralytics_yolo/conversion下运行python3 mapper.py--onnxyolov5s_3.onnx --cal-images ./cal_images--marchnash-m参数解释mapper.py转换代码yolov5s_3.onnx刚才导出的onn模型./cal_images和你训练数据集一样的图片直接从训练数据集中取2张nash-m表示是S100P––deng等待差不多5分钟。出现一个小插曲2026-09-1510:09:21,723 INFO images_y input[1,672,672,1]UINT82026-09-1510:09:21,723 INFO images_uv input[1,336,336,2]UINT82026-09-1510:09:21,724 INFO output0 output[1,84,84,255]FLOAT322026-09-1510:09:21,724 INFO output1 output[1,42,42,255]FLOAT322026-09-1510:09:21,724 INFO output2 output[1,21,21,255]FLOAT322026-09-1510:09:21,725 INFO The hb_compile completes running.S100p给的demo 数据类型是 int32,但是我转出来的却是Float 32。4 改官方的例程在 /app/cdev_demo/bpu/utils/src文件夹下std::vectorfloatdequantizeTensorS32(consthbDNNTensortensor){constautopropstensor.properties;constautoshapeprops.validShape;constintNshape.dimensionSize[0];constintHshape.dimensionSize[1];constintWshape.dimensionSize[2];constintCshape.dimensionSize[3];inttotalN*H*W*C;std::vectorfloatresult(total);uint8_t*basestatic_castuint8_t*(tensor.sysMem.virAddr);// 新增FLOAT32 分支 if(props.tensorTypeHB_DNN_TENSOR_TYPE_F32){// F32 输出按 NHWC 连续读取// 注意这里假设输出是连续的没有 padding// 如果有 padding需要按 stride 遍历constfloat*fptrreinterpret_castconstfloat*(base);for(inti0;itotal;i){result[i]fptr[i];}returnresult;}// 原来的 S32 分支 constfloat*scale_dataprops.scale.scaleData;// per-channel scaleconstint32_t*zp_dataprops.scale.zeroPointData;// per-channel zero-point (optional)intstrideprops.stride[3];// byte stride between channelsfor(inth0;hH;h){intoffset_hh*shape.dimensionSize[2];intidx_hh*W;for(intw0;wW;w){intoffset_w(offset_hw)*props.stride[2];intidx_w(idx_hw)*C;for(intc0;cC;c){intidxidx_wc;intoffsetoffset_wc*stride;floatsscale_data[c];intzpzp_data?zp_data[c]:0;floatvalstatic_castfloat(*(int32_t*)(baseoffset));result[idx](val-zp)*s;}}}returnresult;}/app/cdev_demo/bpu/02_detection_sample/01_ultralytics_yolov5x/buildruncmake..makesunriseubuntu:/app/cdev_demo/bpu/02_detection_sample/01_ultralytics_yolov5x/build$ ./ultralytics_yolov5x[UCP]: log level3[UCP]: UCP version3.13.6[VP]: log level3[DNN]: log level3[HPL]: log level3[UCPT]: log level6./ultralytics_yolov5x[BPU][[BPU_MONITOR]][281472686818656][INFO]BPULib verison(2,2,15)[f21ee84]![DNN]:3.13.6_(4.7.5 HBRT)output_count_:3 feature_size:85 all_results size:3[Saved]Result saved to: result.jpg至此大功告成