高级检索+

基于改进YOLO11n的非结构化环境下枸杞果实与果梗多尺度轻量检测方法

Multi-scale lightweight detection method for wolfberry fruits and pedicels in unstructured environments based on improved YOLO11n

  • 摘要: 针对自然光照变化剧烈、枝叶密集交错遮挡的非结构化农业环境下,枸杞果实与果梗因形态差异大、密集成串导致的检测精度低,实时性差的问题。该研究提出一种基于改进YOLO11n的非结构化环境枸杞果实与果梗多尺度轻量检测方法DFL-YOLO。首先,在骨干网络中引入C3k2-DAttention变形注意力机制模块,通过动态调整感受野形状以适应不同目标形态;其次,设计特征聚焦扩散金字塔网络(feature-focusing diffusion pyramid network, FDPN)构建双向特征流以强化多尺度信息融合,减少果梗的特征流失;再次,采用轻量化共享卷积检测头(lightweight shared convolution head design,LSCHD),降低计算冗余;最后,结合层级自适应幅值剪枝(layer-adaptive magnitude-based pruning,LAMP)策略,实现模型高效压缩,在包含5 760张图像的自建枸杞数据集上进行试验。结果表明,剪枝后的DFL-YOLO模型的精确率、召回率和交并比为0.5时的平均精度(average precision, AP)均值分别为90.0%、76.6%、82.4%。相比YOLO11n基线模型,AP50和AP50-95分别提高了1.7和3.0个百分点,参数量降低至1.45 M(降幅为44%),单幅图像推理耗时仅2.2 ms。田间试验结果表明,DFL-YOLO在非结构化田间环境中精确率为88.2%,召回率为73.8%,AP50为78.7%,满足田间复杂工况下果实与果梗同步高精度检测的要求。该研究在有效降低小目标漏检率的同时优化检测效率,研究结果可为枸杞采摘机器人的视觉感知系统提供可靠的技术支撑。

     

    Abstract: To comprehensively address the critical challenges in unstructured agricultural environments where wolfberry (Lycium barbarum L.) fruits and their corresponding pedicels exhibit extreme morphological scale variations, grow in dense clusters, and are frequently occluded by complex branches and leaves this study proposes DFL-YOLO, a lightweight detection model based on an improved YOLO11 architecture. Existing object detection models typically suffer from high missed detection rates for micro-pedicels, possess excessively large parameter volumes, and struggle to meet the embedded deployment requirements of agricultural mobile harvesting equipment with limited computing power. This research utilizes the YOLO11 model as the baseline to execute structural refinement and lightweight optimization. First, the C3k2-DAttention (deformable attention) mechanism is introduced into the backbone network. By dynamically adjusting the shape of the receptive field, this module actively focuses on and enhances the feature extraction capability for irregular targets, such as elliptical fruits and linear pedicels, effectively suppressing interference from complex foliage backgrounds. Second, to tackle the problem wherein fine pedicel features are easily lost and degraded within deep convolutional networks, a Feature-focusing Diffusion Pyramid Network (FDPN) is designed. By constructing a bidirectional feature flow, the FDPN achieves the highly effective fusion of shallow spatial texture features with deep semantic information, maximally preserving weak geometric features. Third, a Lightweight Shared Convolution Head Design (LSCHD) is adopted to replace the original decoupled detection head. Utilizing a parameter-sharing mechanism, this module compels the network to learn cross-scale generic features, significantly reducing computational costs and parameter redundancy from the perspective of network architecture. Finally, combined with a Layer-adaptive Magnitude-based Pruning (LAMP) strategy, the model is sparsely compressed to accurately eliminate non-critical redundant parameter channels. Extensive experiments were conducted on a self-constructed wolfberry dataset comprising 5 760 images captured in complex environments. The quantitative results demonstrate that the pruned and optimized DFL-YOLO model contains merely 1.45 M parameters and requires 6.3 GFLOPs of computation. Compared with the baseline YOLO11 model, the parameter volume is dramatically reduced by 44% under the premise of maintaining identical computational complexity. Furthermore, the mean average precision at an IoU threshold of 0.5 (AP50) and AP50-95 reach 82.4% and 55.5%, respectively, and the F1 score achieves 82.0%. These core metrics represent significant improvements of 1.7, 3.0, and 1.0 percentage points over the baseline model, respectively, with a single-image inference latency of only 2.2 ms. Compared with mainstream object detection models such as YOLOv5, YOLOv8, and RT-DETR, DFL-YOLO exhibits obvious advantages in both detection accuracy and lightweight performance. Additionally, rigorous field experiments conducted in genuine unstructured environments characterized by strong illumination, severe backlighting, and heavy branch-and-leaf occlusions indicate that DFL-YOLO achieves a precision of 88.2%, a recall of 73.8%, and an AP50 of 78.7% across 313 valid field targets, successfully minimizing missed and false detections to 82 and 31, respectively. Through the dual strategy of structural feature enhancement and systematic model pruning, the improved DFL-YOLO model effectively elevates the detection accuracy of tiny pedicels and overlapping fruits in complex environments while realizing a robust lightweight architecture, thereby achieving an optimal balance between detection efficiency and accuracy. The research findings provide highly reliable technical support for the visual perception systems of automated wolfberry harvesting robots.

     

/

返回文章
返回