高级检索+

利用无人机图像和改进YOLO26的玉米雄穗识别

Recognizing maize tassels using UAV images and improved YOLO26

  • 摘要: 玉米雄穗检测是无人机遥感精准农业中作物表型分析的关键技术。针对无人机垂直俯视视角下玉米雄穗目标现有检测方法存在边缘模糊、密集遮挡、多尺度变化及小目标细节易丢失等问题,该研究以YOLO26n为基线,提出一种改进的无人机玉米雄穗检测方法MFDA-YOLO。针对俯视视角下雄穗边缘模糊、与背景颜色相近的问题,设计多频协同增强模块(multi-frequency collaborative enhancement,MFCE),通过多尺度空洞卷积与频域幅度增强双路协同,配合高效通道注意力(efficient channel attention,ECA)机制从空域与频域两个维度强化雄穗边缘轮廓特征。针对密集遮挡场景下固定插值上采样导致空间细节恢复不足的问题,引入动态轻量上采样算子DySample替换颈部网络固定插值上采样,通过学习采样坐标偏移量实现内容感知的自适应上采样,改善密集雄穗区域的空间细节恢复效果。针对不同生长阶段雄穗目标尺度变化显著、多尺度特征直接拼接融合效果不足的问题,构建差异交互特征聚合模块(differential interactive feature aggregation,DIFA),通过自适应门控权重动态融合浅层细节与深层语义。针对无人机高空俯拍时雄穗成像面积小、下采样过程中高频边缘细节易丢失的问题,引入频率空间卷积(frequency-spatial convolution,FSConv)替换骨干网络标准卷积,结合哈尔离散小波变换(Haar discrete wavelet transform,Haar DWT)分离高低频分量,在降低参数量的同时增强雄穗边缘细节感知能力。在多场景玉米雄穗检测(multi-scene corn tassel detection,MSCTD)数据集上的试验结果表明,MFDA-YOLO的精确率、召回率、mAP@0.5和mAP@0.5:0.95分别为90.68%、89.57%、94.89%和62.74%,较基线模型YOLO26n分别提升1.43、1.18、2.08和1.85个百分点;与Faster R-CNN(ResNet-50)、YOLOv8n、YOLOv9t、YOLOv10n等主流检测模型相比,mAP@0.5分别提升2.59、3.45、2.78和2.57个百分点,推理速度326.68 帧/s。边缘部署试验表明,MFDA-YOLO在NVIDIA Jetson Orin Nano Super 8 GB平台上的mAP@0.5、mAP@0.5:0.95分别为94.60%、62.50%,推理速度为41.8 帧/s,满足25 帧/s的通用田间实时检测标准。该研究可为玉米表型分析、生长监测与育种选择提供技术支持。

     

    Abstract: Maize tassel detection is one of the most important techniques for crop phenotypic analysis in unmanned aerial vehicle (UAV) remote sensing in precision agriculture. However, existing detection is often confined to the tassel targets from a downward-looking UAV perspective, due to their edge blurring, dense occlusion, and multi-scale variations. In this study, an improved YOLO26n, MFDA-YOLO, was proposed to detect maize tassels. A multi-scene corn tassel detection (MSCTD) dataset was constructed using UAV images at the Changyuan Branch of Henan Academy of Agricultural Sciences from July 15 to September 20, 2025. The dataset covered the developmental stages from early tasseling to late flowering under different lighting and weather conditions. A DJI MAVIC 3M UAV equipped with a 5280×3956--pixel camera was operated at an altitude of 12 m with 80% forward overlap and 70% side overlap. A total of 800 original images were cropped into 1280×1280-pixel sub-images. 2840 valid sub-images were retained after quality screening. The original images were divided into training, validation, and test sets at a ratio of 7:2:1. All sub-images derived from the same original image were assigned to the same subset to avoid near-duplicate leakage. The final dataset contained 1988 training images, 568 validation images, and 284 test images. Data augmentation was applied only to the training set, increasing it to 5964 images, while the validation and test sets remained unchanged. Four targeted modifications were introduced. 1) A multi-frequency collaborative enhancement (MFCE) module was designed to strengthen tassel edge features using spatial-domain multi-scale dilated convolution, frequency-domain amplitude enhancement, and efficient channel attention (ECA). 2) DySample was replaced with fixed nearest-neighbor upsampling in the neck network to improve spatial detail recovery in densely occluded scenes using content-aware adaptive sampling. 3) A differential interactive feature aggregation (DIFA) module was developed to dynamically fuse shallow spatial details and deep semantic features by modeling their discrepancy and consistency. 4) Frequency-spatial convolution (FSConv) was replaced with standard strided convolution in the backbone to preserve high-frequency edge information through Haar wavelet decomposition and attention feature enhancement with the parameter count. Experimental results on the MSCTD dataset show that MFDA-YOLO achieved a precision of 90.68%, recall of 89.57%, mAP@0.5 of 94.89%, and mAP@0.5:0.95 of 62.74%, indicating improvements of 1.43, 1.18, 2.08, and 1.85 percentage points over YOLO26n, respectively. Compared with Faster R-CNN, YOLOv8, YOLOv9, and YOLOv10, its mAP@0.5 increased by 2.59, 3.45, 2.78, and 2.57 percentage points, respectively. There was an inference speed of 326.68 FPS with 3.09M parameters. Ablation results showed that the effectiveness of each module with MFCE provided the largest individual gain at low parameter cost. Deployment on an NVIDIA Jetson Orin Nano Super 8 GB using TensorRT FP16 achieved an mAP@0.5 of 94.60%, indicating practical deployability. This finding can provide technical support for maize phenotypic analysis, growth monitoring, and breeding selection.

     

/

返回文章
返回