Abstract:
Maize tassel detection is one of the most important techniques for crop phenotypic analysis in unmanned aerial vehicle (UAV) remote sensing in precision agriculture. However, existing detection is often confined to the tassel targets from a downward-looking UAV perspective, due to their edge blurring, dense occlusion, and multi-scale variations. In this study, an improved YOLO26n, MFDA-YOLO, was proposed to detect maize tassels. A multi-scene corn tassel detection (MSCTD) dataset was constructed using UAV images at the Changyuan Branch of Henan Academy of Agricultural Sciences from July 15 to September 20, 2025. The dataset covered the developmental stages from early tasseling to late flowering under different lighting and weather conditions. A DJI MAVIC 3M UAV equipped with a
5280×
3956--pixel camera was operated at an altitude of 12 m with 80% forward overlap and 70% side overlap. A total of 800 original images were cropped into
1280×
1280-pixel sub-images.
2840 valid sub-images were retained after quality screening. The original images were divided into training, validation, and test sets at a ratio of 7:2:1. All sub-images derived from the same original image were assigned to the same subset to avoid near-duplicate leakage. The final dataset contained
1988 training images, 568 validation images, and 284 test images. Data augmentation was applied only to the training set, increasing it to
5964 images, while the validation and test sets remained unchanged. Four targeted modifications were introduced. 1) A multi-frequency collaborative enhancement (MFCE) module was designed to strengthen tassel edge features using spatial-domain multi-scale dilated convolution, frequency-domain amplitude enhancement, and efficient channel attention (ECA). 2) DySample was replaced with fixed nearest-neighbor upsampling in the neck network to improve spatial detail recovery in densely occluded scenes using content-aware adaptive sampling. 3) A differential interactive feature aggregation (DIFA) module was developed to dynamically fuse shallow spatial details and deep semantic features by modeling their discrepancy and consistency. 4) Frequency-spatial convolution (FSConv) was replaced with standard strided convolution in the backbone to preserve high-frequency edge information through Haar wavelet decomposition and attention feature enhancement with the parameter count. Experimental results on the MSCTD dataset show that MFDA-YOLO achieved a precision of 90.68%, recall of 89.57%, mAP@0.5 of 94.89%, and mAP@0.5:0.95 of 62.74%, indicating improvements of 1.43, 1.18, 2.08, and 1.85 percentage points over YOLO26n, respectively. Compared with Faster R-CNN, YOLOv8, YOLOv9, and YOLOv10, its mAP@0.5 increased by 2.59, 3.45, 2.78, and 2.57 percentage points, respectively. There was an inference speed of 326.68 FPS with 3.09M parameters. Ablation results showed that the effectiveness of each module with MFCE provided the largest individual gain at low parameter cost. Deployment on an NVIDIA Jetson Orin Nano Super 8 GB using TensorRT FP16 achieved an mAP@0.5 of 94.60%, indicating practical deployability. This finding can provide technical support for maize phenotypic analysis, growth monitoring, and breeding selection.