高级检索+

基于自适应融合敏感度感知的绿色小果实检测方法

Green small fruit detection method based on adaptive fusion sensitivity perception

  • 摘要: 为提高果园疏果期绿色小果实的精准检测能力,该研究针对绿色小果实与枝叶背景颜色相近、目标尺寸偏小及对比度偏低等复杂检测难题,提出了一种尺度-对比度联合感知优化的AFS-DETR(Adaptive Fusion and Sensitivity-aware Real-Time DEtection TRansformer)绿色小果实检测模型。首先,在特征提取阶段设计C2AF(C2f_AdaptiveFusion)模块,采用多分支动态卷积与自适应核加权机制,增强细节特征感知能力,减少浅层信息损失。其次,在特征融合阶段引入基于超图理论的HACE(HyperACE)模块,利用其特征间高阶相关性的建模优势,提升自适应多尺度特征融合能力。最后,在特征金字塔传递路径上集成FPADT(FullPAD_Tunnel)轻量门控对齐模块,通过上下文感知优化跨层级特征融合。试验结果表明,在自建黄金梨数据集上,AP50、AP75、APS分别为90.4%、78.4%、56.8%,较基线RT-DETR-R18分别提升1.1、1.9和1.0个百分点。在自建海棠果数据集、公开数据集MinneApple上,AP分别为65.6%和40.9%,均领先对比模型,验证了模型的泛化能力。该模型能够有效满足果园疏果期绿色小果实的精准检测需求,为果园智能化管理提供了关键技术支撑。

     

    Abstract: Green small fruit detection during the orchard thinning stage poses significant challenges due to high color similarity between target fruits and background foliage, small physical dimensions, low image contrast, and frequent dense occlusion. These characteristics caused existing detection frameworks to suffer from shallow feature loss through repeated backbone downsampling, insufficient saliency modeling under low-contrast conditions, and semantic misalignment across feature pyramid levels. An improved detection model, AFS-DETR (Adaptive Fusion and Sensitivity-aware Real-Time DEtection TRansformer), was proposed based on Real-Time DEtection TRansformer (RT-DETR) to improve detection accuracy for green small fruits in complex orchard environments. Three targeted modules were incorporated into the RT-DETR architecture. The C2f_AdaptiveFusion (C2AF) module was designed to replace the standard backbone block with a multi-branch dynamic convolution structure combined with an Adaptive Kernel Weights (AKW) mechanism, enhancing shallow-layer detail feature representation for tiny-target localization. The HyperACE (HACE) module was introduced into the feature pyramid neck to model high-order semantic dependencies across multi-scale features via dynamic hyperedge construction, improving adaptive fusion under low-contrast conditions. The FullPAD_Tunnel (FPADT) gated alignment module was introduced to apply a learnable scalar gate to suppress cross-level semantic misalignment at the feature-pyramid-to-decoder interface with minimal computational overhead. Experiments were conducted on three datasets spanning different fruit species and scene conditions. On the self-built Golden Pear dataset comprising 2,438 high-resolution images captured under diverse orchard conditions including top-view, upward-view, front-lighting, back-lighting, occlusion, and stacking scenarios, AFS-DETR achieved an average precision at intersection-over-union threshold 0.5 (AP50) of 90.4%, AP at threshold 0.75 (AP75) of 78.4%, recall of 84.5%, and F1-score of 86.9%. Small-object average precision (APS) reached 56.8%, representing improvements of 1.1, 1.9, and 1.0 percentage points over the RT-DETR-R18 baseline in AP50, AP75, and APS, respectively. AFS-DETR recorded the highest APS and medium-object average precision (APM) of 85.7% among all compared methods, including YOLOv8-n, YOLOv11-n, YOLOv13-n, Faster R-CNN (Region-based Convolutional Neural Network), Deformable DETR (DEtection TRansformer), RTMDet-Tiny(Real-Time Models for Object Detection), and Dynamic R-CNN. On the public MinneApple benchmark comprising 1,001 images with mixed red and green apple instances at dense small-scale distributions, AFS-DETR achieved a mean average precision (AP) of 40.9% and APS of 27.1%, with the AP ranking first among all compared models; the AP of 40.9% exceeded Deformable DETR (18.7%) by more than a factor of two and surpassed the latest lightweight detectors YOLOv12-n (39.9%) and YOLOv13-n (36.2%). On the self-built Crabapple dataset comprising 1,057 images including nighttime low-illumination scenes, AFS-DETR achieved an AP50 of 89.6%, AP of 65.6%, and APS of 52.0%, outperforming YOLOv11-n (APS of 49.2%), YOLOv13-n (APS of 45.6%), and Faster R-CNN (APS of 35.1%). Ablation experiments on the Golden Pear dataset confirmed that C2AF provided the largest individual performance gain as the feature-extraction foundation; the C2AF and FPADT combination achieved the highest recall of 84.3% and F1-score of 87.5% among all dual-module configurations; and the complete AFS-DETR attained the optimal AP75 of 78.4%, with superior APS verified consistently across all three datasets. Regarding model efficiency, AFS-DETR contained 15.2 M parameters and 46.8 giga floating-point operations (GFLOPs), representing reductions of 23.2% and 17.8% relative to the RT-DETR-R18 baseline, respectively. End-to-end processing speed on an NVIDIA RTX 3090 graphics processing unit (GPU) reached 65.5 frames per second (FPS), satisfying quasi-real-time requirements for field deployment. AFS-DETR effectively addressed the core detection difficulties of green small fruits in complex orchard environments—low foreground-background contrast, fine-scale target dimensions, and dense occlusion—through three complementary and modular structural improvements. The model demonstrated consistent performance advantages across self-built and public benchmark datasets, confirming generalization capability across fruit species and scene variations. The proposed framework provides a reliable and deployable technical solution for precision orchard management, with direct applicability to yield estimation, growth monitoring, and intelligent robotic fruit-thinning operations.

     

/

返回文章
返回