高级检索+

面向少样本病害识别的双提示层次化交互视觉语言模型

Dual-prompt hierarchical interactive vision-language model for few-shot agricultural disease classification

  • 摘要: 针对特色作物种植、新发疫病流行及田间复杂环境等农业生产场景中标注样本稀缺、跨区域症状差异大等现实问题,传统深度学习依赖大规模标注难以落地推广,该文提出一种基于视觉-语言模型的少样本病害识别方法——双提示层次化交互视觉语言模型(dual-prompt hierarchical interactive vision-language model, DPHI)。该模型基于CLIP(contrastive language-image pre-training)架构,设计协同提示交互模块(synergistic prompt interaction, SPI),通过该模块在编码器各层提取可学习提示,并通过类别掩码引导的双向注意力实现视觉与文本的深层语义交互;构建动态原型约束机制(dynamic prototype constraint, DPC),采用EMA(exponential moving average)策略动态更新类别视觉原型以增强特征判别性;并结合多任务协同对比损失优化特征空间。在PlantVillage、PlantDoc与Sugarcane Leaf Disease数据集上的16-shot试验中,DPHI整体性能优于CoOp、MaPLe等主流提示学习方法。此外,在陕西省铜川市实地采集的苹果病害数据集上,DPHI模型在16-shot设置下取得宏平均准确率92.11%、宏平均F1分数92.20%,对各类病害均表现出稳定的识别能力,在样本数量极少的类别上仍能保持较高精度。同时,DPHI模型采用冻结主干、仅训练少量提示参数的范式,在保证识别精度的同时,具有较低的训练与推理开销,可为降低农业病害识别技术应用成本、提升模型落地可行性提供一定参考。

     

    Abstract: Accurate and timely recognition of pests and diseases is often required in precision agriculture. However, conventional deep learning models rely on large-scale annotated data for effective training, leading to weak generalization. The challenges also remain, including the extreme scarcity of high-quality labeled samples, serious symptom variations over regions and crop varieties, strong interference from complex field backgrounds, and highly imbalanced sample distributions among different disease categories. In this study, a high-performance, parameter-efficient, and robust vision-language model was developed for stable and accurate plant disease identification under extremely limited training samples. Real-world deployment was also realized in field environments, particularly for the wide application of intelligent disease diagnosis. A Dual-prompt Hierarchical Interactive vision-language model (DPHI) was constructed using the Contrastive Language-Image Pre-training (CLIP) architecture. Three components were utilized to strengthen deep cross-modal fusion for feature identification. (1) A Synergistic Prompt Interaction (SPI) module was constructed to extract learnable visual and textual prompts hierarchically at each layer of the visual and text encoders. A class-mask-guided bidirectional cross-attention mechanism was applied to realize fine-grained semantic alignment between visual and textual features, while filtering out irrelevant background noise and suppressing inter-class feature confusion. (2) A Dynamic Prototype Constraint (DPC) mechanism was established to maintain stable and consistent category representations, in which the Exponential Moving Average (EMA) strategy was used to dynamically update class visual prototypes. The disturbance was reduced from small-batch sampling bias, noisy data, and outlier samples. (3) A multi-task collaborative contrastive loss function was designed to optimize the shared visual-text feature space, which combined cross-entropy loss, relative distance loss, and textual dispersion loss to simultaneously strengthen intra-class compactness and inter-class separability. Comprehensive experiments were implemented under the 16-shot few-shot setting on three representative public datasets: PlantVillage, PlantDoc, and Sugarcane Leaf Disease Dataset, which covered controlled laboratory environments, complex real-field conditions, and single-crop fine-grained disease classification. The results showed that the DPHI model achieved the highest classification accuracy and F1-score on all three datasets, outperforming state-of-the-art prompt learning, including Context Optimization (CoOp) and Multi-modal Prompt Learning (MaPLe), as well as classic convolutional neural networks and Transformer-based recognition models. An apple disease dataset was collected from Yaozhou District, Tongchuan City, Shaanxi Province, China. The DPHI obtained 92.11% macro-average accuracy and 92.20% macro-average F1-score, indicating the outstanding recognition precision even in classes with extremely scarce samples, such as brown spot disease. The pre-trained CLIP backbone network was frozen to train only a small number of learnable prompt parameters. Llow computational cost and memory usage were attained in both training and inference stages, which greatly improved training efficiency with the low hardware requirements for practical deployment. The DPHI model significantly improved the few-shot learning performance, fine-grained discrimination, and cross-scenario generalization for agricultural disease recognition. The findings can also provide a reliable, parameter-efficient, and low-cost technical solution for intelligent disease diagnosis in real agricultural applications, such as rare cash crops, newly emerging epidemic diseases, and cross-regional popularization. The application cost of intelligent disease recognition can be effectively reduced to ensure the practical feasibility of model deployment in complex field environments.

     

/

返回文章
返回