融合语义与图像感知的棉花病虫害检测方法
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

石河子大学青年创新拔尖人才计划项目(CXBJ202306)和浙江省尖兵领雁+X科技项目(2026C04008)


Cotton Pests and Diseases Detection Method Based on Integration of Semantic and Image Perception
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    针对棉花田间病虫害目标尺度变化显著且单一视觉表征泛化能力不足的问题,本文提出一种融合语义先验与多尺度图像感知的多模态检测方法。以冻结CLIP模型为特征提取基础,借助文本调整模块与参数高效微调网络,将农艺文本转换为结构化提示,实现农学语义与视觉模态深度对齐;同时设计原型引导注意力机制,以可学习原型为桥梁,动态连接像素分布与语义空间。在检测模型层面,以VGG19为主干网络,融合Swin Transformer的全局上下文建模能力,并引入合成融合SFM模块和内块融合IFM模块构建连续尺度特征流,从而缓解传统离散下采样对微小目标特征的截断。试验结果表明,所提模型方法在棉花病虫害数据集上平均精度均值达92.5%,较其他主流模型YOLO v5s和YOLO v8n提升3.6、4.6个百分点。消融试验结果表明,语义引导与多尺度优化策略在性能提升上具有明显的协同效应,且模型参数量仅为1.56×10^6,远低于YOLO v5s和YOLO v8n;此外,Grad-CAM可视化结果表明,模型能够较稳定地聚焦于典型病灶区域。综上,该方法在精度与轻量化之间取得了较好的平衡,可为复杂农情条件下病虫害精准识别及边缘端部署提供参考。

    Abstract:

    Aiming to address the bottlenecks posed by the significant variation in the target scales of cotton field pests and diseases and the limited generalization capability of single-modal visual representations, a multimodal detection method that integrated semantic prior knowledge with multi-scale image perception was proposed.The method used a frozen CLIP model as the foundation for feature extraction.By leveraging a text adaptation module and a parameter-efficient fine-tuning network, it converted agronomic text into structured prompts, thereby achieving deep alignment between agronomic semantics and visual modalities.Additionally, a prototype-guided attention mechanism was designed to dynamically link pixel distributions with the semantic space based on learnable prototypes.At the detection model, VGG19 served as the backbone network, integrating the global context modeling capabilities of swin transformer.Additionally, SFM and IFM modules were introduced to construct a continuous-scale feature flow, thereby mitigating the truncation of fine-scale object features caused by traditional discrete downsampling.Experiments demonstrated that the proposed model achieved an mAP of 92.5% on the cotton pest and disease dataset, outperforming other mainstream models YOLO v5s and YOLO v8n by 3.6 and 4.6 percentage points, respectively.Ablation results further revealed that the semantic guidance and multi-scale optimization strategies exhibited a significant synergistic effect in performance improvement, while the model's parameter count was only 1.56×10^6, which was much lower than that of YOLO v5s and YOLO v8n.Additionally, Grad-CAM visualization results indicated that the model can consistently focus on typical lesion areas.In summary, this method achieved a balance between accuracy and model compactness, providing a reference for the precise identification of pests and diseases under complex agricultural conditions and for edge deployment.

    参考文献
    相似文献
    引证文献
引用本文

李阳,吴科,聂晶,方凯.融合语义与图像感知的棉花病虫害检测方法[J].农业机械学报,2026,57(18):51-61. Li Yang, Wu Ke, Nie Jing, Fang Kai. Cotton Pests and Diseases Detection Method Based on Integration of Semantic and Image Perception[J]. Transactions of the Chinese Society for Agricultural Machinery,2026,57(18):51-61.

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-05-20
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-09-15
  • 出版日期:
文章二维码