面向无人机巡检的玉米病虫害图文协同检测模型
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

河北省重大科技成果转化专项(22287401Z)、国家自然科学基金项目(42171212)和地理信息科学与技术全国重点实验室开放基金项目


Image-Text Collaborative Detection Model for Corn Pests and Diseases in Drone-based Inspections
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    结合无人机巡检与深度学习检测玉米病虫害是智慧农业发展趋势。然而深度学习模型检测性能不仅受图像样本数量影响,还受田间环境、成像质量及模型泛化能力等因素制约,导致单一图像模态难以全面表征病虫害特征。针对上述问题,本文基于YOLO-World模型,融合文本语义信息,提出一种面向无人机巡检的玉米病虫害图文协同检测模型。该模型结合重参数化视觉几何网络、注意力机制及多尺度融合模块,重构了图像特征提取网络,提高了多尺度图像特征提取能力;通过在文本注入图像中加入特征线性调制层和多类注意力,在图像嵌入文本中加入前馈模块、门控机制和残差结构,改进可重参数化的视觉-语言路径聚合网络,提升了图像与文本特征的语义对齐能力与表达精度。从无人机巡检视角采集500幅玉米病虫害图像,包含早期玉米黄叶病和虫害导致的烂叶病2类,并通过数据增强扩展至2000幅,结合对应文本描述构建COCO格式的多模态数据集。试验结果表明,在500幅数据集条件下(分别取25-shot、50-shot训练样本,固定400幅验证样本),本文模型mAP@0.5较单模态YOLO v8系列提升9.40~10.40个百分点,较YOLO-World提升2.80~4.40个百分点;在2000幅数据集上(按比例8:2划分训练集与验证集),mAP@0.5较单模态YOLO v8系列提高5.40~10.70个百分点,较YOLO-World提升3.10个百分点。该模型可为小样本条件下无人机巡检识别玉米病虫害提供新的研究思路。

    Abstract:

    Combining drone inspection with deep learning for maize pest and disease detection represents a trend in the development of smart agriculture.However, the detection performance of deep learning models is influenced not only by the number of image samples but also constrained by factors such as field environment, imaging quality, and model generalization capability, making it difficult for a single image modality to fully characterize pest and disease features.To address the above issues, a vision-language collaborative detection model for maize pests and diseases under drone inspection was proposed based on the YOLO-World framework, integrating textual semantic information.The model reconstructed the image feature extraction network by incorporating a re-parameterized visual geometry network, attention mechanisms, and multi-scale fusion modules, thereby improving the multi-scale image feature extraction capability.By adding a feature linear modulation layer and multi-class attention to the text-injected image, and introducing a feed forward module, gating mechanism, and residual structure to the image-embedded text, the re-parameterizable vision-language path aggregation network was enhanced, which improved the semantic alignment capability and expression accuracy between image and text features.In the experiments, totally 500 corn pest and disease images were captured from a drone inspection perspective, including two categories: early corn yellow leaf disease and leaf rot caused by pests.The dataset was augmented to 2000 images, and a multi-modal dataset in COCO format was constructed by incorporating corresponding textual descriptions.Experimental results showed that on the 500 image dataset (with 25-shot and 50-shot training samples respectively and a fixed validation set of 400 images), the proposed model improved mAP@0.5 by 9.40~10.40 percentage points over the single modal YOLO v8 series and by 2.80~4.40 percentage points over YOLO-World.On the 2000 image dataset (with an 8:2 training-validation split), mAP@0.5 was improved by 5.40~10.70 percentage points over the single modal YOLO v8 series and by 3.10 percentage points over YOLO-World.This model can provide a research approach for identifying corn pests and diseases under few-shot conditions in drone based inspection.

    参考文献
    相似文献
    引证文献
引用本文

承达瑜,苏皓,梁优,赵伟,王建东,宋辞.面向无人机巡检的玉米病虫害图文协同检测模型[J].农业机械学报,2026,57(18):106-116. Cheng Dayu, Su Hao, Liang You, Zhao Wei, Wang Jiandong, Song Ci. Image-Text Collaborative Detection Model for Corn Pests and Diseases in Drone-based Inspections[J]. Transactions of the Chinese Society for Agricultural Machinery,2026,57(18):106-116.

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-02-16
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-09-15
  • 出版日期:
文章二维码