基于多尺度增强与几何自适应融合的食品包装文本检测模型
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

数字赋能国家植物园建设及种质资源收集保护关键技术研究与示范项目(Z221100005222018)


Text Detection Model for Food Packaging Based on Multi-scale Enhancement and Geometric Adaptive Fusion
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    食品包装文本包含产品名称、成分、厂商信息和保质期等关键信息,对食品安全与消费者权益具有重要意义。将现有的场景文本检测方法应用于食品包装文本检测任务中,普遍存在模糊文本、小目标以及艺术字体的漏检和误检等问题,针对这些问题,提出了一种基于多尺度增强与几何自适应融合的食品包装文本检测模型MESA-Text。具体而言,在MESA-Text中首先提出多尺度注意力特征增强融合模块,通过注意力奇偶层交替融合策略,集成门控双分支注意力机制和特征平滑方法,增强融合不同尺度文本特征能力;然后针对文本高宽比失调问题,提出几何稳定模块,融合文本几何特征进行自适应改进,以精准捕捉扭曲、模糊及微小文本特征;最后辅以动态可变形卷积增强不同尺度特征提取,以提升食品包装文本检测效果。消融实验结果表明,引入多尺度注意力特征增强融合模块、几何稳定模块及动态可变形多尺度特征提取模块后,召回率分别提升2.18、0.90、1.14个百分点。采用本文模型方法在食品包装图像上的召回率、精确率和F1值分别达92.12%、91.76%和91.94% ,有效缓解漏检误检问题。

    Abstract:

    Text on food packaging contains key information such as product names, ingredients, manufacturer details, and expiration dates, which are crucial for food safety and consumer rights. However, applying existing scene text detection methods to food packaging often results in missed or incorrect detections, especially for blurred text, small targets, and artistic fonts. To address these challenges, a food packaging text detection model was proposed based on multi-scale enhancement and geometric fusion, named MESA-Text. MESA-Text introduced a multi-scale attention feature enhancement and fusion module, which leveraged an alternating fusion strategy of attention-based even-odd layers, integrating a gated dual-branch attention mechanism and a feature smoothing method to effectively enhance the ability to fuse features from different text scales. To address the problem of imbalanced text aspect ratios, a geometric feature stabilization module was proposed to adaptively refine features by incorporating geometric information, allowing more accurate and robust detection of distorted, blurred, and small-scale text. Additionally, a dynamic deformable convolution module was used to enhance feature extraction across multiple scales, further improving the representation capability of complex text regions and enhancing overall detection performance. The proposed framework was specifically designed to address the diverse appearance characteristics of food packaging text, enabling more reliable localization under complex backgrounds and varying imaging conditions. Ablation experiments demonstrated that incorporating the multi-scale attention enhancement module, geometric stabilization module, and dynamic multi-scale feature extraction module improved recall by 2.18, 0.90, and 1.14 percentage points, respectively. Using the proposed model, MESA-Text achieved 92.12% recall, 91.76% precision, and F1-score of 91.94% on food packaging images, demonstrating superior effectiveness and robustness in challenging food packaging scenarios, while effectively reducing missed and false detections.

    参考文献
    相似文献
    引证文献
引用本文

田萱,肖文昊.基于多尺度增强与几何自适应融合的食品包装文本检测模型[J].农业机械学报,2026,57(17):392-403. Tian Xuan, Xiao Wenhao. Text Detection Model for Food Packaging Based on Multi-scale Enhancement and Geometric Adaptive Fusion[J]. Transactions of the Chinese Society for Agricultural Machinery,2026,57(17):392-403.

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-05-07
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-09-01
  • 出版日期:
文章二维码