追踪提示结合视觉大模型微调的猪只实例分割方法
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

江苏现代农业产业单项技术研发项目(CX(23)3120)


Pig Instance Segmentation Based on Tracking Prompts and Fine-tuned Large Vision Models
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    针对当前主流猪只感知算法存在追踪与分割任务缺乏协同、依赖高成本人工标注的难题,提出了一种端到端猪只追踪与分割框架。 该方法将追踪算法输出的锚框作为视觉大模型 SAM(Segment anything model)的动态提示, 以生成时序连贯且身份一致的个体掩模序列。 通过结合低秩自适应(Low rank adaptation, LoRA)与猪只多尺度特征适配器(PF Adapter),在小规模数据集上对 SAM 进行高效微调,使该模型仅需少量标注样本,即可在性能上超越传统的有监督学习算法,有效克服了对高昂人工标注的依赖。 实验结果表明,微调后的 SAM 模型在两种不同的数据集上的平均像素准确率、交并比和 Dice 相似指数分别达到了 91. 24% 、85. 06% 和 86. 18% ,参数量相较于 SAM 基准模型只增加了 3% 。 改进模型输出的猪只体态和运动轨迹信息可用于猪只评估和母猪分娩预警等,为智慧养殖提供技术支撑。

    Abstract:

    Aiming to address the challenges of lack of coordination between tracking and segmentation tasks and reliance on costly manual annotation in current mainstream pig perception algorithms, a novel end-to-end framework was proposed. Specifically, the proposed method ingeniously leveraged the bounding boxes output by the multi-object tracking algorithm as dynamic spatial prompts for the segment anything model (SAM). This integration facilitated the generation of temporally coherent and identity- consistent individual mask sequences across video frames, bridging the gap between localization and pixel-level segmentation. To adapt the vision foundation model to the specific agricultural domain, the low-rank adaptation (LoRA)was combined with a custom-designed pig multi-scale feature adapter (PF Adapter). The model required only a minimal number of annotated samples to achieve robust feature extraction, ultimately surpassing the performance of traditional fully supervised learning algorithms and effectively overcoming the bottleneck of expensive data annotation. Comprehensive experimental results demonstrated the superiority of the proposed framework. Evaluated across two distinct datasets, the fine- tuned SAM achieved impressive performance metrics, with pixel accuracy (PA )of 91. 24% , an intersection over union (IoU)of 85. 06% , and a dice similarity coefficient of 86. 18% . Remarkably, this significant performance enhancement was achieved with merely a 3% increase in number of parameters compared with the baseline SAM architecture. Furthermore, the pig body posture and movement trajectory information output by the improved model can be used for pig evaluation and sow farrowing warning, providing technical support for smart farming.

    参考文献
    相似文献
    引证文献
引用本文

孙立博,张泽昀,秦文虎.追踪提示结合视觉大模型微调的猪只实例分割方法[J].农业机械学报,2026,57(15):36-45. Sun Libo, Zhang Zeyun, Qin Wenhu. Pig Instance Segmentation Based on Tracking Prompts and Fine-tuned Large Vision Models[J]. Transactions of the Chinese Society for Agricultural Machinery,2026,57(15):36-45.

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-03-17
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-08-01
  • 出版日期:
文章二维码