基于TFViT-BiLSTM与Wav2Vec2.0的蛋鸡发声识别模型
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

现代农业产业技术体系北京市智慧农业创新团队项目(BAIC10-2025-E04)、北京市农林科学院改革与发展项目(GGFZ20250407)和新一代人工智能国家科技重大专项(2021ZD0113803)


Recognition Model of Laying Hens' Vocalizations Based on TFViT-BiLSTM and Wav2Vec2.0
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    针对当前蛋鸡发声识别特征提取较为依赖领域知识、单一声学特征表征能力不足而导致分类精准度受限等问题,本文提出了一种基于TFViT-BiLSTM与Wav2Vec2.0的蛋鸡发声识别模型。通过多头注意力机制对时频视觉转换器(Time-frequency vision transformer, TFViT)提取的频谱图特征与双向长短期记忆网络处理的梅尔倒谱系数进行特征融合;并将融合后的分类预测结果与Wav2Vec2.0模型分类预测结果合并输入XGBoost分类器进行决策融合。通过将ViT标准的全局自注意力机制替换为时频注意力机制,同时利用时频卷积前馈网络替代全连接多层感知机,有效捕捉频谱图局部时频相关性。试验结果表明,构建的模型在蛋鸡音频数据集上整体表现优异,宏准确率、宏召回率及宏F1值分别为97.79%、97.75%和0.9778,显示出较高的识别灵敏度与精确度。在3类典型蛋鸡发声(尖叫声、产蛋声、痛苦声)识别任务中,模型在各类发声上表现略有差异:尖叫声识别效果最优,精确率为98.38%、召回率为99.42%、F1值为0.9884;产蛋声识别效果次之,精确率为98.28%、召回率为98.25%、F1值为0.9807;痛苦声识别表现相对较弱,精确率为96.71%、召回率为95.58%、F1值为0.9643。研究结果可为蛋鸡发声自动识别提供有效的技术支持。

    Abstract:

    Due to the current problem existed in the extraction of vocal recognition features for laying hens relying on domain knowledge, and the single acoustic feature representation ability is insufficient, resulting in limited classification accuracy, a recognition model was proposed based on TFViT-BiLSTM and Wav2Vec2.0. The model performed feature-level fusion by combining the spectrogram features extracted by the time-frequency vision transformer (TFViT) using a multi-head attention mechanism with the Mel-frequency cepstral coefficients processed by a bidirectional long short-term memory network. The fused classification prediction results were then merged with the predictions from the Wav2Vec2.0 model and fed into an XGBoost classifier for decision-level fusion. Specifically, it improved upon the existing ViT model by replacing the standard global self-attention mechanism with a time-frequency attention mechanism, enhancing the model's ability to capture harmonic features of sound. Additionally, a time-frequency convolutional feed-forward network replaced the fully connected multilayer perceptron, effectively capturing the local time-frequency correlations of the spectrogram. The experimental results showed that the constructed model performed excellently on the laying hen audio dataset with a macro accuracy of 97.79%, a macro recall of 97.75%, and a macro F1-score of 0.9778, demonstrating high recognition sensitivity and precision. In the recognition task of three typical types of laying hen vocalizations (screams, laying sounds and distress calls), the model presented differentiated recognition performance for various vocalizations. The model achieved the optimal recognition effect on screams with a precision of 98.38%, a recall of 99.42%, and an F1-score of 0.9884. The recognition performance for laying sounds ranked the second, with a precision of 98.28%, a recall of 98.25%, and an F1-score of 0.9807. Relatively weaker recognition performance was obtained for distress calls, with a precision of 96.71%, a recall of 95.58%, and an F1-score of 0.9643. The research findings can provide effective technical support for the automatic recognition of laying hen vocalizations.

    参考文献
    相似文献
    引证文献
引用本文

余礼根,丁晓丽,邱枫,张海庆,何金,李奇峰.基于TFViT-BiLSTM与Wav2Vec2.0的蛋鸡发声识别模型[J].农业机械学报,2026,57(16):317-326. Yu Ligen, Ding Xiaoli, Qiu Feng, Zhang Haiqing, He Jin, Li Qifeng. Recognition Model of Laying Hens' Vocalizations Based on TFViT-BiLSTM and Wav2Vec2.0[J]. Transactions of the Chinese Society for Agricultural Machinery,2026,57(16):317-326.

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-05-19
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-08-15
  • 出版日期:
文章二维码