基于数据–物理双驱动的强化学习无人船航迹跟踪控制
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

国家自然科学基金项目(52271360、52571405)、兴辽英才计划项目(XLYC2403032)、辽宁省自然科学基金项目(2025-MS-093)和辽宁省科技计划联合计划项目(重点研发计划项目)(2025JH2/101800233)


Reinforcement Learning-based Trajectory Tracking Control for Unmanned Surface Vessels via Data – Physics Dual-driven Approach
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    无人船(Unmanned surface vessels,USVs)作为海洋智慧牧场的核心装备之一,在实际作业过程中需直面海洋环境随机扰动、自身物理约束及动力学参数未知等多重困难。为此,提出一种考虑输入约束与外界环境干扰的数据–物理双驱动强化学习最优控制方法。首先,针对无人船未知动力学系统建模难、传统数据驱动方法可解释性不足的难题,提出一种数据–物理双驱动(Data – physics dual-driven, DPD)的建模方法,构建融合船舶动力学物理定律与循环神经网络(Recurrent neural networks, RNN)数据学习能力的数据–物理双驱动模型,实现未知复杂动力学特性的精准重构与模型物理可解释性的增强。其次,采用反步法前馈控制(Feedforward control, FC)与最优控制(Optimal control, OC)方法相结合的思路设计控制框架,提出一种基于强化学习的优化控制方法,提升无人船与复杂海洋环境的持续交互能力,达到跟踪精度与控制资源损耗的动态平衡,同时在控制器设计中考虑了船体承载能力对控制输入的物理约束。最后,通过Lyapunov稳定性理论证明闭环系统内所有信号是一致最终有界的。通过二级海况下的对比仿真试验结果表明,所提方法的跟踪精度较物理机理驱动(Physics-principle-driven, PPD)方法整体提升42.6%,控制输入成本较PPD方法整体降低26.9%,较传统FC整体降低9.6%,充分验证了方法的有效性和优越性。

    Abstract:

    A data – physics dual-driven ( DPD) reinforcement learning ( RL)-based optimal control method considering input constraints and external environmental disturbances was proposed for unmanned surface vessels ( USVs) applied in intelligent marine ranching. Firstly, to address the challenges of modeling USVs' unknown dynamics and insufficient interpretability of traditional data-driven methods, a DPD model was constructed by integrating the physical laws of ship dynamics with the data learning capability of recurrent neural networks ( RNN), enabling accurate reconstruction of unknown complex dynamic characteristics and enhancing the model's physical interpretability. Secondly, a control framework was designed via the combination of backstepping and optimal control methods. An RL-based optimal control strategy was proposed to improve the USVs' continuous interaction capability with complex marine environments, achieving a dynamic balance between tracking accuracy and control resource consumption. Meanwhile, the physical constraints imposed by the hull's load-bearing capacity on control inputs were incorporated into the controller design. Thirdly, the uniform ultimate boundedness (UUB) of all signals in the closed-loop system was proved by using Lyapunov stability theory. Finally, comparative simulation results under Level 2 sea conditions demonstrated that the proposed method improved tracking accuracy by 42.6% compared with the physics-principle-driven (PPD) method, and reduced control input cost by 26.9% relative to the PPD method and 9.6% compared with traditional backstepping control, which fully verified the effectiveness and superiority of the proposed method.

    参考文献
    相似文献
    引证文献
引用本文

白伟伟,雷家虎,林玉芳,李荣辉.基于数据–物理双驱动的强化学习无人船航迹跟踪控制[J].农业机械学报,2026,57(19):54-62,202. Bai Weiwei, Lei Jiahu, Lin Yufang, Li Ronghui. Reinforcement Learning-based Trajectory Tracking Control for Unmanned Surface Vessels via Data – Physics Dual-driven Approach[J]. Transactions of the Chinese Society for Agricultural Machinery,2026,57(19):54-62,202.

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-03-16
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-10-01
  • 出版日期:
文章二维码