Abstract:Aiming to tackle the issues of mutual overlap and occlusion of pitaya fruits under complex natural growth environments, which are primarily induced by drastic scale variation and dense distribution of fruit targets, an improved collaborative attention-based multi-feature fusion object detection model named SAMF-YOLO v10 was developed on the basis of the original YOLO v10 architecture.In the feature extraction process, the multi-scale feature fusion (MFC) module was used instead of the traditional convolution module, which can enhance the model's ability to express features of different scales and improve the capture of detailed information.Meanwhile, the synergistic effects between spatial and channel attention (SCSA) module was incorporated, it can capture the relationships between features at different spatial positions and in different directions to achieve collaborative enhancement of features, thereby improving the feature expression ability for occluded targets and small targets.The results showed that the SAMF-YOLO v10 model achieved a precision of 90.18%, a recall of 86.28%, and an mAP@0.5 of 93.03% on the pitaya fruit dataset.Compared with the original YOLO v10 model, SAMF-YOLO v10 improved the precision, recall, and mAP@0.5 indicators by 3.17, 4.93, and 3.87 percentage points, respectively.The experimental results demonstrated that the newly proposed model was able to effectively integrate multi-level feature information and critical semantic content, which further enhanced and optimized the overall target detection performance for pitaya fruit images.