Abstract:Aiming to address the challenges of detecting pear fruits in orchard environments, such as variable illumination, multi-scale targets, dense distributions, and severe occlusions, a pear detection method was proposed based on an improved YOLO 11 framework. The proposed approach integrated a self-supervised shape-aware feature enhancement (SS-SAFE) module into the base YOLO 11 network, which learned to predict the aspect ratio of targets through self-supervised learning, thereby strengthening the perception and extraction of distinctive pear-shape features. Additionally, an adaptive kernel convolution (AKConv) was used to replace the standard convolution in the original cross stage partial network with kernel size 2 (C3K2) module, forming a cross stage partial network with kernel size 2 and adaptive kernel convolution (C3K2-AKConv) deep feature extraction structure, improving the network's adaptability to pears of varying scales. Furthermore, the C2PSA (convolutional block with parallel spatial attention) module was replaced by the C2PSA-HSAN (cross-stage partial with hybrid spectral attention network) module, which optimizes the feature interaction and fusion mechanism, significantly enhancing the model's feature fusion ability in complex environments. Experimental results show that the improved model achieves a precision of 88.9%, recall of 81.7%, F1-score of 85.2%, mAP@0.5 of 88.6%, and mAP@0.5:0.95 of 67.2%, outperforming the baseline YOLO 11n by 3.8, 1.9, 2.7, 3.2, 2.9 percentage points, respectively. This study offers a technical reference for the intelligent management and automated harvesting in pear orchards.