Abstract:Aiming to address the challenge of identifying water-deficit conditions in intelligent irrigation for greenhouse-grown lettuce using soil-based cultivation, a multi-modal information fusion algorithm was proposed based on the Transformer self-attention mechanism and causal inference. A model framework was developed to fuse image and time-series features across modalities. Firstly, heterogeneous classification features were extracted from processed images and time-series environmental data. Then, a multi-modal attention fusion module was introduced to achieve deep integration of the two modalities. On this basis, a causal inference-based model distillation technique was employed to reduce noise in the image features, thereby enhancing the neural network's ability to detect water-deficit conditions in lettuce. To evaluate the model's performance, a multi-modal dataset comprising greenhouse lettuce images and environmental data under various water-deficit conditions was constructed. Experimental results demonstrated that the proposed model significantly outperformed mainstream architectures such as ResNet, DenseNet, EfficientNet, ViT, Swin, ConvNeXt, MambaVision, and their respective variants in terms of classification accuracy, transfer learning capability, and generalization. The model achieved an identification accuracy of 94.08%, a precision of 94.81%, and a recall of 94.52% on the test set. The research result can provide practical insights into efficient water management and intelligent irrigation strategies in greenhouse settings, offering notable application value.