Multimodal Remote Sensing Image Segmentation via Foundation Model Domain Adaptation
-
Abstract
Multimodal remote sensing imagery provides richer and more comprehensive information for remote sensing scene understanding than single-modality data owing to the complementary nature of the cross-modal information derived from heterogeneous features. By integrating multiple sensing modalities, such as optical imagery and auxiliary data sources, multimodal approaches capture diverse characteristics of ground objects, thereby improving the robustness and accuracy of scene interpretation tasks. However, acquiring large-scale annotated multimodal remote sensing datasets requires substantial manual effort and time-consuming labeling processes, which significantly limit the availability of labeled data in practical applications. Consequently, the development of effective cross-domain segmentation methods for multimodal remote sensing imagery has become an urgent research problem, particularly in scenarios where labeled samples are available only in the source domain while the target domain remains unlabeled. To further enhance the cross-domain adaptation capability of visual foundation models and address the limitation that most are pretrained on single-modality data, this study proposes a multimodal low-rank domain adaptation network for segmentation, termed MLDANet. In the proposed framework, a domain-adaptive low-rank adaptation module, referred to as DALoRA, is introduced to improve domain generalization. During adversarial domain adaptation training, DALoRA performs domain-specific adaptation through two separate low-rank matrices, which are designed to capture the domain-dependent feature transformations for the source and target domains, respectively. This design effectively alleviates the model’s reliance on source-domain segmentation knowledge while strengthening its feature extraction capability for target-domain images. As a result, the network achieves improved adaptability across domains while maintaining a relatively low training cost. In addition, a modality difference-guided injector (MDI) is proposed to enhance the multimodal information integration within the foundation model. Using inter-modal feature discrepancies as the guiding criterion, the MDI employs an attention mechanism to compute discrepancy-aware weights and injects additional modality information that is beneficial for cross-domain segmentation. Through the fusion of modality feature differences, the injector enables the base model to acquire multimodal representation capabilities that were not present in its original pretraining process. This mechanism effectively compensates for the modality limitation of conventional foundation models and improves the model’s ability to exploit complementary information across different modalities. Extensive comparative experiments demonstrated the effectiveness of the proposed method. In the Potsdam_RGB to Vaihingen cross-domain transfer experiment, MLDANet achieved improvements of 1.93 percentage points in the mean intersection over union (mIoU), 1.75 percentage points in the mean F1-score, and 4.69 percentage points in the overall accuracy compared with the best-performing existing cross-domain segmentation methods. The proposed method further improved the mIoU by 2.54 and 1.14 percentage points in two additional cross-domain experimental settings, consistently outperforming current state-of-the-art approaches in segmentation performance. Furthermore, ablation studies demonstrated that DALoRA enhances the cross-domain adaptation capability of the model and that the MDI strengthens the multimodal feature fusion, jointly improving the segmentation performance of MLDANet for multimodal remote sensing images in scenarios where target-domain annotations are unavailable.
-
-