基础模型域适应的多模态遥感图像分割

Multimodal Remote Sensing Image Segmentation via Foundation Model Domain Adaptation

  • 摘要: 多模态遥感图像因其异质特征的跨模态信息互补性,相比单模态信息在遥感场景解析中能提供更丰富的信息,但获取有标注的多模态遥感图像需要大量人工成本,因此急需对多模态遥感图像的跨域分割方法进行研究。为了进一步增强视觉基础模型的跨域适应能力与针对多数基础模型预训练模态单一的问题,本文提出了基础模型低秩适应与域适应的多模态分割网络(Multimodal Low-rank Domain Adaptation Network,MLDANet)。在该网络中,用于域适应的低秩适应(Domain-Adaptive Low-Rank Adaptation,DALoRA)在域对抗训练时通过两个低秩矩阵分别进行不同域的适应,以更少的训练成本提高网络的适应性,提高网络的特征提取能力。另外,所提出的特征差异引导的注入器(Modality Difference-guided Injector,MDI),以特征差异为基准,借助注意力机制计算差异权重,为基础模型注入更有助于跨域分割的额外模态信息,通过模态特征差异融合的方式,使基础模型具备预训练中缺失的多模态表征能力。对比实验结果表明,MLDANet在Potsdam_RGB迁移至Vaigingen对比实验中,相对于当前跨域分割方法的最优指标,平均交并比(mean Intersection over Union,mIoU)高出1.93个百分点,平均F1分数高出1.75个百分点,总体精度高出4.69个百分点;而在另外两组跨域对比实验中,mIoU分别比当前最优方法高出了2.54和1.14个百分点,分割表现均优于当前最先进的算法。而消融实验结果说明DALoRA提高了模型的跨域适应能力与MDI增强了模型的多模态融合能力,进而增强了MLDANet在目标域图像无标签情况下的多模态遥感图像分割能力。

     

    Abstract: Multimodal remote sensing imagery provides richer and more comprehensive information for remote sensing scene understanding than single-modality data owing to the complementary nature of the cross-modal information derived from heterogeneous features. By integrating multiple sensing modalities, such as optical imagery and auxiliary data sources, multimodal approaches capture diverse characteristics of ground objects, thereby improving the robustness and accuracy of scene interpretation tasks. However, acquiring large-scale annotated multimodal remote sensing datasets requires substantial manual effort and time-consuming labeling processes, which significantly limit the availability of labeled data in practical applications. Consequently, the development of effective cross-domain segmentation methods for multimodal remote sensing imagery has become an urgent research problem, particularly in scenarios where labeled samples are available only in the source domain while the target domain remains unlabeled. To further enhance the cross-domain adaptation capability of visual foundation models and address the limitation that most are pretrained on single-modality data, this study proposes a multimodal low-rank domain adaptation network for segmentation, termed MLDANet. In the proposed framework, a domain-adaptive low-rank adaptation module, referred to as DALoRA, is introduced to improve domain generalization. During adversarial domain adaptation training, DALoRA performs domain-specific adaptation through two separate low-rank matrices, which are designed to capture the domain-dependent feature transformations for the source and target domains, respectively. This design effectively alleviates the model’s reliance on source-domain segmentation knowledge while strengthening its feature extraction capability for target-domain images. As a result, the network achieves improved adaptability across domains while maintaining a relatively low training cost. In addition, a modality difference-guided injector (MDI) is proposed to enhance the multimodal information integration within the foundation model. Using inter-modal feature discrepancies as the guiding criterion, the MDI employs an attention mechanism to compute discrepancy-aware weights and injects additional modality information that is beneficial for cross-domain segmentation. Through the fusion of modality feature differences, the injector enables the base model to acquire multimodal representation capabilities that were not present in its original pretraining process. This mechanism effectively compensates for the modality limitation of conventional foundation models and improves the model’s ability to exploit complementary information across different modalities. Extensive comparative experiments demonstrated the effectiveness of the proposed method. In the Potsdam_RGB to Vaihingen cross-domain transfer experiment, MLDANet achieved improvements of 1.93 percentage points in the mean intersection over union (mIoU), 1.75 percentage points in the mean F1-score, and 4.69 percentage points in the overall accuracy compared with the best-performing existing cross-domain segmentation methods. The proposed method further improved the mIoU by 2.54 and 1.14 percentage points in two additional cross-domain experimental settings, consistently outperforming current state-of-the-art approaches in segmentation performance. Furthermore, ablation studies demonstrated that DALoRA enhances the cross-domain adaptation capability of the model and that the MDI strengthens the multimodal feature fusion, jointly improving the segmentation performance of MLDANet for multimodal remote sensing images in scenarios where target-domain annotations are unavailable.

     

/

返回文章
返回