Abstract:
Hyperspectral and LiDAR fusion classification techniques can realize high-precision classification of features. Currently, supervised traditional and deep learning methods have achieved better classification results, but often require a large number of labeled samples. There are relatively few studies on self-supervised multimodal remote sensing fusion classification based on self-supervision, and the existing self-supervised contrast learning framework uses data augmentation to generate positive sample pairs, which is not applicable to multimodal remote sensing images, and will destroy the spatial distribution and semantic similarity between multimodal data, and the model is too complex to be conducive to the generalization of the downstream fine-tuning task. Therefore, a multimodal remote sensing image fusion classification based on contrastive learning network (MMCLNet) is proposed, which is different from the traditional contrast learning network in that it can fully utilize a large amount of unlabeled data in the pre-training stage without data enhancement operations to learn the discriminative feature representations. At the same time, the well-designed two-branch network reduces the complexity of the network. Moreover, it adopts the multilevel feature fusion in the fine-tuning stage. In addition, a multi-level feature fusion network is used in the fine-tuning stage to fully integrate the heterogeneous features of the two modal data. Extensive experiments using three real multimodal remote sensing image fusion classification datasets demonstrate the advantages of the proposed research method on datasets with a small number of labeled samples.