Abstract:
To address the challenges of edge blurring, significant noise, and target-background confusion in remote sensing image caption generation caused by low resolution, motion blur, and complex illumination, this paper proposes a novel caption generation method that integrates edge Gaussian aggregation with graph structure enhancement. To recover structural information under low-contrast conditions, an edge enhancement module based on an orientation-adaptive Scharr filter is constructed in the encoding stage, effectively extracting detailed features from blurred regions. Meanwhile, a Gaussian prior-based feature refiner is designed to suppress noise and regularize the feature response distribution through learnable Gaussian kernels, significantly improving the saliency of foreground targets. To fully exploit the spatial contextual relationships among remote sensing objects, a graph neural network is introduced to perform structural enhancement on multi-scale features, achieving geometrically consistent feature fusion. In the decoding stage, a cross-layer residual feature aggregation mechanism is developed to fuse outputs from different encoder layers, reducing information redundancy and enhancing the transmission efficiency of key semantic information. Experimental results on three standard datasets—Sydney, UCM, and RSICD—demonstrate that the proposed method effectively improves the description quality and robustness of caption generation, achieving CIDEr scores of 255.36%, 355.63%, and 124.23%, respectively. These results exhibit strong competitive advantages over current state-of-the-art methods, thereby verifying the effectiveness and advancement of the proposed approach.