• 融合边缘高斯聚合与图增强的遥感图像字幕生成

    Integrating edge gaussian aggregation and graph enhancement for remote sensing image captioning

    • 针对遥感图像字幕生成中因分辨率低、运动模糊及光照复杂等因素导致的边缘模糊、噪声显著和目标背景混淆等问题,提出一种融合边缘高斯聚合与图结构增强的字幕生成方法。为恢复弱对比度条件下的结构信息,在编码阶段构建基于方向自适应Scharr滤波器的边缘增强模块,以有效提取模糊区域中的细节特征;同时,设计基于高斯先验的特征细化器,通过可学习高斯核抑制噪声并规整特征响应分布,显著提升前景目标的显著性。为充分利用遥感对象间的空间上下文关系,引入图神经网络对多尺度特征进行结构增强,实现几何一致性引导的特征融合。在解码阶段,设计跨层残差特征聚合机制,融合编码器不同层级输出,以减少信息冗余并强化关键语义信息的传递效率。在Sydney、UCM和RSICD三个标准数据集上的实验结果表明,该方法能够有效提升字幕生成的描述质量与鲁棒性,CIDEr指标分别达255.36%、355.63%和124.23%,相较当前主流方法具有较强的性能优势,验证了所提方法的有效性与先进性。

       

      Abstract: To address the challenges of edge blurring, significant noise, and target-background confusion in remote sensing image caption generation caused by low resolution, motion blur, and complex illumination, this paper proposes a novel caption generation method that integrates edge Gaussian aggregation with graph structure enhancement. To recover structural information under low-contrast conditions, an edge enhancement module based on an orientation-adaptive Scharr filter is constructed in the encoding stage, effectively extracting detailed features from blurred regions. Meanwhile, a Gaussian prior-based feature refiner is designed to suppress noise and regularize the feature response distribution through learnable Gaussian kernels, significantly improving the saliency of foreground targets. To fully exploit the spatial contextual relationships among remote sensing objects, a graph neural network is introduced to perform structural enhancement on multi-scale features, achieving geometrically consistent feature fusion. In the decoding stage, a cross-layer residual feature aggregation mechanism is developed to fuse outputs from different encoder layers, reducing information redundancy and enhancing the transmission efficiency of key semantic information. Experimental results on three standard datasets—Sydney, UCM, and RSICD—demonstrate that the proposed method effectively improves the description quality and robustness of caption generation, achieving CIDEr scores of 255.36%, 355.63%, and 124.23%, respectively. These results exhibit strong competitive advantages over current state-of-the-art methods, thereby verifying the effectiveness and advancement of the proposed approach.

       

    /

    返回文章
    返回