Skip to main navigation Skip to search Skip to main content

STCAFormer: Spatio-temporal cross attention based transformer for fusion of remote sensing images

  • Jun Li
  • , Xiaowei He
  • , Yantai Yang
  • , Qinghong Sheng*
  • , Zhaocong Wu
  • , Bo Wang
  • , Xiao Ling
  • , Xiang Liu
  • , Yu Liu
  • , Fan Gao
  • , Libo Wang
  • , Matthieu Molinier
  • *Corresponding author for this work
  • Nanjing University of Aeronautics and Astronautics
  • Wuhan University
  • PLA Army Engineering University
  • Nanjing University of Information Science & Technology

Research output: Contribution to journalArticleScientificpeer-review

Abstract

Spatio-temporal fusion is an important technique used to address the trade-off between temporal and spatial resolution in remote sensing satellite images. However, existing deep learning-based methods often rely on a single fusion approach, which limits their ability to effectively balance the preservation of spatial details in fine images with the integration of temporal information from coarse images. In this paper, we proposed a hybrid model architecture STCAFormer which combines temporal and spatial fusion attention mechanism, leveraging their complementary strengths to maximize the utilization of temporal and spatial information from coarse and fine images. Specifically, we design a dual-input encoder structure that processes two input images simultaneously. Moreover, inspired by the unified workflow of traditional methods that implement spatio-temporal fusion in pixel space, STCAFormer introduces a fusion stage designed for spatio-temporal fusions in feature space, enabling the effective integration of spatial and temporal features from dual-input encoder. Three novel blocks are designed in this research: a swin-temporal fusion block utilizing cross-attention mechanism to capture and fuse temporal variation features, a multi-scale spatial fusion block that integrates weighted summation for temporal variation and fine spatial features at both local and global scales, and a dual-path up-sampling block designed to achieve high-fidelity image restoration. STCAFormer was compared with six different types of spatio-temporal fusion methods on remote sensing image pairs, including Landsat 8/MODIS and Sentinel-2/3 pairs, from various regions. Experimental results demonstrate that STCAFormer performs better than baseline methods when the regions and satellite sensors remain unchanged. The transfer ability of STCAFormer was also proved to be more competitive than baseline methods. The ablation studies show that the proposed blocks are effective for performance improvement of spatio-temporal fusion The proposed STCAFormer demonstrates robust generalization across Landsat 8/MODIS and Sentinel-2/3, as well as across land-cover types represented in the test datasets. The code of the proposed STCAFormer is available at: https://github.com/Neooolee/STCAFormer.

Original languageEnglish
Article number115594
JournalRemote Sensing of Environment
Volume345
DOIs
Publication statusPublished - Nov 2026
MoE publication typeA1 Journal article-refereed

Funding

This work was supported in part by the National Natural Science Foundation of China under Grant 42301384, Grant 42271448 and Grant 42301489; in part by the “Fourteenth Five-Year” National Key Research and Development Program of China under No. 2024YFC390650; in part by the Natural Science Foundation of Jiangsu Province under Grant BK20220888 and Grant BK20231030; in part by the preresearch project on Civil Aerospace Technologies no. D040307 funded by China's National Space Administration; in part by the Academy of Finland through the Finnish Flagship Programme FCAI: Finnish Center for Artificial Intelligence (Grant No. 320183). The numerical calculations in this paper have been done on the supercomputing system in the Supercomputing Center of Wuhan University. This work is partially supported by High Performance Computing Platform of Nanjing University of Aeronautics and Astronautics.

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 15 - Life on Land
    SDG 15 Life on Land

Keywords

  • Cross attention
  • Deep learning
  • Remote sensing
  • Spatio-temporal fusion
  • Transformer

Fingerprint

Dive into the research topics of 'STCAFormer: Spatio-temporal cross attention based transformer for fusion of remote sensing images'. Together they form a unique fingerprint.

Cite this