Factorized Spatiotemporal Attention Improves Representation Learning in Generalizable Radar Transformers
DOI:
https://doi.org/10.13021/jssr2026.5601Abstract
Millimeter-wave (mmWave) single-chip radars function even in poor weather conditions and can operate at a low cost, but they suffer from angular resolution constraints. Recently, architectures like the Generalizable Radar Transformer (GRT) have emerged as foundational models to infer 3D spatial representations directly from raw 4D radar tensor data. However, standard Transformer attention mechanisms in foundational radar models conflate multi-dimensional spatial features and temporal dynamics, leading to suboptimal feature representation across complex radar tensors. To address this limitation, we investigated a targeted architectural enhancement to the baseline GRT model: divided space-time factorized attention. Our design factorizes multi-dimensional self-attention into sequential spatial self-attention and temporal self-attention. To stabilize learning and improve gradient flow across long sequences, the architecture integrates learnable 1D temporal positional embeddings and adaptive gated residual connections. We evaluated our proposed architecture against the baseline GRT model on selected mmWave radar traces. Experimental evaluation demonstrates that factorizing spatiotemporal attention yields consistent performance improvements across all evaluation metrics: reducing binary cross-entropy (BCE) loss from 0.1365 to 0.1290 (~5.5%), improving 3D Chamfer distance from 1.2816 to 1.2322 (~3.9%), and lowering vertical (height) and range (depth) positioning errors from 13.1263 to 12.6042 (~4.0%) and 3.9556 to 3.8183 (~3.5%), respectively. These findings demonstrate that factorized attention structures provide superior geometric representation for mmWave radar perception.


