DGTFNet: Depth-Guided Tri-Axial Fusion Network for Efficient Generalizable Stereo Matching

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Stereo matching is a crucial task in computer vision that estimates pixel-level disparities from rectified image pairs to reconstruct three-dimensional depth information. It has diverse applications, ranging from augmented reality to autonomous driving. While deep learning-based methods have achieved remarkable progress through 3D CNNs and Transformer-based architectures, their reliance on domain-specific fine-tuning and localized feature extraction often hampers robustness and generalization in real-world scenarios. This letter introduces the Depth-Guided Tri-Axial Fusion Network (DGTFNet), which overcomes these limitations by integrating depth priors from a monocular depth foundation model via the Depth-Guided Cross-Modal Attention (DGCMA) module. Additionally, we propose a Tri-Axial Attention (TAA) module that employs directional strip convolutions to capture long-range dependencies across horizontal, vertical, and spatial dimensions. Extensive evaluations on public stereo benchmarks demonstrate that DGTFNet significantly outperforms state-of-the-art methods in zero-shot evaluations. Ablation studies further validate the contribution of each module in delivering robust and efficient stereo matching.

키워드

Feature extractionCostsFoundation modelsThree-dimensional displaysFusesComputer architectureRobustnessTransformersTrainingFilteringDeep learning for visual perceptioncomputer vision for automationRGB-D perceptionstereo matching
제목
DGTFNet: Depth-Guided Tri-Axial Fusion Network for Efficient Generalizable Stereo Matching
저자
Moon, SeunghunLee, Hae ukKang, Suk-Ju
DOI
10.1109/LRA.2025.3606382
발행일
2025-10
유형
Article
저널명
IEEE Robotics and Automation Letters
10
10
페이지
10791 ~ 10798