CMVDE: Consistent Multi-View Video Depth Estimation via Geometric-Temporal Coupling Approach

  • Son, Hosung
  • Shin, Min-jung
  • Cho, Minji
  • Kim, Joonsoo
  • Yun, Kug-jin
  • ... Kang, Suk-Ju
Citations

WEB OF SCIENCE

1
Citations

SCOPUS

1

초록

In the field of video depth estimation, significant strides have been made with deep learning-based multi-view stereo approaches. However, existing studies struggle to produce consistently accurate depth maps that account for both multi-view geometry and temporal consistency from monocular video contents. To overcome this limitation, we introduce CMVDE, an innovative video depth estimation framework that leverages a multi-view geometric-temporal coupling approach in an end-to-end manner. Our proposed geometric consistency module efficiently generates multi-view geometric features by employing mutual cross-view epipolar attention between adjacent video frames. Additionally, it compresses these features using the novel multi-scale feature compressor, producing an effective input tensor for the subsequent module. Moreover, our framework enhances temporal consistency across consecutive video frames with the temporal consistency module based on convolutional LSTM [1] leveraging previous depth information as geometric guidance. Compared to state-of-the-art models, our approach achieves superior performance in depth quality and consecutive consistency on the ScanNet [2] and 7-Scenes [3] datasets, surpassing previous multi-view video depth estimation methods.

키워드

EstimationFeature extractionCostsThree-dimensional displaysGeometryCouplingsComputational efficiencyEpipolar attentionmulti-view stereomulti-view geometric-temporal couplingvideo depth estimation
제목
CMVDE: Consistent Multi-View Video Depth Estimation via Geometric-Temporal Coupling Approach
저자
Son, HosungShin, Min-jungCho, MinjiKim, JoonsooYun, Kug-jinKang, Suk-Ju
DOI
10.1109/TMM.2024.3397080
발행일
2024
유형
Article
저널명
IEEE Transactions on Multimedia
26
페이지
9710 ~ 9721