LGPose: unseen 6D pose estimation

  • Sun, Shantong
  • Bao, Xu
  • Wang, Yang
  • Gong, Pan
  • Park, Unsang
Citations

WEB OF SCIENCE

0

초록

Unseen object 6D pose estimation is one of the challenging problems in the field of computer vision. The Vision-Language Models (VLMs) extend the detector from unseen categories to seen categories through open vocabulary learning, which provides a new idea for unseen object pose estimation. However, the effective integration of six degrees of freedom information of objects with a vision language model remains an unresolved issue. In this work we present LGPose, a vision-language fusion network for unseen object 6D pose estimation. We incorporate information such as object position and orientation into the prompts to guide object segmentation and object pose estimation. In order to correlate the pose information between images and text, we adopt Transformer to extract image features and fuse the position embeddings of image patches with the text pose information. This helps to deeply integrate vision features with language features. Furthermore, inspired by motion estimation in the field of video compression, we propose a feature refinement model for pose perception, which can extract features that are beneficial for object pose estimation. Through comprehensive qualitative and quantitative experiments on two popular object pose estimation datasets, REAL275 and Toyota-Light, we demonstrate that LGPose outperforms recent unseen object 6D pose estimation approaches.

키워드

Object 6D pose estimationVision-language modelsPrompt guideFeature refinementTRACKING
제목
LGPose: unseen 6D pose estimation
저자
Sun, ShantongBao, XuWang, YangGong, PanPark, Unsang
DOI
10.1016/j.knosys.2026.116736
발행일
2026-10
유형
Article
저널명
Knowledge-Based Systems
351