상세 보기
VHOIP: Video-based Human-Object Interaction recognition with CLIP Prior knowledge
- Baek, Doyeol;
- Choe, Junsuk
WEB OF SCIENCE
1SCOPUS
3초록
In this paper, we introduce a novel approach to recognizing Human-Object Interactions (HOI) in videos, crucial for understanding videos focused on human activities. Traditional methods often fall short of accurately identifying subtle interactions, particularly in dynamic sequences involving multiple individuals and objects. To address these issues, we leverage the CLIP (Contrastive Language-Image Pre-training), renowned for its rich visual and linguistic knowledge. Our method, Video-based HOI recognition with CLIP Prior knowledge (VHOIP), merges the spatial and temporal analysis capabilities of a video-based HOI framework with the detailed interaction understanding from CLIP. This enhancement significantly advances our HOI recognition performances. Through rigorous validation of three different HOI recognition datasets, our method demonstrates remarkable improvements over current state-of-the-art techniques, both qualitatively and quantitatively, indicating the effectiveness of our approach.
키워드
- 제목
- VHOIP: Video-based Human-Object Interaction recognition with CLIP Prior knowledge
- 저자
- Baek, Doyeol; Choe, Junsuk
- 발행일
- 2025-04
- 유형
- Article
- 권
- 190
- 페이지
- 133 ~ 140