VHOIP: Video-based Human-Object Interaction recognition with CLIP Prior knowledge

Citations

WEB OF SCIENCE

1
Citations

SCOPUS

3

초록

In this paper, we introduce a novel approach to recognizing Human-Object Interactions (HOI) in videos, crucial for understanding videos focused on human activities. Traditional methods often fall short of accurately identifying subtle interactions, particularly in dynamic sequences involving multiple individuals and objects. To address these issues, we leverage the CLIP (Contrastive Language-Image Pre-training), renowned for its rich visual and linguistic knowledge. Our method, Video-based HOI recognition with CLIP Prior knowledge (VHOIP), merges the spatial and temporal analysis capabilities of a video-based HOI framework with the detailed interaction understanding from CLIP. This enhancement significantly advances our HOI recognition performances. Through rigorous validation of three different HOI recognition datasets, our method demonstrates remarkable improvements over current state-of-the-art techniques, both qualitatively and quantitatively, indicating the effectiveness of our approach.

키워드

Deep learningRepresentation learningComputer visionHuman-object interaction recognitionAFFORDANCES
제목
VHOIP: Video-based Human-Object Interaction recognition with CLIP Prior knowledge
저자
Baek, DoyeolChoe, Junsuk
DOI
10.1016/j.patrec.2025.02.014
발행일
2025-04
유형
Article
저널명
Pattern Recognition Letters
190
페이지
133 ~ 140