ViViT-HH-SupCon: 고차원 특징 추출(High-Dimensional Feature Extraction Architecture)과 후킹(Hooking) 기반 생성형 AI 비디오 엔진 식별 연구

ViViT-HH-SupCon: Generative AI Video Engine Identification via High-Dimensional Feature Extraction Architecture and Hooking

초록

This study proposes the ViViT-HH-SupCon pipeline architecture, which integrates a spatiotemporal high-dimensional featureextraction structure with a high-dimensional raw feature direct hooking mechanism within the encoder, for multi-originidentification of advanced generative AI videos. Existing low-dimensional embedding-based learning models analyze onlyframe-level spatial noise, which leads to the loss of subtle artifacts and degraded identification accuracy. To address this limitation,the proposed framework maps a high-dimensional feature space based on a ViViT encoder that chronologically integrates inputdata, and utilizes a hooking mechanism designed to directly extract high-dimensional raw features prior to the advancedcompression and abstraction stages of the final output layer. Experimental results across 8 state-of-the-art generative AI videoengines demonstrate that the proposed model achieves a macro F1-score of 0.9220 and an overall accuracy of 93.30%, yieldingperformance margins of 25.83%p and 32.50%p, respectively, over the baseline. Notably, it enhances the precision for previouslylower-performing engines, Veo3 and Vidu_Q1, to 0.9877 and 0.9800, successfully mitigating misclassification patterns.

키워드

Generative AI Video IdentificationVideo ForensicsHigh-Dimensional Feature ExtractionViViT-SupConFingerprint Preservation
제목
ViViT-HH-SupCon: 고차원 특징 추출(High-Dimensional Feature Extraction Architecture)과 후킹(Hooking) 기반 생성형 AI 비디오 엔진 식별 연구
제목 (타언어)
ViViT-HH-SupCon: Generative AI Video Engine Identification via High-Dimensional Feature Extraction Architecture and Hooking
저자
강자원도경화박수용
발행일
2026-07
유형
Y
저널명
방송공학회 논문지
31
4
페이지
709 ~ 721