AI 데이터 품질 관리를 위한 텍스트 구조적 무결성 검증 연구: T-Cohesion Index의 도메인 간 일관성에 관한 초기 검증

A Study on Structural Integrity Verification of Text for AI Data Quality Management: A Preliminary Cross-Domain Validation of the T-Cohesion Index

초록

This study proposes the T-Cohesion Index (T-Index), a new metric that measures the structural integrity of text to ensure the reliability of unstructured data in the era of AI Transformation (AX). To mitigate the performance degradation that content-based detection models suffer when overfitting to specific vocabulary or domains, the study focuses on the geometric connectivity of word co-occurrence networks. The T-Index integrates five topological properties—density, average path length, clustering coefficient, number of connected components, and isolation ratio—to quantify structural fragmentation. Across 819 news documents from two domains, politics (PolitiFact) and health (CoAID), the structural signal significantly distinguished normal from polluted data in both domains (AUC 0.637 and 0.718; both p<0.001; Cohen's d 0.46 and 0.93). However, transferring a threshold calibrated on one domain directly to the other reduced recall to 0.245, indicating that while the structural discriminative signal is consistent across domains, the operating threshold requires domain-specific calibration. These results provide a preliminary cross-domain validation of the T-Index as a pre-screening indicator—rather than a final classifier—for large-scale data pipelines.

키워드

AI Data Quality ManagementStructural IntegrityT-Cohesion IndexInformation Pollution DetectionCross-Domain ValidationAI 데이터 품질 관리구조적 무결성T-Cohesion 지수정보 오염 탐지도메인 간 일관성
제목
AI 데이터 품질 관리를 위한 텍스트 구조적 무결성 검증 연구: T-Cohesion Index의 도메인 간 일관성에 관한 초기 검증
제목 (타언어)
A Study on Structural Integrity Verification of Text for AI Data Quality Management: A Preliminary Cross-Domain Validation of the T-Cohesion Index
저자
조윤나홍원경김진화
DOI
10.36498/kbigdt.2026.11.1.221
발행일
2026-06
유형
Y
저널명
The Korea Journal of BigData
11
1
페이지
221 ~ 235