상세 보기
초록
Large-scale error annotation in learner corpora remains challenging, particularly for under-resourced L1 populations. Although Chinese learners represent the world’s largest group of English learners, publicly accessible, contemporary corpora of their English writing with comprehensive error annotations are scarce, limiting both research and pedagogical applications. This study introduces a human-in-the-loop, AI-assisted framework designed to deliver reliable annotations while also generating pedagogical feedback. The framework employs a streamlined three-domain taxonomy (Grammar, Lexis, and Mechanics) as well as a minimal four-field record structure (error type, span, correction, and explanation). Applied to Ten-thousand English Compositions of Chinese Learners (the TECCL corpus), the annotation pipeline integrates in-context demonstrations from the Chinese Learner English Corpus, chain-of-thought reasoning, and retrieval-augmented generation. In a proof-of-concept evaluation of 100 essays, inter-annotator agreement reached Cohen’s κ = 0.78 overall, with an exact-span match of 85%, and intra-annotator agreement exceeded 90%. Analysis of 164 sentences yielded 187 errors distributed as follows: Grammar 66.3%, Lexis 23.0%, and Mechanics 10.7%. Expert review confirmed high precision for morphosyntactic and lexical errors and identified nine omissions (recall = 95.4%), primarily involving idiomaticity or collocation. The pilot results suggest that the approach can support transparent scaling while preserving research comparability and producing materials suitable for classroom use. Although further validation with larger datasets is required, the framework shows potential for adaptation to other learner populations.
키워드
- 제목
- An AI-Assisted Framework for Error Annotation in Chinese Learner English Corpora
- 저자
- 강병규; Jiangli Wei; 유원호
- 발행일
- 2026-06
- 유형
- Y
- 권
- 23
- 호
- 2
- 페이지
- 431 ~ 448