상세 보기
LLM 모델에 기반한 중국어 결과보어 구문의 자동추출 프레임워크 연구
- 김은하;
- 강병규
초록
This study proposes a large language model (LLM)–based framework for constructing a specialized dataset of Chinese resultative complement constructions (RCCs) to overcome the limitations of existing automatic extraction methods. Despite being a critical yet challenging category for second language learners, RCCs are difficult to identify accurately because their "verb + verb/adjective" surface structures often overlap with non-resultative compounds, leading to significant structural ambiguity. While conventional morphological analyzers like HanLP and Jieba, as well as corpora such as CCL and BCC, often struggle with high-precision extraction, this research introduces a model that utilizes linguistically informed diagnostic criteria to capture the defining properties of RCCs. By prioritizing precision and minimizing error rates across large-scale corpora, this framework moves beyond high-frequency pattern matching to enable the construction of a comprehensive, diverse RCC dataset. Ultimately, this work contributes a robust computational approach to Chinese linguistic research and provides a foundational resource for developing advanced pedagogical materials.
키워드
- 제목
- LLM 모델에 기반한 중국어 결과보어 구문의 자동추출 프레임워크 연구
- 제목 (타언어)
- Constructing a Large-Scale Chinese Resultative Complement Dataset Using Large Language Models
- 저자
- 김은하; 강병규
- 발행일
- 2026-02
- 유형
- Y
- 저널명
- 중국문학
- 권
- 126
- 페이지
- 383 ~ 414