LLM 모델에 기반한 중국어 결과보어 구문의 자동추출 프레임워크 연구

Constructing a Large-Scale Chinese Resultative Complement Dataset Using Large Language Models

초록

This study proposes a large language model (LLM)–based framework for constructing a specialized dataset of Chinese resultative complement constructions (RCCs) to overcome the limitations of existing automatic extraction methods. Despite being a critical yet challenging category for second language learners, RCCs are difficult to identify accurately because their "verb + verb/adjective" surface structures often overlap with non-resultative compounds, leading to significant structural ambiguity. While conventional morphological analyzers like HanLP and Jieba, as well as corpora such as CCL and BCC, often struggle with high-precision extraction, this research introduces a model that utilizes linguistically informed diagnostic criteria to capture the defining properties of RCCs. By prioritizing precision and minimizing error rates across large-scale corpora, this framework moves beyond high-frequency pattern matching to enable the construction of a comprehensive, diverse RCC dataset. Ultimately, this work contributes a robust computational approach to Chinese linguistic research and provides a foundational resource for developing advanced pedagogical materials.

키워드

결과보어중국어 문법자동 추출대규모 언어모델(LLM)코퍼스 분석구조적 중의성복합술어트랜스포머(Transformer)mT5Geminiresultative complementsChinese grammarautomatic extractionlarge language models (LLMs)corpus analysisstructural ambiguitycomplex predicatesTransformermT5Gemini
제목
LLM 모델에 기반한 중국어 결과보어 구문의 자동추출 프레임워크 연구
제목 (타언어)
Constructing a Large-Scale Chinese Resultative Complement Dataset Using Large Language Models
저자
김은하강병규
DOI
10.21192/scll.126..202602.015
발행일
2026-02
유형
Y
저널명
중국문학
126
페이지
383 ~ 414