상세 보기
Comparing Human and LLM Evaluations on AI-Generated Critical Thinking Items: Implications for Valid Applications of Automatic Item Generation
- Kim, Euigyum;
- Khalil, Salah;
- Shin, Hyo Jeong
SCOPUS
0초록
As a core 21st-century skill, critical thinking (CT) has garnered increasing attention in today's information society. Although growing interest has led to the development of various CT assessments and frameworks, research on leveraging large language models (LLMs) for the automatic generation and validation of CT items remains limited. To address this gap, this study examines AI-generated CT items developed based on MACAT's PACIER framework. We employed a human-in-the-loop evaluation approach, in which a human expert and four LLMs independently rated each item on a three-point quality scale and conducted qualitative reviews to identify item-level issues. The results demonstrate marked differences between the human and LLM evaluations. The human reviewer delivered more discerning and variable evaluations, whereas the LLMs exhibited greater uniformity and consistency, but tended to be permissive and generous in their judgments. Notably, the human expert identified subtle flaws that the LLMs failed to detect, such as imprecise terminology, overly suggestive answer choices, and culturally biased content, all of which pose threats to the validity of the assessment. These insights affirm the essential role of human engagement in validating and optimizing the automatic item generation (AIG) process for complex latent constructs such as CT. © 2025 Copyright for this paper by its authors.
키워드
- 제목
- Comparing Human and LLM Evaluations on AI-Generated Critical Thinking Items: Implications for Valid Applications of Automatic Item Generation
- 저자
- Kim, Euigyum; Khalil, Salah; Shin, Hyo Jeong
- 발행일
- 2025
- 유형
- Conference paper
- 저널명
- CEUR Workshop Proceedings
- 권
- 4006