상세 보기
자연어처리 기반 연구 제안서 유사도 분석 프레임워크 개발 및 적용
- 백진주;
- 김주람
초록
Duplicate submission of identical or semantically similar research proposals across institutions or programs undermines the credibility and efficiency of public research and development (R&D) funding. Manual screening has clear limits because proposals comprise largely unstructured text. Identifying substantially similar proposals despite paraphrasing requires considerable expertise and time. This paper proposes a Natural Language Processing (NLP)-based framework that automatically quantifies semantic similarity between proposals by combining sentence embeddings from the Korean sentence bidirectional encoder representations from transformers (Ko-SBERT) model with a weighted aggregation across key sections to produce a proposal-level score. In a case study using proposals actually submitted to government programs, we computed similarities for 4,371 proposal pairs and flagged 87 pairs that simultaneously fell within the top 10% of similarity and shared institutional affiliations. The review list recovered 10 of the 12 previously identified suspicious cases. A Mann-Whitney U test further showed that the suspected pairs had significantly higher similarity scores than the general pairs (p<0.001), supporting the validity of the approach. Beyond detecting potential duplicate proposals for funding, the framework provides section-level contribution scores to explain why pairs are flagged and can be extended to similar-proposal recommendation and automated topic classification for broader R&D administrative support.
키워드
- 제목
- 자연어처리 기반 연구 제안서 유사도 분석 프레임워크 개발 및 적용
- 제목 (타언어)
- Development and Application of a Natural Language Processing-Based Framework for Research Proposal Similarity Analysis
- 저자
- 백진주; 김주람
- 발행일
- 2026-02
- 유형
- Y
- 저널명
- 한국산학기술학회논문지
- 권
- 27
- 호
- 2
- 페이지
- 150 ~ 158