근대 잡지 <개벽>의 데이터과학적 언어 분석

Data scientific Language Analysis of a Modern Periodical GaeByeok

초록

This paper presents a data-scientific analysis of the entire corpus of the modern Korean periodical Gaebyeok, examining both metadata and main article texts. Metadata analysis reveals that unspecified pseudonyms account for about 35% of entries—mostly announcements or miscellaneous items not essential for identifying pen names—despite the periodical’s history of forced discontinuation and author censorship. Main text analysis, using deep learning techniques— subword tokenization, fastText static embeddings, and BERTopic dynamic embeddings—produced notable results. Subword tokenization effectively segmented tokens such as ‘아니-하면,’ ‘무엇-을,’ ‘경제-적,’ and ‘생각-을’ without a conventional morphological analyzer. fastText captured subwords as character n-grams in modern Korean texts with orthographic differences from contemporary Korean, positioning functionally related words such as ‘잇다,’ ‘업다,’ ‘하다,’ and ‘것이다’ in close semantic space. BERTopic identified latent topics, with top clusters including ‘people, arts’, ‘thoughts, mind’, ‘Gaebyeok, readers’, ‘song, sound’, ‘school, Gyeongseong’, ‘Joseon, society’, ‘Japanese, Korean, and ‘conference, government’. An important subcluster on ‘Cheondokyo’, the founding religion of the periodical, should also be incorporated into the topic results.

키워드

개벽필명단어조각정적 임베딩동적 임베딩GaebyeokPseudonymsubwordstatic embeddingsdynamic embeddings
제목
근대 잡지 <개벽>의 데이터과학적 언어 분석
제목 (타언어)
Data scientific Language Analysis of a Modern Periodical GaeByeok
저자
조은경Kate McDowell
DOI
10.29211/soli.2025.56..003
발행일
2025-11
유형
Y
저널명
언어와 정보 사회
56
페이지
55 ~ 82