Exploring impacts of media characteristics on message content using text mining

Citations

SCOPUS

0

초록

Background/Objectives: The purpose of this study is to empirically investigate the semantic similarity of documents posted on different forms of media about specific social issues. Methods/Statistical analysis: Online text data were collected from personal blogs and Internet news published on a major Korean portal site, NAVER. To collect text data from online media, the study used R programming language for web crawling. We examined what effects medium characteristics had on the content of conveyed messages by using a keyword extraction method based on TF-IDF, which is a text mining method, and the cosine similarity measurement method. Findings: The results of this study demonstrate that there were differences in the major keywords extracted from messages conveyed by the three forms of media, but the similarity between keyword-to-keyword matrices extracted from the media was confirmed by a Mantel test, and there were statistically significant degrees of similarity among these matrices. We were therefore able to discover similarities of message content conveyed by each medium. Improvements/Applications: For this study, we used only blog and news data published on a single Korean portal site. The text data better be collected from variety of channels in future studies. © 2020 SERSC.

키워드

Media characteristicOffline mediaOnline mediaSimilarity analysisText miningTF-IDF
제목
Exploring impacts of media characteristics on message content using text mining
저자
Baek, Seung IkKim, J.
발행일
2020
유형
Article
저널명
International Journal of Advanced Science and Technology
29
4 Special Issue
페이지
291 ~ 303