데이터 특성에 따른 스태킹 모델의 유효성 연구

A Study on the Effectiveness of Stacking Model by Data Characteristic

초록

Purpose : This study attempts to explore the performance of various stacking models by applying into two kinds of data, which have very different characteristic. Methods : The Base model includes decision tree, random forest, Naive Bayes, logistic regression, whereas support vector machine is adopted as the Meta model. Two kinds of data are ‘hmeq’ data and ‘bankrupt’ data. The performance is measured by accuracy, sensitivity, specificity, false positive ratio, and false negative ratio. Results : For ‘hmeq’ data which are very well refined, random forest results super performance that all stacking models do not exceed. For ‘bankrupt’ data which are raw, containing much noise, stacking models perform, in overall, better than individual base models. Stacking model 5 outperforms, in particular. Conclusion : The empirical analysis results indicate that when data are highly refined and have a limited number of input variables, stacking approach seems not a good strategy. Choosing a highly hyperparameter-trainable model would be a good choice. On the other hand, when data are almost raw with many input variables, stacking can perform better than individual models by integrating individual model’s ability in reading underlying patterns.

키워드

Stacking modelDecision TreeRandom ForestNaive BayesLogitSVM
제목
데이터 특성에 따른 스태킹 모델의 유효성 연구
제목 (타언어)
A Study on the Effectiveness of Stacking Model by Data Characteristic
저자
조성빈
DOI
10.35373/KMES.30.2.2
발행일
2025-06
유형
Y
저널명
한국경영공학회지
30
2
페이지
17 ~ 30