Bridging Ensemble Performance and Transparency: Knowledge Distillation for Obesity Classification on KNHANES Dataset

  • Kim, Kyungjin; 
  • Lee, Youngro; 
  • Seo, Jongmo; 
  • Oh, Sewon

초록

Obesity is a pressing global health challenge that necessitates robust and interpretable models for predicting Body Mass Index (BMI). In this study, we utilized data from the Korea National Health and Nutrition Examination Survey (KNHANES) to evaluate multiple machine learning models for both BMI regression and binary classification (using threshold 25 kg/m). Among the tested models, the XGBRegressor achieved the highest area under the curve (AUC) in binary classification. However, interpretability remains critical in clinical applications. To enhance transparency, we employed a knowledge distillation approach, using the XGBRegressor as the teacher model and training a single DecisionTreeRegressor as the student model. This distillation process significantly improved the decision tree's performance compared to training it directly on the dataset. Furthermore, the distilled model enabled interpretable, rule-based predictions, highlighting key obesity-related features such as insulin resistance (HOMA-IR).

제목
Bridging Ensemble Performance and Transparency: Knowledge Distillation for Obesity Classification on KNHANES Dataset
저자
Kim, Kyungjin; Lee, Youngro; Seo, Jongmo; Oh, Sewon
DOI
10.1109/EMBC58623.2025.11251879
발행일
2025-07-14
학회명
47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC 2025)
개최지
Copenhagen, Denmark
개최국가
덴마크
학회 개최일
2025-07-14 ~ 2025-07-17