On-premises open-source large language models for privacy-preserving multimodal depression screening

  • Kwon, Soonjun; 
  • Kim, Yihyun; 
  • Jhon, Min; 
  • Park, Jin-Hyun; 
  • Lim, Bahngtaik; 
  • ... Lee, Hwamin; 
  • 외 3명
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Objective: Narrative and speech data can provide valuable signals for depression screening, yet privacy and datagovernance requirements often limit the use of closed-based models in clinical practice. In addition, existing large language model (LLM)-based approaches are largely text-centric, and multimodal integration of acoustic features and structured clinical variables remains limited. This study aimed to develop and externally validate a privacy-preserving multimodal depression screening prediction framework using open-source large language models that integrate sociodemographic information, emotion-memory narratives, and acoustic features. Method: This study analyzed 3536 participants collected at Chonnam National University Hospital. Inputs combined sociodemographic and lifestyle variables, Korean transcripts of happy- and sad-memory narratives, and speech-derived extended Geneva minimalistic acoustic parameter set (eGeMAPS) features. To maintain prompt conciseness, statistically significant features were selected from the 88 eGeMAPS features extracted for each happy- and sad-memory narrative, with Mann-Whitney U tests conducted exclusively on the internal training split to prevent data leakage. Five open-source LLMs (Gemma-3-27B, Qwen-3-32B, Llama-3.3-70B, Phi414B, and gpt-oss-20b) were evaluated under zero-shot prompting, Chain-of-Thought prompting, and supervised fine-tuning. External validation used Extended Distress Analysis Interview Corpus (E-DAIC) (N = 275). Results: Under zero-shot prompting, the best internal F1-score was 0.735 (Gemma-3-27B). Chain-of-Thought prompting improved Llama-3.3-70B (F1-score = 0.708) but reduced performance for other models. Supervised fine-tuning improved all models, yielding internal accuracies of 0.852 to 0.881 and F1-scores of 0.818 to 0.865 across five models, corresponding to F1 gains of 0.12 to 0.30 versus prompting-only approaches. In external validation, accuracy ranged from 0.764 to 0.822 and F1-score ranged from 0.683 to 0.807. Conclusion: This study suggests that multimodal open-source LLMs integrating clinical variables, narrative text, and acoustic features can support privacy-preserving depression screening in an on-premises setting. Supervised fine-tuning provided the most consistent performance improvements, and external validation supported robustness beyond the development cohort.

키워드

Depression Screening; Open-source LLMs; Multi-modal; Digital phenotyping; SEVERITY
제목
On-premises open-source large language models for privacy-preserving multimodal depression screening
저자
Kwon, Soonjun; Kim, Yihyun; Jhon, Min; Park, Jin-Hyun; Lim, Bahngtaik; Jeon, Eunkyoung; Kim, Jae-Min; Kim, Ju-Wan; Lee, Hwamin
DOI
10.1016/j.ijmedinf.2026.106577
발행일
2026-11
유형
Article
저널명
International Journal of Medical Informatics
권
220