상세 보기
Qualitative Analysis of Discrepancy Patterns Between Large Language Models and Human Reviewers in Abstract Screening for Systematic Reviews
WEB OF SCIENCE
0초록
Background: We evaluated the reliability of large language models (LLMs) for abstract screening under real-world review practices and qualitatively characterized model-human discordance to inform safe workflow integration. Methods: We evaluated GPT-4.0, GPT-5.0, and GPT-5.0-mini on two curated systematic review datasets representing contrasting topic densities, defined by the target-to-background ratio (TBR): a core-subject dataset (TBR 69%) in which the target intervention was central, and a peripheral-subject dataset (TBR 4.4%) in which the target intervention was incidental. Final inclusion after full-text review served as the reference standard. We developed a qualitative taxonomy of disagreements, classifying false negatives as intended human leniency, gray-zone ambiguity, or true LLM misses, and false positives as implicit or additional human exclusion rules, gray-zone ambiguity, or nominal inclusions that increase workload only. Results: GPT-5.0-mini achieved the best sensitivity-efficiency trade-off (core-subject: 91% sensitivity with 96.7% workload reduction; peripheral-subject: 83% sensitivity with 92.7% workload reduction) and negative predictive value >99% in both datasets. Disagreement was lower when relevance was central (core-subject: 1.6%, 7/430) with no true LLM misses (0/430). In the peripheral-subject dataset, disagreement was higher (10.6%, 74/696), driven mainly by intended human leniency among false negatives (52/56) and gray-zone ambiguity among false positives (12/18), while true LLM misses remained rare (0.4%, 3/696). Conclusion: Many model-human disagreements reflect topic-and workflow-dependent screening conventions rather than intrinsic model failure. LLM-assisted screening may improve efficiency without compromising reliability when accompanied by appropriate safeguards for ambiguous records.
키워드
- 제목
- Qualitative Analysis of Discrepancy Patterns Between Large Language Models and Human Reviewers in Abstract Screening for Systematic Reviews
- 저자
- Lee, Kyung Hwa; Kim, Hakyoung; Yang, Dae Sik; Yoon, Won Sup; Rim, Chai Hong
- 발행일
- 2026
- 유형
- Article; Early Access
- 언어
- ENG
- 출판사
- Tokyo University of Agriculture
- 발행국가
- 일본
- ISSN
- E 1349-9092
P 0917-5040