Abstract / Summary
Mental health disorders affect 970 million people globally; yet, over 50% do not access timely evaluation due to structural barriers and professional shortages. Chatbots and AI-based conversational agents have emerged as promising tools for mental health screening and assessment.
This study systematically evaluated the effectiveness, accuracy, reliability, and acceptability of chatbots and AI-based conversational agents for mental health screening and assessment in adults.
Systematic search conducted in May 2025 across PubMed/MEDLINE, PsycINFO, Scopus, and Web of Science (2019-2025), following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 guidelines. Eligible studies evaluated chatbots or AI for mental health screening/assessment in adults (≥18 y). Risk of bias was assessed using appropriate tools (Risk of Bias 2, Quality Assessment of Diagnostic Accuracy Studies-2, Joanna Briggs Institute checklists, and Mixed Methods Appraisal Tool). This systematic review was registered with PROSPERO (International Prospective Register of Systematic Reviews; CRD420251072392).
Eighteen studies (2021-2025) were included, with samples ranging from 20 to 3902 participants. Rule-based chatbots demonstrated high reliability (Cronbach α >0.85) and good acceptability (Acceptability of Intervention Measure >19/25). Generative models (large language models) achieved sensitivities of 0.84 to 0.93 and specificities of 0.80 to 0.96 for depression and anxiety, with correlations up to r=0.96 with expert clinicians in suicide risk assessment. Hybrid approaches combining large language models with machine learning achieved exceptional performance for cognitive impairment (F1-score of 92.1%, specificity of 99.6%). Most studies reported high user satisfaction (≥70%), although barriers existed among older populations. Methodological quality was heterogeneous with a moderate risk of bias in critical dimensions.
Chatbots and AI conversational agents demonstrate clinically relevant performance in mental health screening and assessment. However, safe implementation requires clear clinical protocols, professional supervision, integration with electronic health records, and active mitigation of algorithmic bias. These technologies should complement rather than replace clinical judgment.