UPM Institutional Repository

A counsellor-in-the-loop evaluation framework for multi-model assessment of LLM-generated mental health advisories


Citation

Shamshudeen, Shahrul Hazman and Mohd Sharef, Nurfadhlina and Yusoff, Muhamad Saiful Bahri (2026) A counsellor-in-the-loop evaluation framework for multi-model assessment of LLM-generated mental health advisories. Journal of Information and Knowledge Management. art. no. 2650046. ISSN 0219-6492; eISSN: 1793-6926 (In Press)

Abstract

The demand for scalable and empathetic mental health support is driving increased interest in the use of large language models (LLMs) as advisory tools. Very few studies have been published that show how LLMs perform psychologically and demonstrate cross-model variation. We introduce DASS21-EvaLLM, a counsellor-in-the-loop evaluation system as an advisory appropriateness screening instrument for DASS-21 integration with four prominent LLMs (ChatGPT, Gemini, LLaMA and Mistral). The DASS21-EvaLLM provides the ability to rate, annotate and compare responses within a single interface. Using 65 simulated cases of clients and 13 licensed counsellors’ assessments, we considered the advisory quality of LLMs based upon each client’s profile for depression, anxiety and stress according to three specific criteria (accuracy, empathy and clarity), including a novel Weighted Score Index (WSI), for comprehensive and multi-dimensional comparison of advisory performance among LLMs. Overall results show that Gemini gives the highest quality overall as well as the highest level of empathy among LLMs while ChatGPT has the next highest level of advisory quality. Mistral and LLaMA both had specific strengths in certain scenarios, but both lacked emotional engagement and low levels of interpretability overall. Our contributions are: (i) a replicable evaluation protocol and workflow for evaluating LLM-based psychological advisories with counsellor oversight, (ii) a transparent WSI rubric and audit trail for per-criterion scoring and commentary, and (iii) evidence-based guidance for model selection and governance in digital mental health applications. DASS21-EvaLLM is an evaluation and training tool not a diagnostic system that supports safer deployment, improves counselling practice and supervision, and informs the design of responsible, human-centred advisory systems.


Download File

Full text not available from this repository.

Additional Metadata

Item Type: Article
Subject: Computer Science Applications
Subject: Computer Networks and Communications
Subject: Library and Information Sciences
Divisions: Faculty of Computer Science and Information Technology
Faculty of Medicine and Health Science
Institute for Mathematical Research
DOI Number: https://doi.org/10.1142/S0219649226500462
Publisher: World Scientific
Keywords: AI in counselling; anxiety; DASS-21; depression; generative AI; interactive evaluation; large language models; mental health advisory; stress
Sustainable Development Goals (SDGs): SDG 3: Good Health and Well-being, SDG 9: Industry, Innovation and Infrastructure, SDG 16: Peace, Justice and Strong Institutions
Depositing User: Ms. Siti Radziah Mohamed@mahmod
Date Deposited: 26 Aug 2026 04:51
Last Modified: 26 Aug 2026 04:51
Altmetrics: https://www.altmetric.com/details.php?domain=psasir.upm.edu.my&doi=10.1142/S0219649226500462
URI: http://psasir.upm.edu.my/id/eprint/127991
Statistic Details: View Download Statistic

Actions (login required)

View Item View Item