SQM-Trust: An LLM-Based Assistant for Auditing Trustworthiness in ML-Powered Speech and Language Anxiety Assessment
Abstract
Speech and language are increasingly used as non-invasive markers for diagnosing and monitoring anxiety, and recent advancements in machine learning techniques, natural language processing, large language models, and speech processing have accelerated the development of ML-based anxiety assessment tools. However, trustworthiness has received far less attention than accuracy, even though clinical usability depends on the former. This paper introduces the SQM-Trust agent (at the prototype stage), built on a trustworthy framework of 19 items—adapted from STARD-AI (Standards for Reporting Diagnostic Accuracy Studies–Artificial Intelligence), QUADAS-AI (Quality Assessment of Diagnostic Accuracy Studies–Artificial Intelligence), and MINIMAR (Minimum Information for Medical AI Reporting)—spanning study design, fairness, transparency, ML development, and clinical usability, tailored specifically to speech- and language-based anxiety assessment. The SQM-Trust agent is an LLM-based agent (powered by Claude Sonnet 4.6) that interactively audits a developer's report against each trustworthiness item, assigns a 0–2 score, and suggests concrete revisions. We present the current prototype and a preliminary, pre-fine-tuning evaluation against expert-assigned scores on 14 (excerpt, item) pairs drawn from 10 held-out papers—a partial coverage of the 19-item framework rather than full item-by-item scoring of every paper. The SQM-Trust agent could match expert scores exactly for 6 of the 10 papers (60.0\% exact-match accuracy) and obtained a mean absolute error (MAE) of 0.283 on the 0–2 rubric and a quadratic-weighted Cohen's kappa of 0.722. These early results suggest reasonable agreement with expert judgment, with no clear tendency to over- or under-score. The main aim behind developing the SQM-Trust agent is to provide a lightweight, accessible tool to help junior ML developers build and report more trustworthy anxiety-assessment systems before deployment.