Qadir, Sarvech; Ni, Congning; Vaidya, Mihir Sachin; Ryu, Hyeyoung; Mulvaney, Shelagh A.; Kantarcioglu, Murat; Novak, Laurie Lovett; Malin, Bradley; Rose, Susannah Leigh; Yin, Zhijun. (2026).Ìý.ÌýProceedings of the 9th ACM Conference on Fairness, Accountability, and Transparency (FAccT 2026), 6702–6722.Ìý
´¡²õÌýlarge language models (LLMs) are increasingly used to provide mental health support, ensuring that they respond safely and consistently has become an important concern. Most evaluations focus on the quality of a model’s final answer, but this study examined whether the models actually followed the plans they appeared to make while generating those responses. The researchers introduced the concept of execution fidelity, which measures how well a model’s final response aligns with the commitments it expressed during its internal reasoning process, such as showing empathy, offering guidance, setting appropriate boundaries, or framing a problem. Five leading LLMs were evaluated using 75 high-risk mental health scenarios involving depression, anxiety, trauma, and emotional distress. The analysis found that models often differed in how they planned their responses and, in some cases, failed to carry through on important safety-related commitments. Common issues included omitting disclaimers, providing less detailed guidance than originally planned, or replacing specific recommendations with more general emotional reassurance. The researchers also found that greater uncertainty and inconsistency in a model’s planning were associated with a higher likelihood of these gaps. These findings suggest that examining how AI systems develop their responses—not just the final output—may help identify safety risks and improve oversight of LLMs used in mental health settings.
