Reliability Boundaries for Human-Facing Language Models
PDF

Keywords

synthetic patients
educational NLP
psychometrics
language-model safety

Abstract

This internal reference article examines reliability limits of language models in clinical and educational contexts through a design-and-assurance lens. It synthesizes the allocated target literature without reporting new experiments, observations, or performance estimates. The analysis treats the practical unit of review as a model output interpreted through a measurement instrument, learner profile, or professional workflow. That framing keeps technical mechanisms, evidence quality, user consequences, and institutional controls visible in the same argument. Particular attention is given to when fluent generation is fit for support rather than substitution. The review distinguishes what each cited source directly addresses from the cross-domain principles used for internal comparison. It argues that credible adoption depends on traceable requirements, context-sensitive evaluation, explicit uncertainty, and a documented path for human intervention. The result is a structured reference for teams considering educational content, simulated cases, and professional decision support, especially where construct misalignment and misplaced user trust could turn a technically plausible component into an unreliable system. The article is intended to support scoping, design review, and evidence planning; it is not a claim of product readiness or an original empirical study.

PDF

References

Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2001). The PHQ-9: Validity of a brief depression severity measure. Journal of General Internal Medicine, 16(9), 606–613. https://doi.org/10.1046/j.1525-1497.2001.016009606.x

Liu, K., Xiong, H., Zhang, J., & Peng, M. (2026). MOSAIC: A Cognitively Motivated Multi-Agent Framework for Interpretable and Training-Free Empathetic Dialogue. Electronics, 15(10), 2078.

Liu, K., Xiong, H., Zhang, J., & Peng, M. (2026). Unifying Aesthetic Evaluation via Multimodal Annotation and Fine-Grained Sentiment Analysis. Big Data and Cognitive Computing, 10, 37.

Shen, Q., & Han, Y. (2026). The Reliability Illusion in Synthetic Patients: Psychometric Misalignment of Open-weight LLMs on PHQ-9 and GAD-7. In Proceedings of the 10th Workshop on Computational Linguistics and Clinical Psychology (CLPsych 2026) (pp. 88–99). Association for Computational Linguistics.

Shen, Q., Cao, F., Yao, M., Gilda, S., Dorr, B., & Leite, W. (2026). Children’s English Reading Story Generation via Supervised Fine-Tuning of Compact LLMs with Controllable Difficulty and Safety. In Proceedings of the 21st Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2026) (pp. 751–765). Association for Computational Linguistics.

Spitzer, R. L., Kroenke, K., Williams, J. B. W., & Löwe, B. (2006). A brief measure for assessing generalized anxiety disorder: The GAD-7. Archives of Internal Medicine, 166(10), 1092–1097. https://doi.org/10.1001/archinte.166.10.1092

Tao, J., Lyu, R., & Cao, X. (2026). A Deep Learning-Based Automated Content Moderation Framework for Online Platforms. Future-Adaptive Intelligence and Lifelong Systems, 1(1).

Tao, J., Lyu, R., & Cao, X. (2026). A Scalable Data Governance Architecture for Privacy-Aware Intelligent Learning Systems in Lifelong. Future-Adaptive Intelligence and Lifelong Systems, 1(1).

Wang, J., Fan, L., Li, B., & Zhang, L. (2026). Forecasting with Guidance: Representation-Level Supervision for Time Series Forecasting. arXiv preprint arXiv:2603.24262.

Xiong, H., Zhang, J., Wang, Z., Pan, T., & Hu, Q. (2026). VividTalker: A Modular Framework for Expressive 3D Talking Avatars with Controllable Gaze and Blink. In ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).

Zhang, L., & Yong, F. (2026). The impact of government accounting supervision on insider trading in China. Borsa Istanbul Review, 26(1), 1–14. Article 100764.