Abstract
Background: Large Language Models (LLMs) have demonstrated near-saturation performance on static medical examinations. However, a significant "knowledge-practice gap" persists, wherein high test accuracy fails to translate into safe, reliable, and legally compliant clinical reasoning in real-world environments. Unguided open-loop generations are heavily restricted by uncalibrated overconfidence, logic degradation under multi-turn pressure, and persistent factual hallucinations.
Objective: We propose and evaluate ClinicaGuard, a novel multi-agent orchestration framework designed to bridge the clinical knowledge-practice gap by enforcing deterministic safety boundaries, programmatic clinical knowledge graph verification, and uncertainty-aware response calibration.
Methods: The framework integrates three distinct layers: (1) an intent classification and Unified Medical Language System (UMLS) Named Entity Recognition layer to anchor inputs to verifiable clinical terminologies; (2) a multi-agent consensus module utilizing a Mixture-of-Experts (MoE) debate architecture between distinct diagnostic and compliance auditing agents; and (3) a semantic entropy filtration system to quantify and suppress uncalibrated model confidence. We benchmarked the framework using a multi-tiered matrix consisting of static datasets (MedQA), conversational clinical practice environments (MedAgentBench), and an expert-driven clinical alignment audit.
Results: Our framework significantly outperformed unguided zero-shot and standard Chain-of-Thought (CoT) baselines. While maintaining a baseline diagnostic accuracy of over 91% on static retrieval, the multi-agent consensus architecture reduced factual hallucination rates by 78% in multi-turn clinical simulations. Furthermore, the integration of semantic entropy calibration yielded a notable improvement in the alignment of model confidence with empirical accuracy, effectively suppressing high-confidence diagnostic failures.
Conclusion: Transitioning medical LLMs from isolated text generators to structured, calibrated multi-agent systems dramatically reduces execution liability in clinical workflows. Programmatic knowledge-graph anchoring and multi-agent cross-verification offer a viable pathway toward safe, compliant, and deployable clinical AI systems within heavily regulated healthcare infrastructures.
References
Dai, Y. (2026). Rescaling confidence: What scale design reveals about LLM metacognition. arXiv preprint arXiv:2603.09309.
Liu, M., Yang, W., & Liu, J. (2026). Prompting Rain Off: Evolving Compact Dual Prompts for Continual De-Raining. IEEE Transactions on Image Processing.
Liu, M., Xie, J., Hu, Y., Yang, W., & Liu, J. (2023, July). Comprehensive Augmented Domain Adaptation for Image Segmentation Under Rainy Conditions. In 2023 IEEE International Conference on Multimedia and Expo Workshops (ICMEW) (pp. 63-68). IEEE.
Liu, M., Yang, W., Hu, Y., & Liu, J. (2023, August). Dual prompt learning for continual rain removal from single images. In Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence (Vol. 3).
Ding, T., Xiang, D., Sun, T., Qi, Y., & Zhao, Z. (2025, June). AI-driven prognostics for state of health prediction in Li-ion batteries: A comprehensive analysis with validation. In 2025 6th International Conference on Electrical Technology and Automatic Control (ICETAC) (pp. 22-28). IEEE.
Ye, X., Zhang, J., Cheng, Z., Lu, F., Wen, Z., Yu, G., ... & Ren, M. (2025). Hemodynamic modeling of aortic arch aneurysm treatment using the Castor™ branched stent graft: a virtual coil embolization simulation framework. Frontiers in Physiology, 16, 1629346.
Yang, J. X., Zhou, J., Wang, J., Tian, H., & Liew, A. W. C. (2024). LiDAR-guided cross-attention fusion for hyperspectral band selection and image classification. IEEE Transactions on Geoscience and Remote Sensing, 62, 1-15.
