Measuring confidence gating for multimodal decision support under 46% visual-text disagreement: a threshold audit
PDF

Keywords

multimodal decision support
confidence gating
visual-text disagreement
paired simulation
reproducibility

Abstract

We evaluated confidence gating for multimodal decision support under 46% visual-text disagreement. A deterministic paired simulation generated 64 cases and preserved a median slice. Mean decision consistency changed from 0.522 to 0.558; the paired difference was +0.036 (95% interval +0.034 to +0.038). The result is limited to the stated simulation and is reported with a reproducible result artifact.

PDF

References

Dai, Y., Peng, X., Wang, Y., Nakov, P., & Xie, Z. (2026). Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs? arXiv preprint arXiv:2608.05864. arXiv. https://doi.org/10.48550/arXiv.2608.05864

Wang, L., Ouyang, L., Weng, H., Chen, X., Wang, A., & Zhang, K. (2026). Intelligent Disassembly System for PCB Components Integrating Multimodal Large Language Model and Multi-Agent Framework. Processes, 14(2), 227. https://doi.org/10.3390/pr14020227

Jin, Y. (2023). Decision making in large-scale language testing: Intersections of policy, practice and research. Studies in Language Assessment, 64-92. https://doi.org/10.58379/hkkx5020

HIRAGAR, P. R. (2025). Med-AgentX: Multimodal Large Language Model Agents with Explainable Reinforcement Learning for Trustworthy Biomedical Decision Support. https://doi.org/10.21203/rs.3.rs-7773686/v1

Shi, K. (2024). Optimization of Decision Making Capabilities of Intelligent Characters in Strategy Games Based on Large Language Model Agent. Journal of Big Data and Computing, 2(3), 136-140. https://doi.org/10.62517/jbdc.202401320

Mustafayeva, A., Israfilova, E., Baxshiyeva, G., & Aslanova, S. (2026). Hybrid Cnn-gru Model for Real-time Multimodal Decision-making in Image and Text Analysis. https://doi.org/10.21203/rs.3.rs-9257523/v1

Yang, S., Zhang, Y., & Gu, C. (2026). Large Language Model-Enhanced Agent-Based Modeling for Intelligent Crowd Evacuation under Disaster Scenarios. https://doi.org/10.5194/egusphere-egu26-3725

Mohammadi, P., Asgari, S., Rashidi, A., & Adibpour Hassankiadeh, S. (2026). A RAG-Enhanced Multimodal Large Language Model Framework for Culvert Inspection in Utah. https://doi.org/10.2139/ssrn.6429799

Schumacher, I., Bühler, V. M. M., Jaggi, D., & Roth, J. (2024). Artificial intelligence derived large language model in decision-making process in uveitis. International Journal of Retina and Vitreous, 10(1), 63. https://doi.org/10.1186/s40942-024-00581-1

Zhou, M., & Yang, J. (2026). From Street-View Imagery to Auditable Spatial Evidence: A Multimodal Large Language Model Workflow for Urban-Renewal Screening. https://doi.org/10.2139/ssrn.7225622