Stress-testing retrieval self-check for retrieval-augmented answering under 22% distractor density: a paired bootstrap
PDF

Keywords

retrieval-augmented answering
retrieval self-check
distractor density
paired simulation
reproducibility

Abstract

We evaluated retrieval self-check for retrieval-augmented answering under 22% distractor density. A deterministic paired simulation generated 64 cases and preserved a median slice. Mean evidence faithfulness changed from 0.557 to 0.601; the paired difference was +0.044 (95% interval +0.042 to +0.047). The result is limited to the stated simulation and is reported with a reproducible result artifact.

PDF

References

Sang, Y. (2025). AutoCrit: A Meta-Reasoning Framework for Self-Critique and Iterative Error Correction in LLM Chains-of-Thought. 2025 6th International Conference on Machine Learning and Computer Application (ICMLCA), 1177-1180. https://doi.org/10.1109/icmlca66850.2025.11336788

Ji, S., Zhang, X., Huang, Y., & Li, D. (2026). CGS-RAG:Community-Aggregated Graph Semantic Retrieval-Augmented Generation. https://doi.org/10.21203/rs.3.rs-9748268/v1

Dhenia, R. N. K., Sridhar, R., & Kanani, I. J. (2024). Retrieval-Augmented Generation: Enhancing Reliability in Large Language Models. American International Journal of Computer Science and Technology, 6, 46-49. https://doi.org/10.63282/3117-5481/aijcst-v6i2p105

Roy, P., & Singirikonda, A. (2026). Optimization and Constraint Modeling using LLMs with a Retrieval Augmented Generation Process. https://doi.org/10.2139/ssrn.7068218

Joseph, A. (2025). Reducing Hallucinations in Large Language Models Through Integrated Self-Verification and Retrieval-Augmented Generation. Volume 2B: 45th Computers and Information in Engineering Conference (CIE), V02BT02A032. https://doi.org/10.1115/detc2025-169730

Garg, A., & Esposito, J. (2026). Operationalizing Evidence Admissibility in Retrieval-Augmented Generation: The DCA-TRAG Architecture. https://doi.org/10.2139/ssrn.7111618

Kumar, C., Deshmukh, S., Khatter, H., Kumar, B., & Pal, Y. (2026). ACW-RC: Adaptive Confidence-Weighted Retrieval Correction for Robust Augmented Generation. https://doi.org/10.2139/ssrn.6546797

Farizi, A. A., Arsi, P., & Subarkah, P. (2026). Comparative Performance of Retrieval Augmented Generation Tourism Chatbots. Indonesian Journal of Innovation Studies, 27(1). https://doi.org/10.21070/ijins.v27i1.1836

Kasim Vali, D. (2025). Autonomous QA Agent: A Retrieval-Augmented Framework for Reliable Selenium Script Generation. https://doi.org/10.31224/5891

Budha, K., & Lagun, N. (2026). Operationalizing Reliability Gaps in Large Language Models: A Semi-Systematic Evidence Map of Reasoning, Factuality, Evaluation, and Retrieval-Augmented Generation. https://doi.org/10.21203/rs.3.rs-9882260/v1