Rethinking Evaluation in Reliability and Rater Agreement through Bias Diagnosis and Constructed Responses

Keywords

judge quality
contest scoring
reliability weighting
Monte Carlo simulation
inter-rater agreement

Abstract

The reference set connects Research on evaluation system of the judge quality in students' contest… with studies of GenAI for automated essay scoring: A turing test of rubrics, rater… and Radiographic Evaluation of Lumbar Intervertebral Disc Height Index: An Intra and…, offering several competing ways to frame reliability and rater agreement. The analysis follows bias diagnosis, inter-rater reliability, and constructed responses, tracing points of convergence as well as differences in terminology, measurement, and experimental design. This makes hidden assumptions visible without introducing unsupported performance claims. The contribution is a structured reading of the existing evidence, not a new synthetic benchmark claim. It provides criteria for selecting methods, interpreting metrics, and designing follow-up studies that can be independently checked.

References

Zhao, R., Guo, Y., Ma, X., Yin, X., & Tang, J. (2022). Research on evaluation system of the judge quality in students' contest based on data driven. In 2022 3rd International Conference on Education, Knowledge and Information Management (ICEKIM) (pp. 1075-1079). IEEE. https://doi.org/10.1109/ICEKIM55072.2022.00233

Rtam, N. (2024). Evaluation of the departmental inter-rater reliability when scoring thyroid nodules according to the British Thyroid Association Ultrasound-classification model: Is there significant disagreement?. Ultrasound, 32(2), 76-84. https://doi.org/10.1177/1742271x231215500

Gazi, M., Fadairo, A., Minor, M., Acker, J., & Gropen, T.-I. (2019). Abstract WP298: Communication Center Guided Prehospital Stroke Assessment Scoring has High Inter-rater Agreement. Stroke, 50(Suppl_1). https://doi.org/10.1161/str.50.suppl_1.wp298

Söğüt, S., & Büyükkıdık, S. (2026). GenAI for automated essay scoring: A turing test of rubrics, rater agreement, and authorship in L2 writing assessment. Educational Assessment, Evaluation and Accountability. https://doi.org/10.1007/s11092-026-09495-y

Kim, T. (2026). Inter-rater agreement in Korean constructed-response writing assessment: a large-scale empirical baseline from the AI Hub 2024 writing evaluation dataset. Frontiers in Education, 11. https://doi.org/10.3389/feduc.2026.1887497

Gonzales, F. (2025). Inter-Rater Agreement and Competency Gaps in QSEN: A Comparative Study of Student Nurse Self-Assessment and Nurse Educator Evaluation. NURSE EDUCATORS AND PRACTITIONERS JOURNAL, 1(2), 9. https://doi.org/10.64397/nepj.v01i02.2025.a13

Chen, X., Sima, S., Sandhu, H., Kuan, J., & Diwan, A. (2022). Radiographic Evaluation of Lumbar Intervertebral Disc Height Index: An Intra and Inter-Rater Agreement and Reliability Study. . https://doi.org/10.2139/ssrn.4137633

Izzetti, R., Fulvio, G., Nisi, M., Gennai, S., & Graziani, F. (2022). Reliability of OMERACT Scoring System in Ultra-High Frequency Ultrasonography of Minor Salivary Glands: Inter-Rater Agreement Study. Journal of Imaging, 8(4), 111. https://doi.org/10.3390/jimaging8040111

Yun, J. (2023). Relationships among Different Effect-Size Indexes for Inter-Rater Agreement between Human and Automated Essay Scoring. Korean Association For Learner-Centered Curriculum And Instruction, 23(18), 901-919. https://doi.org/10.22251/jlcci.2023.23.18.901

Liao, S.-C., Hunt, E.-A., & Chen, W. (2010). Comparison between Inter-rater Reliability and Inter-rater Agreement in Performance Assessment. Annals of the Academy of Medicine, Singapore, 39(8), 613-618. https://doi.org/10.47102/annals-acadmedsg.v39n8p613