Auditing graph-guided fusion for multimodal relation extraction under 29% cross-modal noise: a paired bootstrap
PDF

Keywords

multimodal relation extraction
graph-guided fusion
cross-modal noise
paired simulation
reproducibility

Abstract

We evaluated graph-guided fusion for multimodal relation extraction under 29% cross-modal noise. A deterministic paired simulation generated 72 cases and preserved a median slice. Mean relation F1 changed from 0.581 to 0.632; the paired difference was +0.051 (95% interval +0.048 to +0.053). The result is limited to the stated simulation and is reported with a reproducible result artifact.

PDF

References

Yang, R., & Gupta, R. (2025). Enhancing Multi-Modal Relation Extraction with Reinforcement Learning Guided Graph Diffusion Framework. In Proceedings of the 31st International Conference on Computational Linguistics (pp. 978-988). Association for Computational Linguistics.

Zhang, Z., Wu, Q., Xia, B., Sun, F., Hu, Z., Sun, Y., & Zhang, S. (2025). Automated Molecular Concept Generation and Labeling with Large Language Models. In Proceedings of the 31st International Conference on Computational Linguistics (pp. 6918-6936).

Hu, J., Dowell, N., Brooks, C., & Yan, W. (2018). Temporal Changes in Affiliation and Emotion in MOOC Discussion Forum Discourse. Lecture Notes in Computer Science, 145-149. https://doi.org/10.1007/978-3-319-93846-2_26

Wang, T., Shi, D., Aguilar, J., & Zurada, J. (2025). AutoClusRE: An automatic clustering-based method for relation extraction and knowledge graph construction. https://doi.org/10.21203/rs.3.rs-7705090/v1

Abdullah, A., & Kim, S. T. (2026). Multimodal Radiology Knowledge Graph Generation Using Vision Language Models (Preprint). https://doi.org/10.2196/preprints.91301

Aidynkyzy, A. (2026). FROM NAMED ENTITY RECOGNITION TO RELATION EXTRACTION: LARGE LANGUAGE MODEL ASSISTED CONSTRUCTION OF THE KAZAKH RELATION EXTRACTION DATASET. Вестник Академии гражданской авиации, 41(2). https://doi.org/10.53364/24138614_2026_41_2_13

Pi, W. (2024). Multimodal Large Language Models for Misinformation Detection and Reasoning. https://doi.org/10.59350/mtep9-gwy69

Yue, P., Tang, H., Li, W., Zhang, W., & Yan, B. (2025). MLKGC: Large Language Models for Knowledge Graph Completion Under Multimodal Augmentation. Mathematics, 13(9), 1463. https://doi.org/10.3390/math13091463

Park, J., Bae, M., Na, J., & Kim, H. J. (2026). Improving Large Molecular Language Model via Relation-aware Multimodal Collaboration. Proceedings of the AAAI Conference on Artificial Intelligence, 40(2), 899-907. https://doi.org/10.1609/aaai.v40i2.37058

Yuan, L., Cai, Y., Wang, J., & Li, Q. (2023). Joint Multimodal Entity-Relation Extraction Based on Edge-Enhanced Graph Alignment Network and Word-Pair Relation Tagging. Proceedings of the AAAI Conference on Artificial Intelligence, 37(9), 11051-11059. https://doi.org/10.1609/aaai.v37i9.26309