Exposing the Cracks: Vulnerabilities of Retrieval-Augmented LLM-based Machine Translation

Authors

  • Yanming Sun University of Macau
  • Runzhe Zhan University of Macau
  • Chi Seng Cheang Singapore Management University
  • Han Wu University of Macau
  • Xuebo Liu Harbin Institute of Technology, Shenzhen
  • Yuyao Niu South China University of Technology
  • Fengying Ye University of Macau
  • Kaixin Lan University of Macau
  • Lidia S. Chao University of Macau
  • Derek F. Wong University of Macau

DOI:

https://doi.org/10.1609/aaai.v40i39.40597

Abstract

REtrieval-Augmented LLM-based Machine Translation (REAL-MT) shows promise for knowledge-intensive tasks like idiomatic translation, but its reliability under noisy retrieval, a common challenge in real-world deployment, remains poorly understood. To address this gap, we propose a noise synthesis framework and new metrics to systematically evaluate REAL-MT’s reliability across high-, medium-, and low-resource language pairs. Using both open- and closed-sourced models, including standard LLMs and large reasoning models (LRMs), we find that models heavily rely on retrieved context, and this dependence is significantly more detrimental in low-resource language pairs, producing nonsensical translations. Although LRMs possess enhanced reasoning capabilities, they show no improvement in error correction and are even more susceptible to noise, tending to rationalize incorrect contexts. Attention analysis reveals a shift from the source idiom to noisy content, while confidence increases despite declining accuracy, indicating poor self-monitoring. To mitigate these issues, we investigate training-free and fine-tuning strategies, which improve robustness at the cost of performance in clean contexts, revealing a fundamental trade-off. Our findings highlight the limitations of current approaches, underscoring the need for self-verifying integration mechanisms.

Published

2026-03-14

How to Cite

Sun, Y., Zhan, R., Cheang, C. S., Wu, H., Liu, X., Niu, Y., … Wong, D. F. (2026). Exposing the Cracks: Vulnerabilities of Retrieval-Augmented LLM-based Machine Translation. Proceedings of the AAAI Conference on Artificial Intelligence, 40(39), 33135–33143. https://doi.org/10.1609/aaai.v40i39.40597

Issue

Section

AAAI Technical Track on Natural Language Processing IV