RAG-R1:Incentivizing the Search and Reasoning Capabilities of LLMs Through Multi-Query Parallelism

Authors

  • Zhiwen Tan Ant Group
  • Jiaming Huang Zhejiang University Ant Group
  • Qintong Wu Ant Group
  • Hongxuan Zhang Nanjing university Ant Group
  • Chenyi Zhuang Ant Group
  • Jinjie Gu Ant Group

DOI:

https://doi.org/10.1609/aaai.v40i39.40603

Abstract

Large Language Models (LLMs), despite their remarkable capabilities, are prone to generating hallucinated or outdated content due to their static internal knowledge. While Retrieval-Augmented Generation (RAG) integrated with Reinforcement Learning (RL) offers a solution, these methods are fundamentally constrained by a single-query mode, leading to prohibitive latency and inherent brittleness. To overcome these limitations, we introduce RAG-R1, a novel two-stage training framework centered around multi-query parallelism. Our framework enables LLMs to adaptively leverage internal and external knowledge during the reasoning process while transitioning from the single-query mode to multi-query parallelism. This architectural shift bolsters reasoning robustness while significantly reducing inference latency. Extensive experiments on seven question-answering benchmarks confirm the superiority of our method, which outperforms the strongest baseline by up to 13.7% and decreases inference time by 11.1%.

Downloads

Published

2026-03-14

How to Cite

Tan, Z., Huang, J., Wu, Q., Zhang, H., Zhuang, C., & Gu, J. (2026). RAG-R1:Incentivizing the Search and Reasoning Capabilities of LLMs Through Multi-Query Parallelism. Proceedings of the AAAI Conference on Artificial Intelligence, 40(39), 33187–33195. https://doi.org/10.1609/aaai.v40i39.40603

Issue

Section

AAAI Technical Track on Natural Language Processing IV