MHB: Medical Hallucination Benchmark for Large Language Models in Complex Clinical Tasks

Authors

  • Jianrong Lu Zhejiang University, China Ant Group, China
  • Junwei Liu Peking University, China Ant Group, China
  • Xingyun Zheng Zhejiang University, China
  • Minghui Yang Ant Group, China
  • Jian Wang Ant Group, China
  • Ping Wang Peking University, China
  • Yechao Zhang Nanyang Technological University, Singapore

DOI:

https://doi.org/10.1609/aaai.v40i45.41243

Abstract

The integration of Large Language Models (LLMs) into clinical applications presents transformative potential but is undermined by the critical risk of hallucination, the generation of plausible but factually incorrect information. Such failures pose a direct threat to patient safety and the integrity of clinical decision-making. To address this challenge, we introduce MHB, a novel and comprehensive benchmark framework designed to evaluate LLM reliability in two complex, high-stakes clinical contexts: multi-turn medical dialogues and clinical case report analysis. The core of our contribution is a systematic methodology for generating adversarial test cases by injecting ``hallucination traps" into realistic medical data, guided by a fine-grained taxonomy of clinical errors. MHB, comprising 4,695 samples and 20,288 evaluation rubrics, underwent a rigorous, two-stage validation by a panel of 60 licensed physicians from top-tier hospitals, ensuring high clinical realism and consistency. This comprehensive assessment of leading LLMs revealed significant, clinically relevant shortcomings across the board. Even the best-performing model, Claude-4-Sonnet, exhibited a hallucination rate of 29.1%, with some open-source models exceeding 57.0%. All models struggled with specific traps, like fabricated medical data or non-existent guidelines, highlighting prevalent systemic weaknesses.

Downloads

Published

2026-03-14

How to Cite

Lu, J., Liu, J., Zheng, X., Yang, M., Wang, J., Wang, P., & Zhang, Y. (2026). MHB: Medical Hallucination Benchmark for Large Language Models in Complex Clinical Tasks. Proceedings of the AAAI Conference on Artificial Intelligence, 40(45), 38971-38978. https://doi.org/10.1609/aaai.v40i45.41243

Issue

Section

AAAI Special Track on AI for Social Impact I