VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction

Authors

  • Hao Wang Waseda University
  • Eiki Murata CyberAgent, Inc. AI Shift, Inc.
  • Lingfang Zhang Waseda University
  • Ayako Sato CyberAgent, Inc.
  • So Fukuda Waseda University
  • Ziqi Yin Waseda University
  • Wentao Hu Waseda University
  • Keisuke Nakao Waseda University
  • Yusuke Nakamura Waseda University
  • Sebastian Zwirner Waseda University
  • Yi-Chia Chen Waseda University
  • Hiroyuki Otomo CyberAgent, Inc.
  • Hiroki Ouchi Nara Institute of Science and Technology CyberAgent, Inc.
  • Daisuke Kawahara Waseda University

DOI:

https://doi.org/10.1609/aaai.v40i12.37938

Abstract

Recent advances in multimodal large language models (MLLMs) have significantly enhanced video understanding capabilities, opening new possibilities for practical applications. Yet current video benchmarks focus largely on indoor scenes or short-range outdoor activities, leaving the challenges associated with long-distance travel largely unexplored. Mastering extended geospatial-temporal trajectories is critical for next-generation MLLMs, underpinning real-world tasks such as embodied-AI planning and navigation. To bridge this gap, we present VIR-Bench, a novel benchmark consisting of 200 travel videos that frames itinerary reconstruction as a challenging task designed to evaluate and push forward MLLMs' geospatial-temporal intelligence. Experimental results reveal that state-of-the-art MLLMs, including proprietary ones, struggle to achieve high scores, underscoring the difficulty of handling videos that span extended spatial and temporal scales. Moreover, we conduct an in-depth case study in which we develop a prototype travel-planning agent that leverages the insights gained from VIR-Bench. The agent’s markedly improved itinerary recommendations verify that our evaluation protocol not only benchmarks models effectively but also translates into concrete performance gains in user-facing applications.

Downloads

Published

2026-03-14

How to Cite

Wang, H., Murata, E., Zhang, L., Sato, A., Fukuda, S., Yin, Z., … Kawahara, D. (2026). VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction. Proceedings of the AAAI Conference on Artificial Intelligence, 40(12), 9747–9756. https://doi.org/10.1609/aaai.v40i12.37938

Issue

Section

AAAI Technical Track on Computer Vision IX