Infinite-Story: A Training-Free Consistent Text-to-Image Generation

Authors

  • Jihun Park Daegu Gyeongbuk Institute of Science and Technology
  • Kyoungmin Lee Daegu Gyeongbuk Institute of Science and Technology
  • Jongmin Gim Daegu Gyeongbuk Institute of Science and Technology
  • Hyeonseo Jo Daegu Gyeongbuk Institute of Science and Technology
  • Minseok Oh Daegu Gyeongbuk Institute of Science and Technology
  • Wonhyeok Choi Daegu Gyeongbuk Institute of Science and Technology
  • Kyumin Hwang Daegu Gyeongbuk Institute of Science and Technology
  • Jaeyeul Kim Daegu Gyeongbuk Institute of Science and Technology
  • Minwoo Choi Daegu Gyeongbuk Institute of Science and Technology
  • Sunghoon Im Daegu Gyeongbuk Institute of Science and Technology

DOI:

https://doi.org/10.1609/aaai.v40i10.37776

Abstract

We present Infinite-Story, a training-free framework for consistent text-to-image (T2I) generation tailored for multi-prompt storytelling scenarios. Built upon a scale-wise autoregressive model, our method addresses two key challenges in consistent T2I generation: identity inconsistency and style inconsistency. To overcome these issues, we introduce three complementary techniques: Identity Prompt Replacement, which mitigates context bias in text encoders to align identity attributes across prompts; and a unified attention guidance mechanism comprising Adaptive Style Injection and Synchronized Guidance Adaptation, which jointly enforce global style and identity appearance consistency while preserving prompt fidelity. Unlike prior diffusion-based approaches that require fine-tuning or suffer from slow inference, Infinite-Story operates entirely at test time, delivering high identity and style consistency across diverse prompts. Extensive experiments demonstrate that our method achieves state-of-the-art generation performance, while offering over 6x faster inference (1.72 seconds per image) than the existing fastest consistent T2I models, highlighting its effectiveness and practicality for real-world visual storytelling.

Published

2026-03-14

How to Cite

Park, J., Lee, K., Gim, J., Jo, H., Oh, M., Choi, W., … Im, S. (2026). Infinite-Story: A Training-Free Consistent Text-to-Image Generation. Proceedings of the AAAI Conference on Artificial Intelligence, 40(10), 8278–8286. https://doi.org/10.1609/aaai.v40i10.37776

Issue

Section

AAAI Technical Track on Computer Vision VII