arXiv · 2609.37317
What Comes Next? Omni-StoryBench for Evaluating Story-Grounded Omnimodal Generation
Abstract
Omnimodal evaluation should go beyond independent text, image, and speech production: individually plausible outputs may not express a coherent shared event. We introduce Omni-StoryBench, a story-grounded omnimodal benchmark evaluating whether models can coherently continue stories across image, narration, and speech. Each instance provides a current storybook page and structured next-page conditions, requiring models to generate the next illustration, narration, and spoken character utterance. Omni-StoryBench contains 900 rigorously validated story transitions from openly licensed children's books, with ground-truth next-page references and speech metadata. We evaluate systems with modality-specific metrics and consistency-centered LLM-as-a-judge rubrics for context preservation, condition following, reference consistency, and cross-modal coherence. Across 32 baseline configurations spanning orchestration, semi-orchestration, and native any-to-any paradigms, we find orchestration with strong VLM planning most reliable, while current native omnimodal models often struggle with output completeness and controllability. Our analysis shows text-side performance is associated with image and speech quality, but image generation and visual continuity form the clearest observed bottleneck among the evaluated configurations. These results position Omni-StoryBench as a system-level benchmark measuring coherent omnimodal generation beyond isolated modality quality.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sieun Hyeon, Yejoon Lee, Mintaek Lim, Woojin Kim, Jaeik Kim, Jaeyoung Do. 2026-09-29. What Comes Next? Omni-StoryBench for Evaluating Story-Grounded Omnimodal Generation. https://arxiv.org/abs/2609.37317
Cite the original work for its findings. Save a collection to share your selection of sources.