Research Papers research paper arxiv causal architecture multi-shot video generation next-shot generation

ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling

HuggingFace Papersby Yawen Luo ,March 26, 20262 min read3 views

🧒Explain Like I'm 5Simple language

Hey there, little explorer! Imagine you have a magic drawing box.

This magic box, called ShotStream, can make up cool cartoon videos for you, super fast!

You know how when you watch a cartoon, it has different scenes, like a cat chasing a mouse, then the mouse hiding? ShotStream can make these scenes, one after another, like a story.

The best part? You can tell it what to draw while it's drawing! "Now make the cat wear a hat!" And poof, it does!

It's like having a super-fast friend who draws your story as you tell it, and makes sure all the pictures look like they belong together. No waiting, just fun stories happening right now!

ShotStream enables real-time interactive multi-shot video generation through causal architecture design, dual-cache memory mechanisms, and two-stage distillation to maintain visual consistency and reduce latency. (45 upvotes on HuggingFace)

Abstract

AI-generated summary

Multi-shot video generation is crucial for long narrative storytelling, yet current bidirectional architectures suffer from limited interactivity and high latency. We propose ShotStream, a novel causal multi-shot architecture that enables interactive storytelling and efficient on-the-fly frame generation. By reformulating the task as next-shot generation conditioned on historical context, ShotStream allows users to dynamically instruct ongoing narratives via streaming prompts. We achieve this by first fine-tuning a text-to-video model into a bidirectional next-shot generator, which is then distilled into a causal student via Distribution Matching Distillation. To overcome the challenges of inter-shot consistency and error accumulation inherent in autoregressive generation, we introduce two key innovations. First, a dual-cache memory mechanism preserves visual coherence: a global context cache retains conditional frames for inter-shot consistency, while a local context cache holds generated frames within the current shot for intra-shot consistency. And a RoPE discontinuity indicator is employed to explicitly distinguish the two caches to eliminate ambiguity. Second, to mitigate error accumulation, we propose a two-stage distillation strategy. This begins with intra-shot self-forcing conditioned on ground-truth historical shots and progressively extends to inter-shot self-forcing using self-generated histories, effectively bridging the train-test gap. Extensive experiments demonstrate that ShotStream generates coherent multi-shot videos with sub-second latency, achieving 16 FPS on a single GPU. It matches or exceeds the quality of slower bidirectional models, paving the way for real-time interactive storytelling. Training and inference code, as well as the models, are available on our

View arXiv page View PDF Project page GitHub 93 Add to collection

Get this paper in your agent:

hf papers read 2603.25746

Don't have the latest CLI?

curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2603.25746 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2603.25746 in a Space README.md to link it from this page.

Collections including this paper 2

Original source

HuggingFace Papers

https://huggingface.co/papers/2603.25746

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

Market NewsFresh

David Hogg's PAC leaves some Dem campaigns fuming

Leaders We Deserve, the PAC founded by David Hogg to elect young progressives in Democratic primaries, is leaving some of the campaigns it endorsed griping about alleged broken promises. Why it matters: Multiple campaigns backed by Hogg's PAC fumed after primary losses that the group dangled hopes of financial commitments that never materialized. First it was Irene Shin: The Washington Post reported last July that Leaders We Deserve backed off a commitment to spend $400,000 on the 38-year-old Virginia state delegate's behalf in a U.S. House special election that was won by now-Rep. James Walkinshaw (D-Va.). Now sources close to the campaign of Robert Peters , a 40-year-old Illinois state senator who finished a distant third in the primary to succeed Rep. Robin Kelly (D-Ill.), are telling a

Axios Tech

4mabout 2 hours ago

CountriesLive

China Is Willing to Coordinate on AI Governance

View the official memo here. China has consistently signaled a willingness to engage on global AI governance since at least 2017. This memo compiles key statements from the Chinese government and prominent figures demonstrating their desire to coordinate on the problem of AI. Chinese Vice Premier Ding Xuexiang, at the 2025 World Economic Forum, said: [ ] The post China Is Willing to Coordinate on AI Governance appeared first on Machine Intelligence Research Institute .

intelligence.org

1mabout 2 hours ago

Research Papers

Exclusive | OpenAI’s Former Research Chief Aims to Automate Manufacturing With AI - WSJ

Exclusive | OpenAI’s Former Research Chief Aims to Automate Manufacturing With AI WSJ

GNews AI manufacturing

1mabout 1 month ago

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 205 connections

Scroll to zoom · drag to pan · click to open

Discussion

No comments yet — be the first to share your thoughts!

More in Research Papers

Research Papers

Exclusive | OpenAI’s Former Research Chief Aims to Automate Manufacturing With AI - WSJ

Exclusive | OpenAI’s Former Research Chief Aims to Automate Manufacturing With AI WSJ

GNews AI manufacturing

1mabout 1 month ago

Research Papers

AI Journey 2025 Conference: exploring the future of artificial intelligence - Азия-Плюс

AI Journey 2025 Conference: exploring the future of artificial intelligence Азия-Плюс

Google News - AI Tajikistan

1m5 months ago

Research Papers

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors

Vision Language Models struggle with fine-grained visual perception tasks due to their language-centric training approach, performing poorly on unnamed visual entities despite having relevant information in their representations. (1 upvotes on HuggingFace)

HuggingFace Papers

3m5 days ago

Research PapersFresh

AI vs Machine Learning: What Actually Separates Them In 2026? - Independent Newspaper Nigeria

AI vs Machine Learning: What Actually Separates Them In 2026? Independent Newspaper Nigeria

Google News: Machine Learning

1mabout 7 hours ago