# Google Research 推出 AI 视频联合导演，用 4 个智能体框架生成连贯的长视频

> 原标题：Google Research Introduces an AI Video Co-Director: 4 Agentic Frameworks for Coherent, Minutes-Long Video Generation

- 来源：MarkTechPost（RSS）
- 发布时间：2026-09-28T02:44:27.000Z
- AIHOT：https://aihot.news/items/cmuknsyrq1yf2ro9h5u9s4a1u
- 原文：https://www.marktechpost.com/2026/09/27/google-research-introduces-an-ai-video-co-director-4-agentic-frameworks-for-coherent-minutes-long-video-generation

## 摘要

Google Research 推出 AI 视频联合导演，由 4 个智能体框架组成，运行在 Gemini 和 Veo 之上，用于解决多镜头 AI 视频中的身份漂移和级联错误问题。

## 正文 · 原文

[Google Research](https://research.google/blog/coherent-long-form-video-generation/) has introduced an **AI video co-director** for long-form video generation. The suite of 4 agentic frameworks turns short clips into coherent, minutes-long stories. It targets identity drift and cascading errors, the 2 failures that break most multi-shot AI video pipelines today.

## **Why Long AI Videos Fall Apart**

Diffusion models render high-fidelity clips in seconds. Stitching those clips into a story is harder. Most agentic pipelines chain modules with independent, handcrafted prompts. That causes [semantic drift](https://arxiv.org/abs/2511.17986), where attire or scenery shifts between shots. It also causes [cascading failures](https://arxiv.org/abs/2606.24976), where one bad upstream asset corrupts every later shot.

Google team frames this as a credit assignment problem. A broken final video is hard to trace back to the prompt that caused it.

## **How the AI Video Co-Director Works**

The system sits on top of [Gemini](https://deepmind.google/models/gemini/pro/) and Veo. It is model-agnostic, so the same layer can drive other generators. Outputs inherit [SynthID](https://deepmind.google/models/synthid/) watermarking from the base models.

### **1\. Co-Director: creative planning as a bandit search**

[Co-Director](https://arxiv.org/abs/2604.24842), accepted at COLM 2026, uses a multi-armed bandit (MAB). An Orchestrator Agent picks a configuration across Creative Strategy, Narrative Mode, and Aesthetic Archetype. A Pre-Production Agent builds the storyboard. Keyframe, Video, and Audio sub-agents produce the media. An MLLM Judge then scores the cut and sends a factored reward back to the bandit.

### **2\. CANVAS: persistent visual memory**

[CANVAS](https://arxiv.org/abs/2604.13452), accepted at EMNLP 2026, tracks characters, locations, and object states as the story evolves. It retrieves stored visual anchors when a scene returns. In Google’s museum heist test, [AutoStudio](https://arxiv.org/abs/2406.01388) lost the thief’s cap and Gemini-3.1-Pro changed the gemstone. CANVAS kept both consistent.

### **3\. A²RD: segment-by-segment long video**

[A²RD](https://arxiv.org/abs/2605.06924) (Agentic Autoregressive Diffusion) is a training-free architecture. Each segment runs a Retrieve, Synthesize, Refine, Update loop against a multimodal video memory. The agent switches between extrapolation for new story beats and interpolation for returning entities. Google shared a [10-minute film](https://www.youtube.com/watch?v=jaLZK8Zb6Kc) generated this way.

### **4\. VQQA: closed-loop prompt refinement**

[VQQA](https://arxiv.org/abs/2603.12310) (Video Quality Question Answering) generates visual questions for each prompt. VLM critiques act as “semantic gradients” that rewrite the text prompt. It needs no access to model internals. A Global Selection step picks the best video across all iterations, not simply the last one.

## **Benchmarks and Results**

Google built 3 new benchmarks. [GenAD-Bench](https://co-director-agent.github.io/genad_bench.html) has 400 ad scenarios across 200 fictional products from 50 brands. HardContinuityBench stresses scene reappearances and prop state changes. [LVBench-C](https://github.com/dxlong2000/AARD) has 120 scenarios where key assets vanish for at least 10 segments before returning.

-   **Co-Director:** 81.4 average on GenAD-Bench and 3.96 of 5 in human ratings, per the [project page](https://co-director-agent.github.io/). Baselines included Veo 3.1, Kling 3.0 Omni, Wan 2.6, and MovieAgent.
-   **CANVAS:** gains of 21.6% in background continuity, 9.6% in character consistency, and 7.6% in props consistency.
-   **A²RD:** up to 30% better consistency and 20% better narrative coherence on 1 to 10 minute videos.
-   **VQQA:** absolute gains of 11.57% on [T2V-CompBench](https://github.com/KaiyueSun98/T2V-CompBench) and 8.43% on [VBench2](https://github.com/Vchitect/VBench) over vanilla generation.

## **How It Compares**

Feature

[Google AI Video Co-Director](https://research.google/blog/coherent-long-form-video-generation/)

[StoryMem](https://kevin-thu.github.io/StoryMem/) (ByteDance, NTU)

[MovieAgent](https://arxiv.org/abs/2503.07314) (Show Lab, NUS)

[AutoStudio](https://arxiv.org/abs/2406.01388)

**Output**

Minutes-long multi-shot video with voiceover and score

Minute-long multi-shot video

Multi-scene, multi-shot video with subtitles and audio

Multi-turn image sequences (no video)

**Architecture**

4 frameworks in a hierarchical multi-agent orchestration layer

Memory-to-Video diffusion model, shot by shot

Multi-agent chain-of-thought planning (director, screenwriter, storyboard artist, location manager)

3 LLM agents plus a Stable Diffusion based agent

**Consistency mechanism**

Persistent visual memory (CANVAS) and multimodal video memory (A²RD)

Keyframe memory bank from earlier shots

Hierarchical planning plus per-character customization

Subject manager plus Parallel-UNet

**Self-correction loop**

Bandit search with MLLM Judge; VQQA prompt refinement with Global Selection

Semantic keyframe selection and aesthetic filtering

Not reported

Not reported

**Model training**

No fine-tuning; orchestrates existing models

LoRA fine-tuning on the base model

Per-character LoRA (ED-LoRA via ROICtrl)

Training-free

**Base generators**

Gemini and Veo (model-agnostic)

Wan2.2

ROICtrl, SVD, HunyuanVideo I2V

Stable Diffusion

**Longest reported output**

10 minutes (A²RD)

About 1 minute

Not specified

N/A (images)

**Code**

[Co-Director](https://github.com/GoogleCloudPlatform/genmedia-izumi-agent/tree/main/demos/backend/ads_codirector) and [A²RD](https://github.com/dxlong2000/AARD) public; CANVAS coming soon

[Public](https://github.com/Kevin-thu/StoryMem)

[Public](https://github.com/showlab/MovieAgent)

[Public](https://github.com/donahowe/AutoStudio)

_Sources: linked papers, project pages, and GitHub repositories. Verified September 27, 2026._

## **Key Takeaways**

-   Google treats long-form video as a global optimization and world-state tracking problem.
-   4 frameworks cover planning, storyboarding, long generation, and self-correction.
-   It runs as an orchestration layer on Gemini and Veo, with SynthID watermarking.
-   A²RD produced a continuous 10-minute film with stable characters and locations.
-   Co-Director scored 81.4 on GenAD-Bench, ahead of a 75.7 random search baseline.

## **FAQ**

-   **What is Google’s AI video co-director?** It is a multi-agent orchestration layer on Gemini and Veo. It plans, generates, and corrects multi-shot videos to keep them consistent.
-   **How long can the videos be?** A²RD was evaluated on videos from 1 to 10 minutes, and Google released a continuous 10-minute demo film.
-   **Can developers use it today?** Partially. Co-Director and A²RD code is on GitHub. CANVAS code is pending, and the full pipeline is not a Google product.

* * *
