Artificial Analysis 发布 AA-Video-T2V v2.0 和无音频版 AA-Video-T2V-Silent v2.0 文生视频基准,统一以 1080p 高码率评估,按 10 种用例、10 种能力和多种视觉风格排名,分别基于超 68,000 票(1,000+ prompts)和超 47,000 票(500 prompts)的人工偏好投票。
We are launching AA-Video-T2V v2.0, our new benchmark for evaluating text to video models, alongside AA-Video-T2V-Silent v2.0 for video generation without audio. Built on a new methodology, it judges every model at 1080p on a regularly refreshed prompt set, and ranks them across 10 use cases, 10 capabilities and a wide range of styles.
Video models are being adopted across more industries and workflows, from film studios to advertising agencies. Our new benchmark not only ranks models overall, but also shows which model is best for specific use case and video generation capability. Use cases are grounded in how consumers and enterprises use video generation. Capabilities draw on lab and academic research, and on how creators and businesses push video models today. We tag every prompt by use case (such as Live-Action Film and Marketing & Advertising) and by the capability it tests (such as Text Rendering and Audio Synchronization), and the overall benchmark samples evenly across both. Because of this, the overall ranking reflects a model's versatility across use cases and well-roundedness across capabilities. We also tag each prompt by visual style, such as photorealistic, 3D render, cartoon and anime, and hand-drawn illustration.
AI video is also moving onto bigger screens and into production, from microdramas to movie theaters, while low barrier to generate is resulting in a proliferation of low quality AI video content. The quality bar keeps rising, so we now judge every clip at 1080p and high bitrate.
We are launching AA-Video-T2V v2.0 with more than 68,000 high quality human preference votes from private evaluators based in US/UK over 1,000 prompts, and AA-Video-T2V-Silent v2.0 with more than 47,000 votes over 500 prompts.
Initial insights from an in-depth analysis of the 10 highest ranking models on the Artificial Analysis AA-Video-T2V v2.0 Leaderboard:
➤ Wan 3.0 ranks #1 overall and leads 10 of the 20 category boards, including Cartoon and Anime style and Animation & Gaming use case, at $12 per minute of video.
➤ Dreamina Seedance 2.5 ranks #2 and is the human performance specialist, #1 on both Human Anatomy and Dialogue & Lip Sync. At $34.12/min it is the most expensive model in the top 10.
➤ MiniMax H3 (768p) ranks #3, statistically tied with Seedance 2.5 at $4.80/min, about 1/7 of the price. It is also #1 on Text Rendering.
➤ FLUX 3 ranks #4, with its strongest results on Text Rendering (#3) and Dialogue & Lip Sync (#2).
➤ Gemini Omni Flash 1.1 ranks #5 and is the graphic 2D and audio specialist, #1 on UI/UX & Motion Design use case, Flat Design style and Audio Synchronization capability.
See below for the use case, capability and style breakdowns 🧵
来源:Artificial Analysis · x.com