精选归档 · 第 5 页

81100 条 · 共 142

6月4日6月4日周四

星期四 · 3 条
09:28
xAI:News(网页)精选
AI 评分 75/100
xAI 发布 Grok Imagine 1.5 预览版(图像转视频模型)

xAI 通过 API 发布了图像转视频模型 grok-imagine-video-1.5-preview(Grok Imagine 1.5 预览版)。该模型能将单张静态图片转为流畅的电影感视频,用户提供起始帧和描述运动的提示词后,模型可生成包含相机移动、氛围和物理效果的动画,并保持对源图像的忠实。支持生成 720p 片段,可使用自然语言指令控制镜头、节奏和音效,并支持逐帧拼接成长场景。模型目前通过 xAI API 提供预览使用。


推荐理由:xAI的新视频模型从单张图像生成电影级短片,支持自然语言控制运镜和氛围,对视频创作者和开发者是个值得一试的工具。
09:06
Elon Musk@elonmusk精选
AI 评分 72/100
Grok Imagine视频生成上线VercelGrok Imagine on VercelVercel 的 AI Gateway 上现已推出 Grok Imagine Video 1.5。该服务支持图生视频并同步音频,一次性完成。示例代码: await generateVideo({ model: 'xai/grok-imagine-video-1.5-preview', prompt: 'a rabbit sprinting through nyc' });

Vercel Developers: Grok Imagine Video 1.5 on AI Gateway. Image-to-video generation with synced audio in one pass. 𝚊𝚠𝚊𝚒𝚝 𝚐𝚎𝚗𝚎𝚛𝚊𝚝...

另有 1 家信源报道X:Elon Musk (@elonmusk, xAI)
推荐理由:Grok Imagine Video 1.5 把同步音频塞进了图生视频,一条 prompt 直接出带声短片,做短视频和创意的可以换上这条流水线了。
09:06

6月3日6月3日周三

星期三 · 2 条
13:38
公众号:火山引擎精选
AI 评分 64/100
Vibe Creating:让创作回归「表达」本身

火山引擎 Seedance 2.0 提出 AI 视频创作新范式 Vibe Creating,核心是让创作者放下技术负担,用故事表达代替复杂 Prompt 参数。该范式强调用富有画面感的语言描述场景、情绪和叙事,模型自行理解意图并完成景别、光影、节奏的诠释,避免过度规定镜头调度。适用于文学作品可视化、影视预演等场景,并配套发布《Vibe Creating 实践手册》及可执行的 Prompt Skill,从创意到高质量提示词一步到位。


推荐理由:火山引擎把 Seedance 2.0 的用法提炼成「Vibe Creating」方法论,核心是教人用故事感代替镜头术语,虽然不涉及模型升级,但附带可直接套用的手册和 Skill,做 AI 短视频的可以当成 Prompt 指南。

6月2日6月2日周二

星期二 · 1 条

6月1日6月1日周一

星期一 · 1 条
18:24
Runway:News(网页)精选
AI 评分 61/100
Runway 在伦敦设立欧洲总部及世界模型研究中心

Runway 宣布在伦敦建立新的欧洲总部和专注于通用世界模型的研究中心。公司计划在未来18个月向英国AI生态投资$100M,到2028年投资额将翻倍以上。过去12个月,其在欧洲的订阅销量增长了50%,企业客户占比超20%。新总部将扩大其在欧洲的研究与商业布局,公司正招聘欧洲负责人以组建跨研究、产品、工程和销售的团队,并深化与BBC、Fremantle、WPP等企业的合作。世界模型是其研究的核心,旨在将生成式AI的应用扩展至机器人、科学研究与工业模拟等领域。

另有 1 家信源报道X:Runway (@runwayml)
推荐理由:Runway 把世界模型研发带到伦敦并承诺 1 亿美元投资,不是新品但战略意义清晰,欧洲的视频创作者和工业仿真团队离顶尖工具更近了,做影视、游戏和机器人的可以关注后续落地。

5月31日5月31日周日

星期日 · 1 条
08:00
HuggingFace Daily Papers(社区热门论文)精选
AI 评分 70/100
τ_0-WM:用于机器人操控的统一视频-动作世界模型

τ_0-World Model (τ_0-WM) 是一个统一的视频-动作世界模型,旨在机器人执行动作前预测并评估其未来后果。模型基于共享的视频扩散主干网络构建,提供两个接口:一个联合预测未来视觉潜在表示与连续动作块的视频动作模型,以及一个能将动作序列展开为多视角未来并预测任务进度分数的动作条件视频模拟器。τ_0-WM 使用约27,300小时的多元数据训练,包括真实机器人遥操作、UMI风格交互、自我中心人类视频等。推理时,模型通过测试时计算采样动作候选,并利用去噪一致性和基于模拟器的修正来筛选低质量动作,在长时程和精细机器人操控任务上表现出优于相关基准的性能。


推荐理由:机器人操作领域的大一统尝试,把视频预测和动作生成放在一个扩散模型里,还用27万小时数据训练,做具身智能的可以看看这个架构。

5月30日5月30日周六

星期六 · 1 条
01:38
Google Blog:AI(RSS)精选
AI 评分 74/100
Gemini Omni 与 Gemini 3.5 的 11 个实战展示

Google 在 2026 年 Google I/O 大会上发布了新一代多模态模型 Gemini Omni 与 Gemini 3.5,并同步提供了 11 个视频,集中演示了这两款模型在实际场景中的能力。


推荐理由:Google 官方放出的这组视频演示,直接展示了 Gemini Omni 和 3.5 的实际表现,比参数和 benchmark 更直观,做多模态应用的可以逐帧研究。

5月28日5月28日周四

星期四 · 2 条
05:52
Google Gemini@GeminiApp精选
AI 评分 77/100
Gemini Omni轻松转换视频视觉风格Easily transform your videos into new visual styles with Gemini Omni.Just upload a video or photo and ask Gemini to apply a look or style to your final output.使用 Gemini Omni 轻松将您的视频转换为新的视觉风格。 只需上传视频或照片,并要求 Gemini 为您的最终输出应用某种外观或风格。
另有 13 家信源报道IT之家(RSS)The Decoder:AI News(RSS)Google Blog:AI(RSS)X:阿易 AI Notes (@AYi_AInotes)X:Google DeepMind (@GoogleDeepMind)X:Gemini (@GeminiApp)X:Google AI (@GoogleAI)X:Google AI for Developers (@googleaidevs)X:Jeff Dean (@JeffDean)X:Kim (@kimmonismus)X:Logan Kilpatrick (@OfficialLoganK)X:Rohan Paul (@rohanpaul_ai)X:Sundar Pichai (@sundarpichai)
推荐理由:Gemini 终于把图像风格迁移做到视频上了,并且直接集成到 Omni 里,不需要任何剪辑软件,对短视频创作者是个小但实用的更新。

5月27日5月27日周三

星期三 · 1 条
05:28
Google AI@GoogleAI精选
AI 评分 75/100
Gemini Omni 视频提示词使用指南http://x.com/i/article/2059377716965888000Mastering Gemini Omni: The Ultimate Video Prompting GuideLast week, we introduced Gemini Omni—our newest model designed to create anything from any input, starting with video.You can experience the speed and creativity of Gemini Omni Flash today across @geminiapp, @GoogleFlow, @GoogleFlowMusic, and on @YouTube Shorts and Create.To help you push the boundaries of what’s possible, here are five tips to get the most out of Gemini Omni’s advanced video generation capabilities.1. Leverage Real-World KnowledgeYou don’t need to over-explain the world to Gemini Omni. It’s built with Gemini’s deep understanding of history, science, and culture, so it can reliably create outputs that look, feel, and move realistically. Skip the granular descriptions. Use cultural touchstones, historical eras, or scientific terms directly in your prompt.Example Prompts:• [The video shows items of the alphabet. An unusual item starting with each letter is shown sitting on a table (like a Capybara for C, disco globe for D and Lava Lamp for L). All 26 letters must be represented by 26 items with matching lower thirds displaying the letter. Only one item and lower third at a time. Each lower third must look like a black marker written on a slip of paper in the bottom left. Rapid fire, roughly 9 frames per item at 24FPS. Last frame is a slip of paper "THE END." The whole video is accompanied by calm smooth music]• [Astronaut's POV on Mars]• [A marble rolling fast on a chain reaction style track, continuous smooth shot]2. Take Control of Text RenderingGemini Omni not only has advanced text rendering capabilities, it even allows you seamlessly integrate text into your visuals. You can specify typography, spatial placement, animation styles, and complex visual effects like double exposures all perfectly synced to the action in your video.Example Prompts:• [word by word, one word on the screen at a time: did, you, know, that, this, model, can, do, pretty, good, text!? Each word appears with a different animated style, perfect pacing to a rhythm, sizzle reel]• [Overlay motion-tracked, minimalist text commentary onto the physical environment of the video. This text represents [the subject] deadpan, immediate inner monologue that’s observant, slightly absurd, and life-contemplating. Think “intrusive thoughts.” Clean, white, lowercase sans-serif text (like Helvetica or Inter). The text hovers in 3D space, connected to the subjects being commented on via ultra-thin, crisp, white leader lines]3. Direct Your Camera Like a ProThink like a cinematographer. Gemini Omni responds incredibly well to precise videography directions, camera types, and framing instructions. Try integrating these terms into your next prompt:Example prompts:• Shots & Angles: "One continuous shot", "oner", "static", "locked off", or "fixed angle."• Camera Movements: "Push in", "punch in", "pan left", or "dolly zoom."• Camera Styles: "Natural smartphone zoom", "vintage film camera", or "grainy webcam style."4. Edit Iteratively (and keep what works)Every great video is made in the edit. With Gemini Omni, you don't need to rewrite your entire prompt from scratch to fix a single mistake. Ask for specific, targeted updates, like changing a background or swapping a caption. Omni will preserve the core structure of your video across multiple amends, letting you focus only on what needs tweaking.Example prompts:• [Transport the violin to a new environment]• [Make the violin invisible]• [Change the camera angle so it’s looking over the violinist’s shoulder]5. Change the Action on the FlyWant to alter a character's pacing or emotion mid-scene? You can directly prompt Gemini Omni to modify how a subject moves or interacts with their environment without breaking the continuity of the character model.Example prompts:• [Make the character walk on their tiptoes]• [Speed up the pacing]• [Have them leap into the air]Start CreatingThe director’s chair is yours. Try out these prompting techniques with Gemini Omni Flash, and tag @GoogleAI to show us what you create!Google 发布了其多模态模型 Gemini Omni 的视频生成功能使用指南。该模型可通过 Gemini 应用、Google Flow 等平台体验。指南包含五项提示词技巧:利用模型已有的现实世界知识进行简洁描述;精确控制文本在视频中的渲染与排版;使用专业镜头指令(如推拉摇移)像电影摄影师一样调度画面;通过迭代编辑高效修改视频;以及在生成中直接调整角色的动作节奏或情绪。其核心在于通过精准的提示词引导模型生成复杂且可控的视频内容。另有 13 家信源报道IT之家(RSS)The Decoder:AI News(RSS)Google Blog:AI(RSS)X:阿易 AI Notes (@AYi_AInotes)X:Google DeepMind (@GoogleDeepMind)X:Gemini (@GeminiApp)X:Google AI (@GoogleAI)X:Google AI for Developers (@googleaidevs)X:Jeff Dean (@JeffDean)X:Kim (@kimmonismus)X:Logan Kilpatrick (@OfficialLoganK)X:Rohan Paul (@rohanpaul_ai)X:Sundar Pichai (@sundarpichai)
推荐理由:Google 官方放出的视频提示技巧,没有废话全是可复制的 prompt,想玩 Gemini Omni 的创作者可以直接抄作业。

5月26日5月26日周二

星期二 · 3 条
22:34
Runway:News(网页)精选
AI 评分 68/100
Project Luxo:跨越AI媒体的恐怖谷

Runway通过Project Luxo研究发现,AI生成视频已跨越“恐怖谷”。他们向创意生态从业者展示了《The Rogue》等AI短片及广告样片,评估显示观众开始关注故事本身,而非技术瑕疵。所有作品均由单人团队制作,耗时从3周到4小时不等。Runway认为,这标志着AI媒体成熟——当技术足够好以至于“隐形”,观众沉浸于故事时,便实现了这一跨越。


推荐理由:Runway 用短片和一次百万播放广告测试宣称 AI 视频已越过恐怖谷,观众开始投入故事而非找瑕疵。这对内容生产的心理门槛是一次重塑,但一次推广式的成功不等于行业已稳定跨过。
11:18
HuggingFace Daily Papers(社区热门论文)精选
AI 评分 70/100
WBench:面向交互式世界模型评估的多轮基准

WBench 是一个用于系统评估交互式世界模型的多轮基准。它提出了一个五维评估框架,涵盖视频质量、场景设定遵循度、交互指令遵循度、一致性与物理符合性。该基准包含 289 个测试案例与 1,058 轮交互,覆盖了多样化的场景、风格、主体及第一/第三人称视角。评估使用 22 个结合专业视觉模型与大型多模态模型的自动子指标,所有指标均经过人工校验。对 20 个 SOTA 模型的评测发现,目前尚无模型在所有维度上表现均优。


推荐理由:视频世界模型的评估终于有了统一尺度,WBench 从画面质量到物理一致性覆盖五个维度,289 个测试用例把 20 个模型拉平一看,没有谁全面领先,做这方向的值得拿来跑一遍。
08:00
HuggingFace Daily Papers(社区热门论文)精选
AI 评分 72/100
GE-Sim 2.0:面向机器人操作的全面闭环视频世界模拟器路线图

GE-Sim 2.0是一个用于机器人操作的闭环视频世界模拟器。它基于动作条件视频生成框架,并使用数千小时涵盖遥操作与接触交互等真实世界数据进行重新训练,提升了动作跟随与轨迹覆盖能力。其核心新增三个模块:从视频潜变量解码本体感受状态的“状态专家”;为生成轨迹评分并提供成功信号与奖励的“世界评判”;以及能实现快速轨迹生成的加速框架。该模型仅2B参数,在WorldArena排行榜上位列第一,优于专用模型与闭源生成器,其训练出的策略能转化为实际世界性能提升。


推荐理由:过去机器人策略训练卡在仿真到真机的鸿沟上,GE-Sim 2.0 把视频生成、状态提取和自动评估闭环了,策略迭代效率可能翻倍,搞具身智能的很值得蹲一下。

5月23日5月23日周六

星期六 · 2 条
06:39
ViggleAI@ViggleAI精选
AI 评分 75/100
动作捕捉与角色动画制作更轻松motion capture and character animation have never been easier.keep building & more features coming soon!动作捕捉和角色动画制作从未如此简单。 持续构建,更多功能即将推出!

PINOC: A walkthrough of what PINOC does: 🧵 1. Upload a motion video, get clean skeletal animation. Export as .fbx/.glb, ready ...


推荐理由:Viggle 把视频转骨骼动画这件事做到了零成本,无动捕设备、直接导出 FBX,对独立动画师和小团队挺友好,值得试试看。
01:50
Ethan Mollick@emollick精选
AI 评分 76/100
Gemini Omni原生视频编辑能力解析I think people don't realize why Gemini Omni is different than other video AIs. It is fully multimodal, so it can edit video natively, tooI took the famous "train " movie from 1896 & made it a bullet train, LEGO, added a time traveler, a centipede, muppets... (see reflections?)我认为人们没有意识到Gemini Omni与其他视频AI的不同之处。它是完全多模态的,因此也能原生编辑视频。 我拿了1896年著名的"火车"电影,把它变成了高铁、乐高,加入了时间旅行者、蜈蚣、布偶……(看到倒影了吗?)
另有 13 家信源报道IT之家(RSS)The Decoder:AI News(RSS)Google Blog:AI(RSS)X:阿易 AI Notes (@AYi_AInotes)X:Google DeepMind (@GoogleDeepMind)X:Gemini (@GeminiApp)X:Google AI (@GoogleAI)X:Google AI for Developers (@googleaidevs)X:Jeff Dean (@JeffDean)X:Kim (@kimmonismus)X:Logan Kilpatrick (@OfficialLoganK)X:Rohan Paul (@rohanpaul_ai)X:Sundar Pichai (@sundarpichai)
推荐理由:Ethan Mollick 用几个例子把 Gemini Omni 的真正能力讲清楚了,原生多模态让视频编辑不再是生硬叠加,而是理解场景后的重构,做视频的该看。

5月22日5月22日周五

星期五 · 2 条
14:03
公众号:龙猫LongCat(美团)精选
AI 评分 88/100
LongCat-Video-Avatar 1.5 正式开源

美团正式开源 LongCat-Video-Avatar 1.5,在唇形同步、物理合理性、长视频稳定性、多人互动和高效推理上实现全面跃升。该模型采用 DMD 蒸馏将生成步骤压缩至 8 步,推理效率提升约 15 倍,生成 10 秒视频仅需约 1 分钟;Wav2Vec2 升级为 Whisper-large,并引入 GRPO 逐帧偏好对齐优化手部稳定性。


推荐理由:开源数字人模型首次在多人场景区分、长视频稳定性和推理效率上同时达到商业产品水平,评测透明度高。
02:45
Runway:News(网页)精选
AI 评分 74/100
Runway发布Aleph 2.0视频编辑模型及Edit Studio应用

Runway于2026年5月21日发布了视频编辑模型Aleph 2.0及其新产品Edit Studio。Aleph 2.0支持编辑最长30秒的1080p视频,具备精准局部编辑能力,可只改变指定内容而完全保留原视频其余部分。该模型引入了基于单帧图像的精确控制,并支持一次性跨多个镜头应用编辑。Edit Studio是基于这些新能力构建的应用,旨在帮助用户高效地将现有视频素材转化为所需版本,例如更换产品、调整背景或修复拍摄瑕疵。该功能现已向所有付费Runway桌面网页端用户开放,使用优惠码可享受套餐折扣。


推荐理由:精准局部编辑是过去一年 AI 视频工具最大的短板,Aleph 2.0 把这事做对了,预览控制加多镜头编辑让商业视频迭代成本大幅下降。

5月21日5月21日周四

星期四 · 1 条
02:14
Google Gemini@GeminiApp精选
AI 评分 72/100
Gemini Omni让视频创作编辑更轻松Creating, remixing, and editing a video is easier than ever with Gemini Omni.It offers a fluid, conversational way to create and edit. Just upload a video from your camera roll and ask Gemini to make changes.使用Gemini Omni创建、混剪和编辑视频比以往任何时候都更容易。 它提供了一种流畅的对话式创作和编辑方式。只需从相册上传视频,并让Gemini进行修改即可。
另有 13 家信源报道IT之家(RSS)The Decoder:AI News(RSS)Google Blog:AI(RSS)X:阿易 AI Notes (@AYi_AInotes)X:Google DeepMind (@GoogleDeepMind)X:Gemini (@GeminiApp)X:Google AI (@GoogleAI)X:Google AI for Developers (@googleaidevs)X:Jeff Dean (@JeffDean)X:Kim (@kimmonismus)X:Logan Kilpatrick (@OfficialLoganK)X:Rohan Paul (@rohanpaul_ai)X:Sundar Pichai (@sundarpichai)
推荐理由:Gemini Omni把视频编辑做成了对话,虽然不算革命性更新,但对随手剪片的普通人来说,不用学剪辑软件就是最大的可用性。