Reka AI 发布 19B 全能模型 Rho-1,单模型处理文本、图像、视频与机器人控制
Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model
Reka AI 发布 190 亿参数全能模型 Rho-1 的研究预览版,在单一神经网络内处理并生成文本、图像、视频和机器人控制动作。该模型将全部模态作为 token 放入同一上下文窗口,无需工具调用或外部模型,可实时生成连续视频并即时响应新指令,同一套权重既预测摄像头图像也驱动机器人运动。
Reka AI has released a research preview of Rho-1. The 19-billion-parameter omni-model processes and generates text, images, video, and robot control actions in a single neural network. Unlike most AI systems that route tasks to specialized models, Rho-1 runs all modalities as tokens in one shared context window with no tool calls or external models. The model generates continuous video in real time and responds to new instructions on the fly without restarting.
The same weights that predict camera images also drive robot movements. To work around scarce robot training data, Reka AI built an inverse dynamics model that pulls control signals from ordinary internet videos. Rho-1 trained on 320 H100 GPUs over about three months.
Reka AI isn't new to multimodal AI. In April 2024, the company shipped Reka Core, a multimodal language model that competed with GPT-4, Claude 3, and Gemini Ultra on benchmarks. The release fits a broader push in AI research toward so-called world models.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
来源:The Decoder:AI News · the-decoder.com