We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain specialists across all of these capabilities. To enable this, we develop a scalable environment and reinforcement learning (RL) infrastructure. On the environment side, we host hundreds of concurrent instances across heterogeneous per-domain backends, serving both trajectory data collection and online RL. To enable efficient and scalable RL training, an asynchronous framework keeps explicit control over the cross-domain training distribution and remains stable under noisy environment feedback and off-policy drift. Experimental results show that MintAct achieves state-of-the-art performance (48.9 on OSWorld-Verified) across a wide range of benchmarks at comparable model sizes.
MintAct:面向数字环境的统一视觉智能体
AI 导读
MintAct 是一系列视觉语言模型,统一了 UI grounding、跨移动端/桌面/Web 的多步导航与视觉工具使用,提供 2B、4B、8B 三种规模,性能可匹配各领域专用模型。其配套的可扩展环境与异步 RL 框架支持数百个并发实例和在线强化学习,在 OSWorld-Verified 上取得 48.9 的 SOTA 成绩。
HuggingFace Daily Papers(社区热门论文)
49
AI 编辑部评分,满分 100MintAct:面向数字环境的统一视觉智能体
MintAct 是一系列视觉语言模型,统一了 UI grounding、跨移动端/桌面/Web 的多步导航与视觉工具使用,提供 2B、4B、8B 三种规模,性能可匹配各领域专用模型。其配套的可扩展环境与异步 RL 框架支持数百个并发实例和在线强化学习,在 OSWorld-Verified 上取得 48.9 的 SOTA 成绩。
原文 · 保持原样,未翻译原文 · 未翻译
来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org