现有的计算机使用基准测试无法捕捉真实世界计算机使用的现实性、复杂性和长周期需求,限制了它们揭示前沿智能体局限性的能力。我们推出了 OSWorld 2.0,这是一个包含 108 个长周期计算机使用工作流的基准测试,涵盖日常和专业任务,旨在捕捉复杂且具有挑战性的真实世界现象。
每个任务都代表一个真实的端到端工作流,人类用户完成其中位时间约为 1.6 小时,而使用 Claude Opus 4.7 并开启最大思考模式时,平均需要 318 次工具调用,相比之下 OSWorld 1.0 中约为 30 次。
OSWorld 2.0 针对真实工作流中常见但以往基准测试中代表性不足的挑战性现象,涵盖了交互设计挑战(如流式交互和动态环境)以及智能体模式挑战(如跨来源推理、隐式状态推断和视觉空间精度)。任务基于真实的输入工件,并与真实的、有状态的用户画像数据进行交叉引用,同时包含单独的安全报告以审计安全敏感的执行过程。
在我们主要的 500 步二值完成度指标下,Claude Opus 4.8 在开启最大思考模式和批量工具调用时表现最佳,但仅完成了 20.6% 的任务,部分得分为 54.8%;GPT-5.5 的 token 效率高得多,但得分停滞在 13% 左右。
这些结果表明,当前智能体距离专业级别的计算机使用仍相去甚远:它们并非在基本的 GUI 控制或编码上出错,而是会丢失约束条件、错过任务中途到达的信息、猜测而非询问用户,并且跳过验证环节,在任务依赖于它们必须恢复的隐藏状态时最为挣扎。
Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents. We introduce OSWorld 2.0, a benchmark of 108 long-horizon computer-use workflows across everyday and professional tasks, designed to capture complex and challenging real-world phenomena. Each task represents a realistic end-to-end workflow that takes human users a median of about 1.6 hours to complete and requires an average of 318 tool calls with Claude Opus 4.7 using maximum thinking, compared with about 30 in OSWorld 1.0.
OSWorld 2.0 targets challenge phenomena that are common in real workflows yet underrepresented in prior benchmarks, spanning interaction-design challenges such as streaming interaction and dynamic environments, as well as agent-pattern challenges such as cross-source reasoning, implicit-state inference, and visual-spatial precision. Tasks are grounded in authentic input artifacts and cross-referenced against realistic stateful user profile data, and include separate safety reports auditing safety-sensitive execution. Under our primary binary-completion metric at 500 steps, Claude Opus 4.8 with maximum thinking and batched tool calls scores best but still completes only 20.6% of tasks at a 54.8% partial score; GPT-5.5 is far more token-efficient yet plateaus near 13%.
These results show that current agents are still far from professional-level computer use: rather than stumbling on basic GUI control or coding, they lose track of constraints, miss information that arrives mid-task, guess rather than ask the user, and skip verification, struggling most when a task hinges on hidden state they must recover.