腾讯混元发布 WebCraftBench 真实应用评测基准

Arena.ai · @arena · X·2026-09-24 04:30·2小时前
AI 导读

腾讯混元团队推出 WebCraftBench,基于 Code Arena 内部复刻版收集的 369 条真实请求构建,涵盖非正式语言、不完整指令等真实使用场景。该基准测试 AI 构建应用的外观、功能与需求达成度,在 17 个模型上的排名与 Code Arena 排行榜相关系数达 0.89;在 197 对人工验证样本上,与人类偏好匹配率 85.3%。

Arena.ai@arena
45AI 编辑部评分,满分 100

腾讯混元发布 WebCraftBench 真实应用评测基准

2026-09-24 04:30· 2小时前
AI 导读

腾讯混元团队推出 WebCraftBench,基于 Code Arena 内部复刻版收集的 369 条真实请求构建,涵盖非正式语言、不完整指令等真实使用场景。该基准测试 AI 构建应用的外观、功能与需求达成度,在 17 个模型上的排名与 Code Arena 排行榜相关系数达 0.89;在 197 对人工验证样本上,与人类偏好匹配率 85.3%。

Real-world app requests don’t come as perfect specs. Congrats to the @TencentHunyuan team on WebCraftBench!

Built from 369 requests collected from their internal replica of Code Arena, it embraces the messy details of real human usage: informal language, incomplete instructions, different project scopes, and diverse preferences.

Their benchmark tests how AI-built apps look, work, and deliver on those requests. Across 17 models, its rankings strongly correlate with our Code Arena leaderboard (0.89).

What could you uncover from real human-AI conversations and Arena’s leaderboard history? Explore our open datasets at http://huggingface.co/lmarena-ai/datasets!

Tencent HyThe AI says the site is done. Then the homepage errors, the buttons overlap, and it added a login flow you never asked for. 🙃 Introduces WebCraftBench: agents ...

来源:Arena.ai· x.com