Real-world app requests don’t come as perfect specs. Congrats to the @TencentHunyuan team on WebCraftBench!
Built from 369 requests collected from their internal replica of Code Arena, it embraces the messy details of real human usage: informal language, incomplete instructions, different project scopes, and diverse preferences.
Their benchmark tests how AI-built apps look, work, and deliver on those requests. Across 17 models, its rankings strongly correlate with our Code Arena leaderboard (0.89).
What could you uncover from real human-AI conversations and Arena’s leaderboard history? Explore our open datasets at http://huggingface.co/lmarena-ai/datasets!