跳到正文
Arena.ai· @arena · X·· 2 小时前AI 评分58
AI 导读

Arena 宣布以 31 亿美元估值完成 2 亿美元 B 轮融资,并发布衡量 AI 智能体安全与对齐的 Alignment Index。

正文

Today, we're announcing our Series B: $200M at a $3.1B valuation, and the release of Arena's Alignment Index.

AI is advancing faster than our ability to evaluate it. The world needs a neutral third party to measure how safe and aligned AI actually is once it's in the hands of real people. That's the role Arena is stepping into today.

The Alignment Index is built from real-world agent traces, starting with three signals: Unauthorized Action, False Attribution, and Deceptive Completion, with results for over 20 frontier models launching today.

Hear more from our CEO @ml_angelopoulos below, and read the full breakdown, and more details about our company growth in thread.

This funding will allow us to go further on safety signals, agentic capabilities, and modalities. It will also allow us to grow the company to match the scale of our mission. Today, we’re still a nimble team of 90, and we're looking for people to join us across research, product, engineering, and more. Come build with us.

As AI gets more capable and autonomous, understanding how well it can safely and effectively work for people only matters more. That's what Arena is built to measure and advance.

Thank you to our community, our investors, and everyone building with us!

Take a look at our open roles if you want to join us (link in bio).

引用Arena.ai@arena
Introducing the Arena Alignment Index, our new benchmark measuring safety and alignment of AI agents in real-world use. Built from 90K+ real-world agent sessions across 27 models, the index measures three critical signals: - Unauthorized Action (UA): Taking actions beyond the user's instructions or permissions - False Attribution (FA): Attributing statements or actions that are contradicted by user-provided evidence - Deceptive Completion (DC): Claiming a task was completed when it was not. Key findings: - OpenAI models currently lead the Alignment Index - Rogue actions are rare, but can have serious consequences when they occur - Agents can mislead users about task progress - Misalignment risks increase with conversation length - Safety and alignment are improving across model generations As shown in the leaderboard below (sorted by lab), @OpenAI’s GPT-6.1-Sol leads the Arena Alignment Index with a score of 87.9, followed by @AnthropicAI’s Claude-Opus-5.5 at 83.2 and @SpaceXAI's Grok-4.7 at 82.7. OpenAI also has the best observed rates across all three signals: 0.89% Unauthorized Action, 1.98% False Attribution, and 2.34% Deceptive Completion. Across all four labs, newer models consistently outperform their predecessors, suggesting broad progress in agent safety and alignment. As agents take on longer, more complex, and higher-stakes tasks, measuring not just what they can accomplish, but how safely and reliably they act, becomes increasingly important. This marks an important step toward making safety and alignment a core part of how Arena evaluates AI. The index is an initial starting point, and we'll continue expanding the index with additional safety signals and models over time. More analysis below👇
在 X 查看被引用的帖子

来源:Arena.ai · x.com