JevBench 发布:面向类型化决策的新基准

Rohan Paul · @rohanpaul_ai · X·2026-09-20 06:19·1小时前
AI 导读

新基准 JevBench 发布,专为输出受限软件决策而非开放式文本的模型设计,紧随 TypeSafe 9 月 15 日发布 Jev——输入应用状态与固定选项,返回带概率的类型化答案。该基准综合智能、校准、速度与成本,用几何平均防止单一维度优势掩盖短板。GPT-5.6 Luna 在难题准确率上明显高于 Jev 1.13.0,但 Jev 因延迟、校准和成本更优而在综合分上领先。

Rohan Paul@rohanpaul_ai
35AI 编辑部评分,满分 100

JevBench 发布:面向类型化决策的新基准

2026-09-20 06:19· 1小时前
AI 导读

新基准 JevBench 发布,专为输出受限软件决策而非开放式文本的模型设计,紧随 TypeSafe 9 月 15 日发布 Jev——输入应用状态与固定选项,返回带概率的类型化答案。该基准综合智能、校准、速度与成本,用几何平均防止单一维度优势掩盖短板。GPT-5.6 Luna 在难题准确率上明显高于 Jev 1.13.0,但 Jev 因延迟、校准和成本更优而在综合分上领先。

A new benchmark called JevBench just dropped.

for models whose output is a bounded software decision rather than open-ended prose and follows TypeSafe’s 15 Sept release of Jev, which takes application state plus fixed choices and returns a typed answer with probabilities instead of prose.

This benchmark's score deliberately combines Intelligence, Calibration, Speed and Cost because deployment can fail even when raw accuracy is high.

e.g. GPT-5.6 Luna records substantially higher hard-case accuracy than Jev 1.13.0, yet Jev leads the composite because the benchmark also prices latency, calibration and cost.

The geometric mean prevents exceptional performance on 1 axis from fully compensating for a weak one.

The result is evidence about a narrow typed-decision workload, not evidence that Jev is generally more capable than GPT-5.6 Luna.

来源:Rohan Paul· x.com