跳到正文
TypeSafe AI· @typesafeai · X·· 2 小时前AI 评分38
AI 导读

Jev 对研究智能体来说是一个巨大的增强! @edwardirby 构建了一个 @youdotcom 智能体来监控商业风险,并将 Jev 与普通的 LLM 评判器进行了对比。 • 报告质量相同 • 但 LLM 反复无常:对同一威胁的评分从 0.35 -> 0.68 -> 0.50 • LLM 漏掉了 11 项调查中的 5 项。Jev 一项都没漏。 • Jev 便宜 250 倍,快 3-6 倍 详细文章 👇

正文

Jev is a huge power-up for research agents!

@edwardirby built a @youdotcom agent that monitors business risks, and compared Jev to an ordinary LLM judge.

  • Same report quality
  • But the LLM flip-flopped: 0.35 -> 0.68 -> 0.50 on the same threat
  • the LLM missed 5 of 11 investigations. Jev missed none.
  • Jev was 250x cheaper, 3-6x faster

Write-up 👇

引用Edward Irby@edwardirby
How to build a reliable risk agent without a frontier model ($0.02/sweep): • @youdotcom search • @typesafeai Jev for typed judgments • @QwenDevs for proposals & synthesis • MCP for integration Full architecture & cost breakdown below 👇 https://x.com/i/article/2105364784459997184
在 X 查看被引用的帖子

来源:TypeSafe AI · x.com