TypeSafe AI· @typesafeai · X·· 2 小时前AI 评分38
AI 导读
Jev 对研究智能体来说是一个巨大的增强! @edwardirby 构建了一个 @youdotcom 智能体来监控商业风险,并将 Jev 与普通的 LLM 评判器进行了对比。 • 报告质量相同 • 但 LLM 反复无常:对同一威胁的评分从 0.35 -> 0.68 -> 0.50 • LLM 漏掉了 11 项调查中的 5 项。Jev 一项都没漏。 • Jev 便宜 250 倍,快 3-6 倍 详细文章 👇
正文
Jev is a huge power-up for research agents!
@edwardirby built a @youdotcom agent that monitors business risks, and compared Jev to an ordinary LLM judge.
- Same report quality
- But the LLM flip-flopped: 0.35 -> 0.68 -> 0.50 on the same threat
- the LLM missed 5 of 11 investigations. Jev missed none.
- Jev was 250x cheaper, 3-6x faster
Write-up 👇
How to build a reliable risk agent without a frontier model ($0.02/sweep): • @youdotcom search • @typesafeai Jev for typed judgments • @QwenDevs for proposals & synthesis • MCP for integration Full architecture & cost breakdown below 👇 https://x.com/i/article/2105364784459997184在 X 查看被引用的帖子
来源:TypeSafe AI · x.com