Don't sleep on using Jev-as-a-Judge for agent evaluation.
This is one of the most impressive Jev use cases I have found so far.
Jev is a natural fit as a Judge, but it doesn't mean you use it everywhere.
Similarly, you shouldn't use frontier models for evals everywhere.
I'm running lots of tests on this atm, but early results point to an optimized flow (balancing accuracy and cost) that combines Jev and frontier models.
Concretely, use Jev in high-confidence situations, and escalate to a frontier model (GPT-6 or Opus 5.5) in low-confidence verdicts.
Entire write-up coming soon. Let me know if you have questions as I build the full guide.