JevBench is Benchmark Heaven's own benchmark for Jev-class decision models: state and a bounded rubric in, a typed answer out.
Compare Jev alternativesChoose a Jev-class model by use case
Scored 23 Sept 2026 · protocol jevbench::v1.4 · 534 public + 308 sealed aggregate decisions · one request at a time from a server in Germany · harness, public tasks & scoring rules (MIT) · results JSON sha256 808909c4ecc8… · v1.0 results
JevBench v1.4.1 ranking
JevBench v1.4.1
JevBench Score: 77 ranked systems
Official· four axes 0–100, equal-weight harmonic mean · What changed in v1.4 ↓
- 1Jev 1.13.0API63.3I 53 · C 76 · S 83 · K 52 · $0.040
- 2JevK5 v0.2.062.0I 49 · C 75 · S 91 · K 60 · ~$0.022 est.
- 3Hopper59.4I 48 · C 79 · S 87 · K 59 · ~$0.024 est.
- 4Winnow-12B Q855.6I 48 · C 65 · S 82 · K 53 · ~$0.037 est.
- 5reflex 4B54.0I 47 · C 70 · S 68 · K 60 · ~$0.022 est.
- 6djev52.2I 47 · C 55 · S 91 · K 58 · $0.026 ann.
- 7Jev-Omni51.3I 47 · C 64 · S 82 · K 53 · ~$0.037 est.
- 8metask-jev-4b47.8I 45 · C 67 · S 89 · K 55 · ~$0.033 est.
- 9SemIf47.7I 44 · C 67 · S 84 · K 59 · ~$0.022 est.
- 10Jobe Qwen3.5-4B46.9I 44 · C 66 · S 86 · K 60 · ~$0.022 est.
- 11local-jev Qwen3.5-4B46.8I 44 · C 73 · S 75 · K 56 · ~$0.030 est.
- 12system-one-openAPI45.1I 44 · C 55 · S 77 · K 65 · ~$0.015 est.
- 13spark-s1-4b-v644.6I 45 · C 48 · S 81 · K 58 · ~$0.025 est.
- 14jqv44.4I 46 · C 72 · S 75 · K 47 · ~$0.056 est.
- 15Qwen3-Reranker-4B43.5I 45 · C 65 · S 79 · K 49 · $0.050
- 16decider-35b-a3b41.2I 47 · C 65 · S 81 · K 45 · ~$0.067 est.
- 17Raw Qwen3 4B Instruct 2507 direct logits41.0I 46 · C 29 · S 88 · K 60 · ~$0.022 est.
- 18OpenSourceJev40.9I 42 · C 60 · S 82 · K 64 · ~$0.016 est.
- 19ZeroEntropy zerank-240.2I 42 · C 76 · S 79 · K 50 · $0.047
- 20decision-machine-1API39.9I 41 · C 68 · S 93 · K 54 · $0.035
- 21Raw Phi-4 mini direct logits38.0I 42 · C 59 · S 89 · K 50 · ~$0.048 est.
- 22JEV Qwen3.5-9B Base NVFP437.7I 47 · C 68 · S 93 · K 43 · ~$0.077 est.
- 23OpenJev36.9I 45 · C 55 · S 83 · K 45 · ~$0.066 est.
- 24kev 4B36.1I 42 · C 40 · S 76 · K 62 · ~$0.019 est.
- 25Decision 2B35.8I 39 · C 74 · S 84 · K 63 · ~$0.018 est.
- 26Qwen3.5-9B Jev-like data-mix v235.2I 47 · C 61 · S 82 · K 42 · ~$0.083 est.
- 27GPT-6 LunaAPI35.1I 96 · C 92 · S 74 · K 37 · $0.127
- 28SimpleJev Qwen3.8-27BAPI34.6I 52 · C 74 · S 71 · K 39 · ~$0.104 est.
- 29NInfer Qwen3.8-Flash-Next mixed34.0I 50 · C 79 · S 88 · K 39 · ~$0.109 est.
- 30GPT-6 LunaAPI33.3I 97 · C 93 · S 73 · K 36 · $0.135
- 31open-alternative-jev33.2I 39 · C 59 · S 83 · K 60 · ~$0.022 est.
- 32jev-local32.5I 45 · C 64 · S 69 · K 43 · ~$0.077 est.
- 33Decision Fast32.5I 37 · C 65 · S 82 · K 76 · ~$0.0063 est.
- 34decider-2b30.7I 39 · C 43 · S 83 · K 61 · ~$0.020 est.
- 35jeff30.6I 37 · C 68 · S 63 · K 77 · ~$0.0060 est.
- 36Laya30.3I 36 · C 64 · S 71 · K 86 · ~$0.0029 est.
- 37lev-350m28.5I 35 · C 71 · S 85 · K 76 · ~$0.0063 est.
- 38openjev-sglangAPI27.7I 49 · C 69 · S 77 · K 36 · ~$0.131 est.
- 39Von27.5I 34 · C 76 · S 70 · K 78 · ~$0.0055 est.
- 40NInfer Qwen3.8-27B NVFP426.9I 51 · C 76 · S 80 · K 35 · ~$0.145 est.
- 41NInfer Qwen3.8-27B NVFP426.3I 51 · C 67 · S 80 · K 35 · ~$0.145 est.
- 42kev 8B25.6I 42 · C 40 · S 75 · K 44 · ~$0.073 est.
- 43JevOne25.5I 48 · C 78 · S 88 · K 36 · ~$0.137 est.
- 44SimpleJev Qwen3.6-35B-A3BAPI24.9I 46 · C 60 · S 75 · K 38 · ~$0.116 est.
- 45kev 0.6B24.8I 34 · C 50 · S 76 · K 76 · ~$0.0063 est.
- 46Raw Qwen3 8B direct logits23.7I 46 · C 24 · S 86 · K 42 · ~$0.087 est.
- 47system-one23.4I 44 · C 33 · S 84 · K 41 · ~$0.089 est.
- 48OpenDecision21.6I 32 · C 57 · S 80 · K 75 · ~$0.0066 est.
- 49LitJev19.5I 46 · C 77 · S 67 · K 34 · ~$0.163 est.
- 50openJev Verdict 1.419.0I 29 · C 72 · S 78 · K 82 · ~$0.0039 est.
- 51kev 0.5B18.9I 31 · C 50 · S 77 · K 76 · ~$0.0063 est.
- 52Bespoke Nimble 9B18.7I 46 · C 56 · S 79 · K 33 · ~$0.166 est.
- 53GPT-5.6 LunaAPI18.5I 93 · C 87 · S 78 · K 28 · $0.242
- 54openJev Verdict18.1I 30 · C 47 · S 77 · K 83 · ~$0.0037 est.
- 55Raw Qwen3 1.7B direct logits18.1I 33 · C 24 · S 90 · K 65 · ~$0.015 est.
- 56reflex-27b17.8I 46 · C 77 · S 67 · K 32 · ~$0.181 est.
- 57djev15.2I 72 · C 88 · S 75 · K 27 · ~$0.274 est.
- 58GLiNER2 large15.1I 31 · C 25 · S 62 · K 73 · ~$0.0077 est.
- 59OpenJev14.8I 58 · C 58 · S 76 · K 28 · ~$0.255 est.
- 60Qwen3.5-0.8B Decision Model14.5I 28 · C 68 · S 49 · K 76 · ~$0.0065 est.
- 61Gemini 3.1 Flash-LiteAPI14.3I 54 · C 59 · S 82 · K 27 · $0.264
- 62open-jev-deberta-v3-large12.6I 26 · C 67 · S 66 · K 74 · ~$0.0073 est.
- 63smalljev semantic-v912.3I 26 · C 59 · S 80 · K 58 · ~$0.025 est.
- 64GLiNER211.8I 27 · C 25 · S 72 · K 83 · ~$0.0037 est.
- 65Open-Jev 9B11.2I 44 · C 62 · S 72 · K 28 · ~$0.249 est.
- 66Open-Jev 2B10.0I 42 · C 55 · S 73 · K 28 · ~$0.249 est.
- 67GLiNER2.5 multi9.8I 23 · C 57 · S 68 · K 82 · ~$0.0039 est.
- 68SimpleJev7.5I 21 · C 49 · S 57 · K 68 · ~$0.011 est.
- 69GLiNER2.5 small7.2I 20 · C 51 · S 78 · K 82 · ~$0.0039 est.
- 70Raw Qwen3 0.6B direct logits7.1I 23 · C 21 · S 90 · K 74 · ~$0.0074 est.
- 71DeepSeek V4.1 FlashAPI4.8I 94 · C 96 · S 72 · K 17 · $0.594
- 72Mirror2.1I 14 · C 26 · S 71 · K 73 · ~$0.0077 est.
- 73Mixedbread mxbai-rerank-base-v20.4I 7 · C 84 · S 88 · K 68 · $0.012
- 74BAAI bge-reranker-v2-m30.2I 5 · C 84 · S 90 · K 73 · $0.0077
- 75Alibaba GTE Reranker ModernBERT-base0.2I 5 · C 79 · S 91 · K 70 · $0.010
- 76Certo v10.0I 0 · C 83 · S 94 · K 100 · ~$0.0010 est.
- 77Open Jev JSON Canvas0.0I 48 · C 0 · S 84 · K 46 · ~$0.065 est.
- classifier.dev (honorable mention)API70.8I 52 · C 72 · S 88 · K 84 · ~$0.0033 est.
- Qwen3.8 27B (partial run)API0.0I 40 · C 94 · S 61 · K 0 · ~$2.669 est.
- swanOne (partial run)—I – · C – · S – · K – · ~$0.111 est.
- Needle 3, options as tools (partial run)—I – · C – · S – · K – · ~$0.014 est.
- Needle 3 (partial run)—I – · C – · S – · K – · ~$0.024 est.
Score = 4 / (1/I + 1/C + 1/S + 1/K) (each 0–100; × (axis / 50)² for Intelligence, Speed or Cost below 50)
- Jev (TypeSafe, closed)
- Jev rebuild
- Instruction model, JSON schema
- Small tool-calling model
- Service built on Jev
- Zero-shot classifier
- Closed decision API
- Reranker (neutral adapter)
- Raw-logit control (base model)
- Native-logit decision engine
- Shown, not ranked
Axes, accuracy, latency and cost
Every system with its four axes, public and sealed accuracy and the gap between them. On a phone the name column stays put while the table scrolls sideways. † = a note on that system — tap it to read.
| # | System | JevBench Score | Intelligence | Calibration | Speed | Cost axis | Public accuracy 534 | Sealed accuracy 308 | Public − sealed gap | Cost / 1,000 | p50 latency | Endpoint |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 63.3 | 53.1 | 76.3 | 83.3 | 52.0 | 86.6% | 36.7% | +49.9 pp | $0.040 | 0.65 s | API | |
| 2 | 62.0 | 48.9 | 74.5 | 91.1 | 59.5 | 85.3% | 33.1% | +52.2 pp | ~$0.022est. | — | unknown | |
| 3 | Hopper†Author discloses heavy public-benchmark-directed development (26 model/prompt configs, 20+ calibration-map variants observed against the public half); scan of 17 released files vs 231 public tasks found 0 state/instruction matches; training corpus not released so overlap not independently verifiable.by HopitAI | 59.4 | 48.0 | 79.1 | 86.8 | 58.7 | 82.3% | 34.1% | +48.2 pp | ~$0.024est. | 0.13 s | RunPod GPU |
| 4 | Winnow-12B Q8†The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.by Eldan Ring | 55.6 | 48.3 | 64.8 | 82.3 | 52.9 | 85.7% | 33.1% | +52.6 pp | ~$0.037est. | 0.23 s | RunPod GPU |
| 5 | reflex 4B†The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.by kshetrajna12 | 54.0 | 47.5 | 70.4 | 68.0 | 59.7 | 79.2% | 28.2% | +51.0 pp | ~$0.022est. | 1.80 s | RunPod GPU |
| 6 | djev†The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated). v1.4: hosted api.djev.dev was paused by its operator ("Serving is paused by the administrator"); sealed tier measured on the public djev runtime (Davipar/djev-dev 3ce907e, same weights, default mode) self-hosted on an H100; Speed/Cost kept from v1.3.by Maisa (David Villalón) · Maisa, diffusion-gemma | 52.2 | 47.0 | 55.4 | 91.4 | 57.6 | 84.0% | 29.9% | +54.1 pp | $0.026announced | 0.24 s | API |
| 7 | Jev-Omni†Akhilaaa3/Jev-Omni revision c050d51354147985d13286cf4acf90f562f2c631, the author's own load_model.py (merged text decision model + 256-way head) and his own predict(), transformers 5.17.0 / torch 2.8.0 from the pod image; built on the CPU and moved to CUDA with every nn.Linear weight cast to bfloat16 first - the same cast his reference loader jev_omni.py applies - because our 46 GB GPU cannot hold his fp32 copy; on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.by akhilaaa3 · akhilaaa3, Gemma-4-12B merged | 51.3 | 46.8 | 64.1 | 81.5 | 53.0 | 88.7% | 32.1% | +56.6 pp | ~$0.037est. | 0.22 s | RunPod GPU |
| 8 | metask-jev-4b†Model card discloses 44.8k+16.1k+390 training rows incl. synthetic families intentionally mirroring JevBench hard families, plus repeated evaluation on all 231 public items (benchmark-directed development disclosed). Scan of 61 files: 0 exact matches.by Wayfind (metask-ai) | 47.8 | 44.7 | 66.9 | 89.1 | 54.5 | 79.7% | 27.6% | +52.1 pp | ~$0.033est. | 0.07 s | RunPod GPU |
| 9 | SemIfby Theodore Lee (TheoLeeCJ) · formerly OpenJev (Qwen3.5-4B, TheoLeeCJ | 47.7 | 44.4 | 66.8 | 83.7 | 59.5 | 81.0% | 26.3% | +54.7 pp | ~$0.022est. | 0.20 s | RunPod GPU |
| 10 | Jobe Qwen3.5-4B†No trained weights/LoRA/calibration fit; release explicitly rejects fitted temperature/order averaging. Scan of 37 files vs 231 public tasks: 0 matches.by MantisShrimpdev · frozen | 46.9 | 44.1 | 66.1 | 85.6 | 59.5 | 81.0% | 25.6% | +55.3 pp | ~$0.022est. | 0.13 s | RunPod GPU |
| 11 | local-jev Qwen3.5-4B†Scan of 72 files: 0 matches. Author discloses choosing JSON layout after observing results on the 231 public items (benchmark-directed choice, disclosed).by Amith Chandrappa (amithgc) | 46.8 | 44.4 | 73.3 | 75.0 | 55.8 | 80.5% | 26.0% | +54.5 pp | ~$0.030est. | 0.71 s | RunPod GPU |
| 12 | 45.1 | 44.2 | 54.9 | 77.0 | 64.8 | 73.2% | 27.6% | +45.6 pp | ~$0.015est. | 0.65 s | author demo | |
| 13 | spark-s1-4b-v6†Abhishek085/spark-s1-4b-v6 revision 93d49ddbfb29212e3296635a75a3e80cf69da027, code github.com/abhishek085/open-spark-jev 30ac6d89b7fa36c644cf86aac68f35c1d276a919, the author's own MenuScorer.decide with his fitted calibration.json temperature, bf16, base Qwen/Qwen3.5-4B, transformers 5.17.0 / torch 2.8.0 from the pod image, flash-linear-attention 0.5.2 installed, causal_conv1d not installable here (no wheel builds against this toolchain), on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.by Abhishek Rai (abhishek085) · Open Spark Jev, abhishek085 | 44.6 | 45.1 | 47.6 | 81.0 | 57.9 | 79.2% | 26.6% | +52.6 pp | ~$0.025est. | 0.31 s | RunPod GPU |
| 14 | jqv†A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.by hjmurmur (Octalab) · Qwen3-32B zero-shot | 44.4 | 46.4 | 71.6 | 74.6 | 47.5 | 80.1% | 28.2% | +51.8 pp | ~$0.056est. | 0.75 s | RunPod GPU |
| 15 | Qwen3-Reranker-4B†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.by Qwen | 43.5 | 44.6 | 65.2 | 78.7 | 49.2 | 68.0% | 29.9% | +38.1 pp | $0.050 | 0.13 s | RunPod GPU |
| 16 | decider-35b-a3b†The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.by Mapika | 41.2 | 47.2 | 65.3 | 80.8 | 45.3 | 83.1% | 31.5% | +51.6 pp | ~$0.067est. | 0.29 s | RunPod GPU |
| 17 | Raw Qwen3 4B Instruct 2507 direct logits†Neutral raw-logit control.by Alibaba Qwen / neutral reproduction | 41.0 | 46.4 | 29.1 | 87.6 | 59.7 | 69.7% | 27.3% | +42.4 pp | ~$0.022est. | 0.08 s | RunPod GPU |
| 18 | OpenSourceJev†DM submission, measure-only (JevBench publishing HOLD in force). Same author as the existing simplejev-qwen3.5-0.8b row (sabeel111/Featherless AI). Round-4 audit of an earlier commit could not be measured (no public GGUF, Windows-only DLL loader, unpinned llama.cpp build); this round the author published the exact unsloth Q4_K_M GGUF (hash/size independently verified) and we built llama.cpp CUDA from current upstream master on Linux ourselves -- its ABI matched the ctypes bindings exactly, so only a loader file-naming fix was needed (documented diff), no code/scoring/calibration change. Our public-231 subset exactly reproduced the author-reported table: easy 48/48, standard 67/72, hard 66/111, schema 231/231. Calibration (noul temperature) fit only on Google BoolQ, not JevBench. Repo docs name 3 public task IDs while describing benchmark-directed algorithm fixes on the public half (disclosed); 0 exact state/instruction text matches in a released-file scan.by sabeel111 · Qwen3.5-4B Q4_K_M, native llama.cpp | 40.9 | 41.8 | 60.3 | 82.0 | 64.0 | 78.4% | 26.3% | +52.1 pp | ~$0.016est. | — | unknown |
| 19 | ZeroEntropy zerank-2†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.by ZeroEntropy | 40.2 | 42.1 | 75.8 | 79.0 | 49.8 | 70.1% | 28.6% | +41.6 pp | $0.047 | 0.13 s | RunPod GPU |
| 20 | decision-machine-1†A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.APIby milliseconds.ai (Baptiste Laget) | 39.9 | 41.3 | 68.3 | 92.9 | 53.7 | 67.5% | 25.6% | +41.9 pp | $0.035 | 0.17 s | API |
| 21 | Raw Phi-4 mini direct logits†Neutral raw-logit control, not JevBench-directed.by Microsoft / neutral reproduction | 38.0 | 41.8 | 58.8 | 88.8 | 49.6 | 65.8% | 29.2% | +36.6 pp | ~$0.048est. | 0.06 s | RunPod GPU |
| 22 | JEV Qwen3.5-9B Base NVFP4†Byte-identical to upstream March-2026 NVFP4 checkpoint, predates JevBench v1.2, no task-specific training added. Had 3 preliminary scope/policy scoring errors corrected during review (pooled-ECE, full-run latency, cost estimate); score above is final corrected value.by WilfLin | 37.7 | 46.8 | 67.7 | 93.3 | 43.3 | 75.3% | 29.5% | +45.8 pp | ~$0.077est. | 0.02 s | RunPod GPU |
| 23 | OpenJevby razorback16 / Codiv · DiffusionGemma 26B-A4B NVFP4, razorback16 | 36.9 | 45.4 | 55.0 | 83.2 | 45.5 | 81.8% | 28.6% | +53.2 pp | ~$0.066est. | 0.24 s | RunPod GPU |
| 24 | kev 4B†Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 306/308 sealed items answered validly (failures count as wrong)by Jared Palmer · research preview | 36.1 | 42.1 | 39.6 | 75.7 | 61.8 | 66.2% | 22.4% | +43.8 pp | ~$0.019est. | 0.55 s | RunPod GPU |
| 25 | Decision 2B†Flymy-ai/decision-2b-preview revision df57b75db927acc9ad91ec8115508c1e487086eb (checkpoint minicpm5_reduced_v16_4k_v59), base openbmb/MiniCPM5-2B revision 12a3808a956f869c767195e9266b59c4d21d92e2, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.by FlyMy.AI (@denti) · FlyMy.AI, v59 | 35.8 | 38.8 | 74.1 | 84.3 | 62.5 | 75.3% | 26.0% | +49.4 pp | ~$0.018est. | 0.19 s | RunPod GPU |
| 26 | Qwen3.5-9B Jev-like data-mix v2†The author disclosed development on the public JevBench set and public-result comparisons; released-data overlap scan found no matches, but 764 gap and 382 replay training rows are unreleased.by jsaurabh | 35.2 | 47.4 | 61.3 | 82.0 | 42.4 | 78.4% | 29.2% | +49.1 pp | ~$0.083est. | — | unknown |
| 27 | 35.1 | 95.8 | 92.0 | 73.7 | 36.9 | 99.1% | 92.9% | +6.3 pp | $0.127 | 1.44 s | API | |
| 28 | SimpleJev Qwen3.8-27B†Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.APIby Featherless AI | 34.6 | 51.6 | 74.5 | 71.2 | 39.5 | 86.6% | 35.7% | +50.9 pp | ~$0.104est. | 1.01 s | author demo |
| 29 | NInfer Qwen3.8-Flash-Next mixed†Engine scan (1,758 files) vs 231 public tasks: 0 exact matches. Submitter discloses no training for Flash-Next but repeated consultation of public items and a public-hard temperature sweep.by Igor L. / NInfer contributors | 34.0 | 49.5 | 78.6 | 88.2 | 38.9 | 89.6% | 34.1% | +55.5 pp | ~$0.109est. | 0.08 s | RunPod GPU |
| 30 | 33.3 | 97.4 | 93.5 | 72.6 | 36.0 | 99.6% | 95.5% | +4.1 pp | $0.135 | 1.48 s | API | |
| 31 | open-alternative-jev†With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.by IkerMoel · Qwen3.5-4B, IkerMoel | 33.2 | 38.6 | 58.7 | 83.5 | 59.6 | 74.0% | 24.4% | +49.7 pp | ~$0.022est. | 0.21 s | RunPod GPU |
| 32 | jev-local†The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.by us (GitHub) · Qwen3.5-9B | 32.5 | 45.2 | 64.2 | 69.2 | 43.3 | 74.9% | 29.5% | +45.3 pp | ~$0.077est. | 1.05 s | RunPod GPU |
| 33 | Decision Fast†Flymy-ai/decision-fast-preview revision 4225d41c66119fe28e95a2631bb0103decae6d56 (checkpoint qwen3_06b_headfirst_ep2a_v53), base Qwen/Qwen3-0.6B-Base revision da87bfb608c14b7cf20ba1ce41287e8de496c0cd, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.by FlyMy.AI (@denti) · FlyMy.AI, v53a | 32.5 | 37.1 | 65.3 | 81.6 | 76.1 | 63.2% | 25.6% | +37.6 pp | ~$0.0063est. | 0.24 s | RunPod GPU |
| 34 | decider-2b†The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.by Mapika | 30.7 | 38.5 | 43.5 | 83.2 | 61.0 | 71.0% | 24.7% | +46.3 pp | ~$0.020est. | 0.26 s | RunPod GPU |
| 35 | jeff†Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.by Logan Markewich · Logan Markewich, GLiFormer 400M | 30.6 | 36.8 | 67.9 | 63.5 | 76.6 | 62.8% | 33.1% | +29.7 pp | ~$0.0060est. | 0.94 s | CPU |
| 36 | Laya†The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.by Convai Innovations · Convai Innovations, ModernBERT-large 421M | 30.3 | 36.1 | 63.7 | 71.1 | 86.2 | 58.4% | 30.8% | +27.6 pp | ~$0.0029est. | 0.79 s | CPU |
| 37 | lev-350m†Weights franckverrot/lev-350m revision ab08ad8b8f346994d983152917e114224f6adac7, code github.com/franckverrot/lev c48a945dbf629998d7458dcc5c16f58df964db94, the author's own lev.serve /v1/systemone endpoint with its shipped calibration temperature, base LiquidAI/LFM2.5-350M, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.by Franck Verrot (franckverrot) · Franck Verrot, LFM2.5-350M | 28.5 | 34.8 | 70.6 | 85.3 | 76.1 | 58.4% | 25.0% | +33.4 pp | ~$0.0063est. | 0.17 s | RunPod GPU |
| 38 | 27.7 | 49.4 | 69.2 | 77.1 | 36.5 | 85.3% | 33.1% | +52.2 pp | ~$0.131est. | 0.68 s | author demo | |
| 39 | Von†The author disclosed that its temperature calibration map used the 231 public JevBench items; monotonic scaling does not change accuracy. Estimated cost is USD 0.00551 per 1,000 decisions.by wfzyx (Victor Hugo) · wfzyx, Option-Marker 395M | 27.5 | 34.5 | 75.7 | 70.5 | 77.8 | 57.1% | 27.9% | +29.2 pp | ~$0.0055est. | — | unknown |
| 40 | NInfer Qwen3.8-27B NVFP4†T=1.5 is an offline recomputation from raw logits of the same run, not a second execution. T=1.5 was chosen by sweeping public hard items (development-set calibrated, disclosed).by Igor L. / NInfer contributors · T=1.5 | 26.9 | 51.5 | 76.0 | 80.1 | 35.2 | 83.1% | 33.1% | +50.0 pp | ~$0.145est. | 0.37 s | RunPod GPU |
| 41 | NInfer Qwen3.8-27B NVFP4†Raw T=1.0 row. Author discloses repeated public-item consultation and public-hard tuning (applies to both NInfer 27B rows).by Igor L. / NInfer contributors | 26.3 | 51.5 | 67.2 | 80.1 | 35.2 | 83.1% | 33.1% | +50.0 pp | ~$0.145est. | 0.37 s | RunPod GPU |
| 42 | kev 8B†Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.by Jared Palmer · research preview | 25.6 | 41.8 | 40.2 | 74.9 | 44.0 | 71.4% | 21.8% | +49.7 pp | ~$0.073est. | 0.59 s | RunPod GPU |
| 43 | 25.5 | 47.6 | 77.6 | 88.5 | 35.9 | 89.6% | 33.8% | +55.8 pp | ~$0.137est. | 0.09 s | RunPod GPU | |
| 44 | SimpleJev Qwen3.6-35B-A3B†Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.APIby Featherless AI | 24.9 | 45.7 | 59.8 | 75.0 | 38.1 | 81.4% | 28.2% | +53.1 pp | ~$0.116est. | 0.85 s | author demo |
| 45 | kev 0.6B†Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 307/308 sealed items answered validly (failures count as wrong)by Jared Palmer · research preview | 24.8 | 34.2 | 50.0 | 75.6 | 76.1 | 66.7% | 24.0% | +42.6 pp | ~$0.0063est. | 0.59 s | RunPod GPU |
| 46 | Raw Qwen3 8B direct logits†Neutral raw-logit control.by Alibaba Qwen / neutral reproduction | 23.7 | 45.7 | 24.1 | 86.3 | 41.9 | 68.4% | 26.3% | +42.1 pp | ~$0.087est. | 0.08 s | RunPod GPU |
| 47 | system-oneby Sean Goedecke · Qwen3-8B, Sean Goedecke | 23.4 | 43.6 | 32.8 | 84.4 | 41.5 | 71.9% | 24.4% | +47.5 pp | ~$0.089est. | 0.17 s | RunPod GPU |
| 48 | OpenDecision†A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.by Deepan Wadhwa · ModernBERT-large zero-shot | 21.6 | 31.8 | 57.1 | 79.9 | 75.3 | 53.2% | 25.6% | +27.6 pp | ~$0.0066est. | 0.34 s | RunPod GPU |
| 49 | LitJev†The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.by Zhengxu Yu · Qwen3.8-27B | 19.5 | 46.3 | 76.6 | 66.7 | 33.6 | 86.1% | 30.8% | +55.3 pp | ~$0.163est. | 2.03 s | RunPod GPU |
| 50 | openJev Verdict 1.4†Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.by Hemant (heman10x) | 19.0 | 29.4 | 72.0 | 78.1 | 82.4 | 57.6% | 27.9% | +29.7 pp | ~$0.0039est. | 0.31 s | CPU |
| 51 | kev 0.5B†Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release. 307/308 sealed items answered validly (failures count as wrong)by Jared Palmer | 18.9 | 30.5 | 49.7 | 77.0 | 76.1 | 49.4% | 27.3% | +22.1 pp | ~$0.0063est. | 0.43 s | RunPod GPU |
| 52 | Bespoke Nimble 9B†Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.by Bespoke Labs | 18.7 | 46.3 | 56.4 | 78.7 | 33.4 | 79.7% | 28.9% | +50.8 pp | ~$0.166est. | 0.39 s | RunPod GPU |
| 53 | 18.5 | 93.1 | 87.4 | 77.5 | 28.5 | 97.4% | 89.0% | +8.4 pp | $0.242 | 0.97 s | API | |
| 54 | openJev Verdict†The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.by Hemant (heman10x) · heman10x, ModernBERT-base 151M | 18.1 | 30.0 | 47.0 | 76.7 | 83.1 | 55.4% | 24.7% | +30.7 pp | ~$0.0037est. | 0.28 s | CPU |
| 55 | Raw Qwen3 1.7B direct logits†Neutral raw-logit control.by Alibaba Qwen / neutral reproduction | 18.1 | 33.2 | 24.3 | 89.7 | 64.9 | 54.1% | 26.0% | +28.1 pp | ~$0.015est. | 0.07 s | RunPod GPU |
| 56 | reflex-27b†The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.by kshetrajna12 · Qwen3.8-27B | 17.8 | 46.4 | 77.2 | 67.5 | 32.3 | 87.0% | 29.5% | +57.5 pp | ~$0.181est. | 1.89 s | RunPod GPU |
| 57 | djev†Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill. 220/308 sealed items answered validly (failures count as wrong)by David Villalon / Maisa · thinking | 15.2 | 71.6 | 87.8 | 75.2 | 26.9 | 87.4% | 60.1% | +27.4 pp | ~$0.274est. | 0.43 s | RunPod GPU |
| 58 | GLiNER2 large†The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.by Fastino | 15.1 | 31.1 | 24.8 | 61.7 | 73.3 | 56.7% | 28.6% | +28.1 pp | ~$0.0077est. | 1.10 s | CPU |
| 59 | OpenJev†OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.by razorback16 · thinking, BF16 | 14.8 | 58.1 | 58.1 | 76.1 | 27.8 | 88.7% | 42.2% | +46.5 pp | ~$0.255est. | 0.46 s | RunPod GPU |
| 60 | Qwen3.5-0.8B Decision Model†JevLite SystemOne on CPU; bundled per-question calibration; no operator endpoint or network access. Local CPU measurement on the existing 534-decision v1.3 set plus 308 sealed v1.4 decisions; no operator endpoint received sealed text.by Mourad Ghafiri | 14.5 | 28.1 | 68.2 | 49.2 | 75.7 | 59.3% | 34.7% | +24.6 pp | ~$0.0065est. | 7.15 s | CPU |
| 61 | 14.3 | 54.5 | 59.3 | 81.8 | 27.4 | 87.0% | 38.6% | +48.4 pp | $0.264 | 0.76 s | API | |
| 62 | open-jev-deberta-v3-large†297/308 sealed items answered validly (failures count as wrong)by Kotoba Labs · local CPU | 12.6 | 25.6 | 66.6 | 66.0 | 74.0 | 52.4% | 29.5% | +22.8 pp | ~$0.0073est. | 1.77 s | CPU |
| 63 | smalljev semantic-v9†The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.by Aditya (isHeSatoshi) | 12.3 | 25.7 | 59.2 | 79.8 | 57.9 | 60.6% | 26.9% | +33.7 pp | ~$0.025est. | 0.41 s | RunPod GPU |
| 64 | 11.8 | 27.4 | 25.2 | 71.8 | 83.1 | 58.0% | 29.2% | +28.8 pp | ~$0.0037est. | 0.31 s | CPU | |
| 65 | Open-Jev 9B†The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.by Zefan Cai (@Zefan_Cai) | 11.2 | 44.2 | 61.8 | 72.0 | 28.1 | 77.5% | 29.9% | +47.6 pp | ~$0.249est. | 0.75 s | RunPod GPU |
| 66 | Open-Jev 2B†The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.by Zefan Cai (@Zefan_Cai) | 10.0 | 42.3 | 55.3 | 73.5 | 28.1 | 64.5% | 26.3% | +38.2 pp | ~$0.249est. | 0.66 s | RunPod GPU |
| 67 | 9.8 | 23.1 | 57.2 | 67.8 | 82.4 | 48.9% | 32.8% | +16.1 pp | ~$0.0039est. | 0.43 s | CPU | |
| 68 | SimpleJev†143 pinned files scanned: 0 exact matches. No JevBench-specific fine-tuning.by sabeel111 / Featherless AI · Qwen3.5-0.8B, CPU | 7.5 | 21.5 | 49.1 | 57.5 | 68.3 | 54.5% | 34.7% | +19.8 pp | ~$0.011est. | 3.99 s | CPU |
| 69 | GLiNER2.5 small†The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.by Fastino · Fastino, 74M | 7.2 | 20.5 | 50.7 | 77.8 | 82.4 | 45.9% | 28.6% | +17.3 pp | ~$0.0039est. | 0.11 s | CPU |
| 70 | Raw Qwen3 0.6B direct logits†Neutral raw-logit control.by Alibaba Qwen / neutral reproduction | 7.1 | 22.8 | 20.7 | 89.9 | 73.9 | 48.1% | 25.6% | +22.4 pp | ~$0.0074est. | 0.07 s | RunPod GPU |
| 71 | 4.8 | 94.0 | 95.5 | 71.6 | 16.8 | 97.8% | 94.8% | +3.0 pp | $0.594 | 1.42 s | API | |
| 72 | 2.1 | 13.6 | 26.0 | 70.8 | 73.3 | 41.1% | 9.1% | +32.0 pp | ~$0.0077est. | 0.90 s | author demo | |
| 73 | Mixedbread mxbai-rerank-base-v2†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.by Mixedbread | 0.4 | 6.8 | 84.1 | 87.5 | 67.9 | 37.2% | 34.4% | +2.8 pp | $0.012 | 0.07 s | RunPod GPU |
| 74 | BAAI bge-reranker-v2-m3†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.by BAAI | 0.2 | 5.0 | 84.2 | 89.5 | 73.4 | 39.4% | 27.9% | +11.5 pp | $0.0077 | 0.03 s | RunPod GPU |
| 75 | Alibaba GTE Reranker ModernBERT-base†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.by Alibaba-NLP | 0.2 | 4.8 | 78.9 | 90.6 | 69.6 | 33.8% | 33.4% | +0.3 pp | $0.010 | 0.05 s | RunPod GPU |
| 76 | Certo v1†The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.by AltSlate Labs | 0.0 | 0.1 | 83.0 | 94.0 | 100.0 | 31.6% | 29.5% | +2.1 pp | ~$0.0010est. | 0.02 s | RunPod GPU |
| 77 | Open Jev JSON Canvas†Returns only a final label, not a probability distribution — calibration counts as 0 in the composite. Scan of 39 files: 0 matches.by JoshuaSP | 0.0 | 48.2 | 0.0 | 84.1 | 45.6 | 84.4% | 31.2% | +53.2 pp | ~$0.065est. | 0.22 s | RunPod GPU |
| — | classifier.dev†Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.APIby mrmps (@michael_chomsky) · fast tier honorable mention · not ranked | 70.8 | 51.6 | 72.4 | 87.6 | 84.3 | 85.3% | 34.4% | +50.9 pp | ~$0.0033est. | 0.39 s | API |
| — | partial · not ranked | 0.0 | 40.4 | 93.6 | 61.3 | 0.0 | 71.9% | 21.8% | +50.1 pp | ~$2.669est. | 5.75 s | API |
| — | swanOne†Draft vocabulary derived via AGPL-3.0 generator (provenance recorded separately). Submitter consulted all 231 public tasks and swept temperature over 111 public hard tasks. Score corrected during review from pooled-534 ECE/latency to hard-tier/242-item block. v1.4: Not re-measured: needs its own H100 NVL pod (99 GiB weights); RunPod balance ran low during this job. Ready-to-run recipe kept.by swanOne submitter partial · not ranked | — | — | — | — | — | 88.7% | — | — | ~$0.111est. | 0.34 s | RunPod GPU |
| — | Needle 3, options as tools†V1.4: Not re-measured: same as needle-3.by Cactus Compute · post-hoc adapter mode partial · not ranked | — | — | — | — | — | 22.1% | — | — | ~$0.014est. | 3.78 s | CPU |
| — | Needle 3†V1.4: Not re-measured: ~100-250 s per item on a rented CPU pod (19 s on Sandy); needs a dedicated CPU host.by Cactus Compute · Cactus, 2-bit, local CPU partial · not ranked | — | — | — | — | — | 22.5% | — | — | ~$0.024est. | 1.69 s | CPU |
API = the operator's endpoint received sealed item text during evaluation; the answers and item-level results are not published. The sealed text and answers remain private; only system-level aggregates appear here. Cost is per 1,000 decisions. Hover endpoint, cost and API labels for their recorded details.
All 72 system notes and disclosures
- † JevK5 v0.2.0: Author says no JevBench items or outputs were used for training, tuning, or selection; public results are reported. Scan found only one generic instruction shared by 8 public hard items; unreleased teacher/replay corpora were unavailable.
- † Hopper: Author discloses heavy public-benchmark-directed development (26 model/prompt configs, 20+ calibration-map variants observed against the public half); scan of 17 released files vs 231 public tasks found 0 state/instruction matches; training corpus not released so overlap not independently verifiable.
- † Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
- † reflex 4B (kshetrajna12): The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
- † djev (Maisa, diffusion-gemma): The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated). v1.4: hosted api.djev.dev was paused by its operator ("Serving is paused by the administrator"); sealed tier measured on the public djev runtime (Davipar/djev-dev 3ce907e, same weights, default mode) self-hosted on an H100; Speed/Cost kept from v1.3.
- † Jev-Omni (akhilaaa3, Gemma-4-12B merged): Akhilaaa3/Jev-Omni revision c050d51354147985d13286cf4acf90f562f2c631, the author's own load_model.py (merged text decision model + 256-way head) and his own predict(), transformers 5.17.0 / torch 2.8.0 from the pod image; built on the CPU and moved to CUDA with every nn.Linear weight cast to bfloat16 first - the same cast his reference loader jev_omni.py applies - because our 46 GB GPU cannot hold his fp32 copy; on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
- † metask-jev-4b: Model card discloses 44.8k+16.1k+390 training rows incl. synthetic families intentionally mirroring JevBench hard families, plus repeated evaluation on all 231 public items (benchmark-directed development disclosed). Scan of 61 files: 0 exact matches.
- † Jobe Qwen3.5-4B (frozen): No trained weights/LoRA/calibration fit; release explicitly rejects fitted temperature/order averaging. Scan of 37 files vs 231 public tasks: 0 matches.
- † local-jev Qwen3.5-4B: Scan of 72 files: 0 matches. Author discloses choosing JSON layout after observing results on the 231 public items (benchmark-directed choice, disclosed).
- † spark-s1-4b-v6 (Open Spark Jev, abhishek085): Abhishek085/spark-s1-4b-v6 revision 93d49ddbfb29212e3296635a75a3e80cf69da027, code github.com/abhishek085/open-spark-jev 30ac6d89b7fa36c644cf86aac68f35c1d276a919, the author's own MenuScorer.decide with his fitted calibration.json temperature, bf16, base Qwen/Qwen3.5-4B, transformers 5.17.0 / torch 2.8.0 from the pod image, flash-linear-attention 0.5.2 installed, causal_conv1d not installable here (no wheel builds against this toolchain), on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
- † jqv (Qwen3-32B zero-shot): A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
- † Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- † decider-35b-a3b (Mapika): The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
- † Raw Qwen3 4B Instruct 2507 direct logits: Neutral raw-logit control.
- † OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp): DM submission, measure-only (JevBench publishing HOLD in force). Same author as the existing simplejev-qwen3.5-0.8b row (sabeel111/Featherless AI). Round-4 audit of an earlier commit could not be measured (no public GGUF, Windows-only DLL loader, unpinned llama.cpp build); this round the author published the exact unsloth Q4_K_M GGUF (hash/size independently verified) and we built llama.cpp CUDA from current upstream master on Linux ourselves -- its ABI matched the ctypes bindings exactly, so only a loader file-naming fix was needed (documented diff), no code/scoring/calibration change. Our public-231 subset exactly reproduced the author-reported table: easy 48/48, standard 67/72, hard 66/111, schema 231/231. Calibration (noul temperature) fit only on Google BoolQ, not JevBench. Repo docs name 3 public task IDs while describing benchmark-directed algorithm fixes on the public half (disclosed); 0 exact state/instruction text matches in a released-file scan.
- † ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- † decision-machine-1 (milliseconds.ai): A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
- † Raw Phi-4 mini direct logits: Neutral raw-logit control, not JevBench-directed.
- † JEV Qwen3.5-9B Base NVFP4: Byte-identical to upstream March-2026 NVFP4 checkpoint, predates JevBench v1.2, no task-specific training added. Had 3 preliminary scope/policy scoring errors corrected during review (pooled-ECE, full-run latency, cost estimate); score above is final corrected value.
- † kev 4B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 306/308 sealed items answered validly (failures count as wrong)
- † Decision 2B (FlyMy.AI, v59): Flymy-ai/decision-2b-preview revision df57b75db927acc9ad91ec8115508c1e487086eb (checkpoint minicpm5_reduced_v16_4k_v59), base openbmb/MiniCPM5-2B revision 12a3808a956f869c767195e9266b59c4d21d92e2, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
- † Qwen3.5-9B Jev-like data-mix v2: The author disclosed development on the public JevBench set and public-result comparisons; released-data overlap scan found no matches, but 764 gap and 382 replay training rows are unreleased.
- † GPT-6 Luna (low reasoning effort): OpenAI direct API baseline; reasoning effort low; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
- † SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
- † NInfer Qwen3.8-Flash-Next mixed: Engine scan (1,758 files) vs 231 public tasks: 0 exact matches. Submitter discloses no training for Flash-Next but repeated consultation of public items and a public-hard temperature sweep.
- † GPT-6 Luna (default medium reasoning effort): OpenAI direct API baseline; reasoning effort default medium; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
- † open-alternative-jev (Qwen3.5-4B, IkerMoel): With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
- † jev-local (Qwen3.5-9B): The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
- † Decision Fast (FlyMy.AI, v53a): Flymy-ai/decision-fast-preview revision 4225d41c66119fe28e95a2631bb0103decae6d56 (checkpoint qwen3_06b_headfirst_ep2a_v53), base Qwen/Qwen3-0.6B-Base revision da87bfb608c14b7cf20ba1ce41287e8de496c0cd, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
- † decider-2b (Mapika): The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
- † jeff (Logan Markewich, GLiFormer 400M): Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
- † Laya (Convai Innovations, ModernBERT-large 421M): The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
- † lev-350m (Franck Verrot, LFM2.5-350M): Weights franckverrot/lev-350m revision ab08ad8b8f346994d983152917e114224f6adac7, code github.com/franckverrot/lev c48a945dbf629998d7458dcc5c16f58df964db94, the author's own lev.serve /v1/systemone endpoint with its shipped calibration temperature, base LiquidAI/LFM2.5-350M, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
- † Von (wfzyx, Option-Marker 395M): The author disclosed that its temperature calibration map used the 231 public JevBench items; monotonic scaling does not change accuracy. Estimated cost is USD 0.00551 per 1,000 decisions.
- † NInfer Qwen3.8-27B NVFP4 (T=1.5): T=1.5 is an offline recomputation from raw logits of the same run, not a second execution. T=1.5 was chosen by sweeping public hard items (development-set calibrated, disclosed).
- † NInfer Qwen3.8-27B NVFP4: Raw T=1.0 row. Author discloses repeated public-item consultation and public-hard tuning (applies to both NInfer 27B rows).
- † kev 8B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
- † JevOne: Scan vs 231 public tasks: 0 matches. Training corpus/provenance not disclosed — overlap unknown.
- † SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
- † kev 0.6B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 307/308 sealed items answered validly (failures count as wrong)
- † Raw Qwen3 8B direct logits: Neutral raw-logit control.
- † OpenDecision (ModernBERT-large zero-shot): A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
- † LitJev (Qwen3.8-27B): The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
- † openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
- † kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release. 307/308 sealed items answered validly (failures count as wrong)
- † Bespoke Nimble 9B (Bespoke Labs): Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
- † openJev Verdict (heman10x, ModernBERT-base 151M): The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
- † Raw Qwen3 1.7B direct logits: Neutral raw-logit control.
- † reflex-27b (Qwen3.8-27B): The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
- † djev (thinking): Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill. 220/308 sealed items answered validly (failures count as wrong)
- † GLiNER2 large (Fastino): The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
- † OpenJev (thinking, BF16): OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
- † Qwen3.5-0.8B Decision Model (Mourad Ghafiri): JevLite SystemOne on CPU; bundled per-question calibration; no operator endpoint or network access. Local CPU measurement on the existing 534-decision v1.3 set plus 308 sealed v1.4 decisions; no operator endpoint received sealed text.
- † open-jev-deberta-v3-large (local CPU): 297/308 sealed items answered validly (failures count as wrong)
- † smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
- † GLiNER2 (Fastino, gliner2.5-base): A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
- † Open-Jev 9B (Zefan Cai): The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
- † Open-Jev 2B (Zefan Cai): The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
- † GLiNER2.5 multi (Fastino, 287M): The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
- † SimpleJev (Qwen3.5-0.8B, CPU): 143 pinned files scanned: 0 exact matches. No JevBench-specific fine-tuning.
- † GLiNER2.5 small (Fastino, 74M): The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
- † Raw Qwen3 0.6B direct logits: Neutral raw-logit control.
- † Mirror: 171 intended HTTP 422 context rejections (over 512-token limit) counted once each as misses; only 363/534 valid distributions returned. Found via Gmail submission (Lewis). 102/308 sealed items answered validly (failures count as wrong)
- † Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- † BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- † Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- † Certo v1 (AltSlate Labs): The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
- † Open Jev JSON Canvas (JoshuaSP): Returns only a final label, not a probability distribution — calibration counts as 0 in the composite. Scan of 39 files: 0 matches.
- † classifier.dev (fast tier): Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.
- † swanOne: Draft vocabulary derived via AGPL-3.0 generator (provenance recorded separately). Submitter consulted all 231 public tasks and swept temperature over 111 public hard tasks. Score corrected during review from pooled-534 ECE/latency to hard-tier/242-item block. v1.4: Not re-measured: needs its own H100 NVL pod (99 GiB weights); RunPod balance ran low during this job. Ready-to-run recipe kept.
- † Needle 3, options as tools (post-hoc adapter mode): V1.4: Not re-measured: same as needle-3.
- † Needle 3 (Cactus, 2-bit, local CPU): V1.4: Not re-measured: ~100-250 s per item on a rented CPU pod (19 s on Sandy); needs a dedicated CPU host.
Rows without a † have no note beyond the shared provenance: every row was measured or re-run with its recorded recipe, and deviations are in its run manifest.
Artifact: v1.4.1 results JSON · SHA-256 808909c4ecc8… · JevBench v1.4.1 release and method
Compare two systems
Pick any two. Four radars: the score axes, accuracy per tier including the sealed set, and accuracy by family on the v1.2 hard tier and on the sealed set. Further out is better on every spoke; the link keeps the pair.
- A: Jev 1.13.0 — Jev (TypeSafe, closed) · Score 63.3 (#1)
- B: JevK5 v0.2.0 — Jev rebuild · Score 62.0 (#2)
The four score axes
Accuracy per tier, incl. sealed
Hard tier by family (v1.2 topics)
Sealed set by family
All values as a table
| Spoke | A: Jev 1.13.0 | B: JevK5 v0.2.0 |
|---|---|---|
| The four score axes | ||
| Intelligence | 53.1 | 48.9 |
| Calibration | 76.3 | 74.5 |
| Speed | 83.3 | 91.1 |
| Cost | 52.0 | 59.5 |
| Accuracy per tier, incl. sealed | ||
| Easy | 100% | 100% |
| Standard | 99% | 96% |
| Judge | 95% | 95% |
| Hard | 74% | 70% |
| Sealed | 37% | 33% |
| Hard tier by family (v1.2 topics) | ||
| Adversarial | 100% | 92% |
| Ambiguous | 79% | 79% |
| Judge | 79% | 82% |
| Long policy | 61% | 50% |
| Multi-hop | 86% | 77% |
| Probability | 80% | 60% |
| Routing | 100% | 100% |
| Temporal / numeric | 27% | 37% |
| Trade-off | 92% | 83% |
| Trap | 100% | 100% |
| Sealed set by family | ||
| Ambiguous / abstain | 30% | 43% |
| Judge | 34% | 37% |
| Long policy | 28% | 23% |
| Multi-hop | 45% | 39% |
| Paraphrase | 64% | 14% |
| Probability | 50% | 39% |
| Safety judge | 38% | 38% |
| Temporal / numeric | 29% | 27% |
| Trade-off | 38% | 23% |
| Trap / adversarial | 42% | 58% |
What changed in v1.4
- Fresh sealed decisions keep the benchmark moving as public items saturate. Sealed items contribute 20% of Intelligence:
I = 0.8 × I_v1.3 + 0.2 × I_sealed, whereI_sealed = 100 × max(0, (acc_sealed − 0.293) / (1 − 0.293)). Public and sealed scores are published only as aggregates. - Calibration blends toward the sealed-inclusive result at the approved weight:
C = C_v1.3 + (C_v1.4 − C_v1.3) × min(1, 0.2 / 0.35). - The
k = 1generalization penalty reduces Intelligence when public accuracy exceeds sealed accuracy by more than 25 percentage points:I × (1 − max(0, gap − 25) / 100). It rewards systems that generalize beyond the public half. - The four axes use an equal-weight harmonic mean (
p = −1). Intelligence below 50 keeps its quadratic penalty; Speed and Cost each have a separate Jev-class gate below 50. Speed and Cost axis calculations are unchanged from v1.3.0. - The visible API flag discloses when an operator endpoint received held-out item text, without answers. Existing system notes preserve disclosures such as Hopper's public-half development and JevK5's public-set selection.
What the run says
- Jev 1.13.0 leads v1.4.1 with 63.3: Intelligence 53.1, Calibration 76.3, Speed 83.3, Cost 52.0 ($0.040 per 1,000 decisions).
- The best open or open-planned rebuild, JevK5 v0.2.0, is #2 at 62.0 — 1.3 points behind.
- GPT-6 Luna has the highest Intelligence (97.4) but places #30: Speed 72.6, Cost 36.0 — the harmonic mean does not let accuracy buy back a weak axis.
- The sealed set is hard for everyone: the best sealed accuracy among ranked systems is 95.5% (GPT-6 Luna, #30); chance is 29.3%. Large public-minus-sealed gaps above 25 points reduce Intelligence.
- classifier.dev scores 70.8 but is not ranked: it is a service running another entrant's model.
- swanOne, Qwen3.8 27B, Needle 3, options as tools, Needle 3 did not complete every tier; they are listed without a rank.
Jev alternatives, open source and self-hosting
The chart and table above compare the tested systems, not marketing claims. These are the practical answers readers most often need before choosing a Jev-class decision model.
What are open-source alternatives to Jev?
The highest-ranked open entrants in this run are Hopper (#3, 59.4), Winnow-12B Q8 (#4, 55.6), reflex 4B (#5, 54.0), djev (#6, 52.2). “Open” here means the tested row publishes code or weights; check the licence and exact configuration in the board before adopting one.
Which Jev-class models can I self-host in the EU or use for GDPR-sensitive work?
Open entrants with released code or weights can run on infrastructure you choose, including EU infrastructure. That can support data residency, but neither open source nor an EU server makes a deployment GDPR-compliant by itself. Assess your data, contracts, retention, subprocessors and security for the complete setup. See Benchmark Heaven's broader EU-hosting comparison.
jev-router.com offers self-hosted open decision models. Neutrality disclosure: it is run by the authors of this benchmark; it receives no scoring advantage and is not a ranked entrant.
How is JevBench scored?
The official score is the equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost. Version v1.4.1 measures 534 public and 308 sealed decisions: 20% of Intelligence comes from the sealed set, a public-minus-sealed gap above 25 points costs Intelligence, and Intelligence, Speed or Cost below 50 each pull the score down quadratically. What changed in v1.4 · method and tiers.
How do I submit my model?
Open an issue in the JevBench repository with a reproducible endpoint or runnable code, the exact model and licence, and whether public JevBench items were used during development. New entrants use the same frozen harness and appear in a new version. For private data, see the custom evaluation options.
What a decision costs
Every price here is US dollars per 1,000 decisions — not per 1,000 tokens. One decision is a whole question — state, rubric and options — about 950 input tokens for Jev 1.13.0, so at its $0.042 per million input tokens 1,000 decisions cost $0.0399.
How costs are estimated
One decision is a whole question, not a token. Jev 1.13.0 reads 950 input tokens per decision on average over the 534 v1.2 decisions. At its public tariff of $0.042 per MILLION input tokens (output tokens are free, https://docs.typesafe.ai/models), 1,000 decisions therefore cost 950 x 1,000 x $0.042 / 1,000,000 = $0.0399. That is what the Cost column shows: $0.0399 per 1,000 decisions, not per 1,000 tokens.
Systems with a public tariff (per token or per request) are priced at that tariff times the tokens we measured. Systems without one — open weights, author demos, models we ran locally — are priced as if a large inference provider hosted them: the OpenRouter list price of the same weights; if OpenRouter does not list them, the nearest larger sibling; if no model of that size class is on OpenRouter, the DeepInfra list price of the same weights or of the nearest larger model of the same class. We do not use per-minute GPU rental or our own CPU time — providers buy capacity in bulk or own the hardware, and price accordingly. Price × tokens per decision = $ per 1,000 decisions, marked “est.”.
- JevK5 v0.2.0 — ~$0.022 est. per 1,000 decisions: Reconstructed from the frozen v1.3 Cost axis; same speed/cost measurement, not a new price observation
- Hopper — ~$0.024 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
- Winnow-12B Q8 — ~$0.037 est. per 1,000 decisions: OpenRouter google/gemma-3-12b-it hosted reference list price $0.05/M in, $0.0/M out (the nearest publicly hosted 12B Gemma sibling; Winnow reads answer logits in one forward pass and generates no answer tokens) x 393 input and 0 output tokens per decision (input tokens measured (the system's own count))
- reflex 4B — ~$0.022 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.0/M out (the exact base weights; one pass, no generated output) x 377 input and 0 output tokens per decision (input tokens measured (the system's own count))
- Jev-Omni — ~$0.037 est. per 1,000 decisions: OpenRouter Gemma 3 12B input rate list price $0.05/M in, $0.0/M out (a 12B one-pass model with no generated output; the same reference the author uses in his own model card for this model) x 384 input and 0 output tokens per decision (input tokens measured (the system's own count))
- metask-jev-4b — ~$0.033 est. per 1,000 decisions: hosted 4B reference rate USD 0.04/M input, USD 0/M output over the exact measured prompt-token counts of all 534 attempts; no generated answer tokens
- SemIf — ~$0.022 est. per 1,000 decisions: deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (same weights (not on OpenRouter), as open-alternative-jev in v1.1.2) x 396 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1244 in / 0 out tokens per hard decision
- Jobe Qwen3.5-4B — ~$0.022 est. per 1,000 decisions: DeepInfra Qwen3.5-4B hosted reference list price $0.03/M in, $0.0/M out (same underlying weights; one forward pass, no generated tokens) x 396 input and 0 output tokens per decision (input tokens measured (the system's own count))
- local-jev Qwen3.5-4B — ~$0.030 est. per 1,000 decisions: hosted 4B reference rate list price $0.04/M in, $0.0/M out (one forward pass over measured input tokens and no generated answer tokens) x 397 input and 0 output tokens per decision (input tokens measured (the system's own count))
- system-one-open — ~$0.015 est. per 1,000 decisions: deepinfra google/gemma-4-E4B-it list price $0.02/M in, $0.1/M out (Gemma 4 E2B is not listed; the nearest larger sibling, Gemma 4 E4B, is listed only on DeepInfra) x 383 input and 2 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra google/gemma-4-E4B-it $0.02/M in, $0.1/M out x 1235 in / 2 out tokens per hard decision
- spark-s1-4b-v6 — ~$0.025 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B size-class reference list price $0.03/M in, $0.0/M out (a 2-4B one-pass model with no generated output; the board's 4B open-weights reference, as for reflex 4B and decider-2b) x 505 input and 0 output tokens per decision (input tokens measured (the system's own count))
- jqv — ~$0.056 est. per 1,000 decisions: OpenRouter qwen/qwen3-32b list price $0.08/M in, $0.0/M out (the exact base model this system reads logits from; nothing is generated) x 359 input and 0 output tokens per decision (input tokens measured (the system's own count))
- decider-35b-a3b — ~$0.067 est. per 1,000 decisions: OpenRouter Qwen3.6-35B-A3B list price list price $0.1/M in, $0.0/M out (the closest public hosted 35B-A3B direct-logit model; no output is generated) x 312 input and 0 output tokens per decision (input tokens measured (the system's own count))
- Raw Qwen3 4B Instruct 2507 direct logits — ~$0.022 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
- OpenSourceJev — ~$0.016 est. per 1,000 decisions: Reconstructed from the frozen v1.3 Cost axis; same speed/cost measurement, not a new price observation
- Raw Phi-4 mini direct logits — ~$0.048 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
- JEV Qwen3.5-9B Base NVFP4 — ~$0.077 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
- OpenJev — ~$0.066 est. per 1,000 decisions: openrouter google/gemma-4-26b-a4b-it list price $0.09/M in, $0.3/M out (DiffusionGemma 26B-A4B is not listed; the same-size Gemma 4 26B-A4B MoE sibling is (size class moe_26B-A4B)) x 380 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter google/gemma-4-26b-a4b-it $0.09/M in, $0.3/M out x 1222 in / 0 out tokens per hard decision
- kev 4B — ~$0.019 est. per 1,000 decisions: DeepInfra Qwen3.5-4B size-class reference list price $0.03/M in, $0.0/M out (a 4B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
- Decision 2B — ~$0.018 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B size-class reference list price $0.03/M in, $0.0/M out (a 2-4B one-pass model with no generated output; the board's 4B open-weights reference, as for reflex 4B and decider-2b) x 269 input and 0 output tokens per decision (input tokens measured (the system's own count))
- Qwen3.5-9B Jev-like data-mix v2 — ~$0.083 est. per 1,000 decisions: Reconstructed from the frozen v1.3 Cost axis; same speed/cost measurement, not a new price observation
- SimpleJev Qwen3.8-27B — ~$0.104 est. per 1,000 decisions: OpenRouter Gemma 4 26B-A4B size-class reference list price $0.09/M in, $0.0/M out (a public 27B dense model served as a direct-logit classifier; no output is generated) x 809 input and 0 output tokens per decision (input tokens measured (the system's own count))
- NInfer Qwen3.8-Flash-Next mixed — ~$0.109 est. per 1,000 decisions: OpenRouter qwen/qwen3.8-flash hosted list reference list price $0.15/M in, $0.0/M out (same underlying Flash-Next weights; native one-pass option-logit readout) x 366 input and 0 output tokens per decision (input tokens measured (the system's own count))
- open-alternative-jev — ~$0.022 est. per 1,000 decisions: deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (as open-alternative-jev) x 383 input and 1 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1235 in / 1 out tokens per hard decision
- jev-local — ~$0.077 est. per 1,000 decisions: OpenRouter qwen/qwen3.5-9b list price $0.1/M in, $0.0/M out (the exact base weights; scored by log-probabilities, nothing is generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
- Decision Fast — ~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output, the same reference the kev 0.5B/0.6B rows use) x 280 input and 0 output tokens per decision (input tokens measured (the system's own count))
- decider-2b — ~$0.020 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.0/M out (no hosted ~2B Qwen3.5 is listed, so the 4B price is used and errs high; one pass, no output) x 312 input and 0 output tokens per decision (input tokens measured (the system's own count))
- jeff — ~$0.0060 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 272 input and 0 output tokens per decision (input tokens measured (the system's own count))
- Laya — ~$0.0029 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 205 input and 0 output tokens per decision (input tokens measured (the system's own count))
- lev-350m — ~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output, the same reference the kev 0.5B/0.6B rows use) x 280 input and 0 output tokens per decision (input tokens measured (the system's own count))
- openjev-sglang — ~$0.131 est. per 1,000 decisions: openrouter qwen/qwen3.6-35b-a3b list price $0.1/M in, $0.9/M out (same base weights) x 610 input and 2 output tokens per decision [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3.6-35b-a3b $0.1/M in, $0.9/M out x 2272 in / 2 out tokens per hard decision
- Von — ~$0.0055 est. per 1,000 decisions: Reconstructed from the frozen v1.3 Cost axis; same speed/cost measurement, not a new price observation
- NInfer Qwen3.8-27B NVFP4 — ~$0.145 est. per 1,000 decisions: OpenRouter Qwen3.8-27B hosted list reference list price $0.2/M in, $0.0/M out (same underlying weights; native one-pass option-logit readout) x 366 input and 0 output tokens per decision (input tokens measured (the system's own count))
- NInfer Qwen3.8-27B NVFP4 — ~$0.145 est. per 1,000 decisions: OpenRouter Qwen3.8-27B hosted list reference list price $0.2/M in, $0.0/M out (same underlying weights; native one-pass option-logit readout) x 366 input and 0 output tokens per decision (input tokens measured (the system's own count))
- kev 8B — ~$0.073 est. per 1,000 decisions: OpenRouter qwen/qwen3-8b list price list price $0.117/M in, $0.0/M out (the same-size Qwen3-8B weights; kev generates no output tokens) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
- JevOne — ~$0.137 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
- SimpleJev Qwen3.6-35B-A3B — ~$0.116 est. per 1,000 decisions: OpenRouter Qwen3.6-35B-A3B list price list price $0.1/M in, $0.0/M out (the same base weights served as a direct-logit classifier; no output is generated) x 809 input and 0 output tokens per decision (input tokens measured (the system's own count))
- kev 0.6B — ~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
- Raw Qwen3 8B direct logits — ~$0.087 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
- system-one — ~$0.089 est. per 1,000 decisions: openrouter qwen/qwen3-8b list price $0.117/M in, $0.455/M out (same weights, listed on OpenRouter) x 412 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3-8b $0.117/M in, $0.455/M out x 1258 in / 1 out tokens per hard decision
- OpenDecision — ~$0.0066 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 329 input and 0 output tokens per decision (input tokens measured (the system's own count))
- LitJev — ~$0.163 est. per 1,000 decisions: OpenRouter Qwen3.8-27B (as the reflex-27b row) list price $0.214/M in, $0.0/M out (the exact base weights; nothing is generated) x 418 input and 0 output tokens per decision (input tokens measured (the system's own count))
- openJev Verdict 1.4 — ~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
- kev 0.5B — ~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
- Bespoke Nimble 9B — ~$0.166 est. per 1,000 decisions: openrouter qwen/qwen3.5-9b list price $0.1/M in, $0.15/M out (a LoRA merge of Qwen3.5-9B; the base weights are listed on OpenRouter (size class dense_9B), as in the v1.1.3 row) x 970 input and 1 output tokens per decision (input tokens measured (the system's own count))
- openJev Verdict — ~$0.0037 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]
- Raw Qwen3 1.7B direct logits — ~$0.015 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
- reflex-27b — ~$0.181 est. per 1,000 decisions: OpenRouter Qwen3.8-27B list price list price $0.214/M in, $0.0/M out (the exact public base weights used as a direct-logit classifier; no output is generated) x 481 input and 0 output tokens per decision (input tokens measured (the system's own count))
- djev — ~$0.274 est. per 1,000 decisions: same-size hosted reference x 749 measured input and 690 measured output tokens per attempted decision across all 534, failures included
- GLiNER2 large — ~$0.0077 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
- OpenJev — ~$0.255 est. per 1,000 decisions: same hosted reference x 1778 billed input and 315 thought output tokens per decision
- Qwen3.5-0.8B Decision Model — ~$0.0065 est. per 1,000 decisions: same-size DeepInfra Qwen3.5-0.8B reference tariff; measured JevLite tokenizer usage; estimated, not charged
- open-jev-deberta-v3-large — ~$0.0073 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) $0.01/M in, $0.0/M out x 1235 in / 0 out tokens per hard decision
- smalljev semantic-v9 — ~$0.025 est. per 1,000 decisions: submitted Qwen/Qwen2.5-3B-Instruct hosted reference list price $0.04/M in, $0.0/M out (the author's documented reference for the same approximate size class; one forward pass, nothing generated) x 329 input and 0 output tokens per decision (input tokens measured (the system's own count))
- GLiNER2 — ~$0.0037 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]
- Open-Jev 9B — ~$0.249 est. per 1,000 decisions: OpenRouter Qwen3.5-9B list price read 2026-09-21 list price $0.1/M in, $0.0/M out (the exact 9B base and a conservative same-family proxy for the unlisted 2B; the decision head generates no output tokens) x 1439 input and 0 output tokens per decision (input tokens measured (the system's own count))
- Open-Jev 2B — ~$0.249 est. per 1,000 decisions: OpenRouter Qwen3.5-9B list price read 2026-09-21 list price $0.1/M in, $0.0/M out (the exact 9B base and a conservative same-family proxy for the unlisted 2B; the decision head generates no output tokens) x 1439 input and 0 output tokens per decision (input tokens measured (the system's own count))
- GLiNER2.5 multi — ~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
- SimpleJev — ~$0.011 est. per 1,000 decisions: Same nonzero hosted size-class reference and exact prompt-token accounting; see RESULT.md
- GLiNER2.5 small — ~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
- Raw Qwen3 0.6B direct logits — ~$0.0074 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
- Mirror — ~$0.0077 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
- Certo v1 — ~$0.0010 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 86 input and 0 output tokens per decision (input tokens measured (the system's own count))
- Open Jev JSON Canvas — ~$0.065 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
- classifier.dev — ~$0.0033 est. per 1,000 decisions: ESTIMATE from the published paid plan (the free tier was used): classifier.dev Pro $20/month for 200,000 fast classifications a day (https://classifier.dev/pricing, read 2026-09-19) = $0.0033 per 1,000 decisions at full use; one decision = one classification. Lower use costs more per decision: at a tenth of that allowance it is $0.033 per 1,000, and the free tier (20,000 fast classifications a day, which is what this run used) costs nothing.
- swanOne — ~$0.111 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
- Qwen3.8 27B — ~$2.669 est. per 1,000 decisions: openrouter qwen/qwen3.8-27b list price $0.214/M in, $2.55/M out (same weights; our run used a flat-rate Chutes subscription) x 416 input and 393 output tokens per decision [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3.8-27b $0.214/M in, $2.55/M out x 1592 in / 1833 out tokens per hard decision
- Needle 3, options as tools — ~$0.014 est. per 1,000 decisions: same per-token price as Needle 3 (openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out) x 383 input and 20 output tokens per decision, over the 314 easy/standard/judge decisions it ran (no hard-tier run). The v1.2 score lab had no price for this row and scored it 100; fixed. [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]
- Needle 3 — ~$0.024 est. per 1,000 decisions: openrouter meta-llama/llama-3.2-1b-instruct list price $0.027/M in, $0.201/M out (no generative model under 1B is listed; the smallest listed one (1B) errs high; about 20 generated tokens for one tool call) x 383 input and 20 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out x 1235 in / 20 out tokens per hard decision
Reference prices by size class ($ per million input / output tokens)
- dense 2-4B: deepinfra Qwen/Qwen3.5-4B $0.03 / $0.15; deepinfra google/gemma-4-E4B-it $0.02 / $0.1; openrouter google/gemma-3-4b-it $0.05 / $0.1; openrouter meta-llama/llama-3.2-3b-instruct $0.05 / $0.33
- dense 27B: openrouter qwen/qwen3.5-27b $0.195 / $1.56; openrouter qwen/qwen3.6-27b $0.3 / $2; openrouter qwen/qwen3.8-27b $0.214 / $2.55
- dense 9B: openrouter qwen/qwen3.5-9b $0.1 / $0.15
- encoder classifier <=0.6B: BAAI/bge-large-en-v1.5 (335M) $0.01; Qwen/Qwen3-Embedding-0.6B $0.01; intfloat/e5-large-v2 (335M) $0.01; intfloat/multilingual-e5-large (560M) $0.01; thenlper/gte-base (110M) $0.005
- generative <=1B: deepinfra meta-llama/Llama-3.2-1B-Instruct $0.005 / $0.01; openrouter meta-llama/llama-3.2-1b-instruct $0.027 / $0.201
- moe 26B-A4B: deepinfra google/gemma-4-26B-A4B-it $0.07 / $0.34; openrouter google/gemma-4-26b-a4b-it $0.09 / $0.3
- moe 35B-A3B: deepinfra Qwen/Qwen3.6-35B-A3B $0.1 / $0.95; openrouter qwen/qwen3.5-35b-a3b $0.1625 / $1.3; openrouter qwen/qwen3.6-35b-a3b $0.1 / $0.9
Sources: OpenRouter https://openrouter.ai/api/v1/models and DeepInfra https://api.deepinfra.com/models/list (both read 2026-09-19).
Correction, v1.2.3 (20 September 2026): every price recomputed, each decision counted once
usd_per_1000_v11_tiers = 1000 x (mean input tokens per decision x $/M in + output tokens charged x $/M out) / 1e6, over all 314 v1.1 decisions (72 easy + 242 standard+judge), each decision counted exactly once and priced exactly once. A metered row uses the provider's own tariff and its own measured token counts, including the requests whose answer could not be parsed; an estimated row uses the reference tariff for its weights or size class and, when the run reports no usage, the input tokens of the gemini-3.1-flash-lite run on the same prompts over the same 314 decisions. usd_per_1000 = (v11 x 314 + hard x 220) / 534.
- The v1.1 and v1.1.3 aggregations built their cost average from a row list that contained the 242-decision standard+judge run twice (once as the standard tier, once as the judge tier) and the 72 easy decisions once: 556 rows instead of 314. The standard and judge tiers were therefore over-weighted in the price, which made the affected rows look 1.5-3.3 % more expensive than they are.
- Rows without their own token counts were priced at the input tokens of the gemini-3.1-flash-lite run measured on the 242 standard+judge decisions only (452 per decision) and that figure was applied to all 314 v1.1 decisions, which excludes the shorter easy tier. Over all 314 decisions the same run averages 383.41 input tokens, which is the figure used from v1.2.3 on. This made the affected rows look 4-11 % more expensive.
- A metered row's price left out the requests whose answer came back unparseable. Those requests returned HTTP 200 with generated tokens and were billed, and JevBench already counts them as wrong answers, so from v1.2.3 they are priced too. Only DeepSeek V4.1 Flash had any (9 of its 314 v1.1 decisions); its price rises by 2.6 %.
- No tariff was wrong. The hard-tier costs, and classifier.dev's flat plan price, were already correct.
No tariff, measurement, item, answer or rank changed. The prices before and after:
- Jev 1.13.0 — $0.0406 → $0.0399 (-1.72 %)
- SemIf — $0.0230 → $0.0224 (-2.31 %)
- system-one-open — $0.0157 → $0.0149 (-5.13 %)
- OpenJev — $0.0672 → $0.0656 (-2.36 %)
- GPT-5.6 Luna — $0.2473 → $0.2419 (-2.17 %)
- openjev-sglang — $0.1346 → $0.1313 (-2.48 %)
- Bespoke Nimble 9B — $0.1085 → $0.1049 (-3.28 %)
- Gemini 3.1 Flash-Lite — $0.2682 → $0.2638 (-1.65 %)
- DeepSeek V4.1 Flash — $0.5788 → $0.5937 (+2.57 %)
- system-one — $0.0915 → $0.0894 (-2.29 %)
- openJev Verdict — $0.0039 → $0.0037 (-5.21 %)
- GLiNER2 — $0.0039 → $0.0037 (-5.21 %)
- open-jev-deberta-v3-large — $0.0077 → $0.0073 (-5.20 %)
- Qwen3.8 27B — $2.7110 → $2.6691 (-1.55 %)
- Needle 3, options as tools — $0.0162 → $0.0144 (-11.39 %)
- Needle 3 — $0.0249 → $0.0238 (-4.36 %)
Who could not be measured, and why
An exclusion is an availability fact about our run — hardware, access, terms — never a quality verdict. Partial runs are in the table above, greyed and without a rank; so are the honorable mentions, which are complete runs that simply are not ranked.
- open-jev (Dasein Labs) — MLX on Apple Silicon only. Its own README says Linux containers cannot reach the Apple GPU, so a RunPod NVIDIA GPU cannot run it.
- open-jev (JoshuaSP) — A DiffusionGemma 26B-A4B serving wrapper rather than new trained weights. It was demonstrated on an H100; no public endpoint exists and no suitable 80 GB RunPod host was available in this round.
- mini-jev (Mikhail Rakutko (r-ms)) — Public weights exist and fit a normal GPU, but the implementation covers Choice/Noul and explicitly does not measure Score. A faithful full-suite adapter would require new interface work rather than a mechanical endpoint adapter.
- system-one-gemma (Akash Kamat) — The adapter is public, but its Gemma base is gated behind Google’s licence terms. We do not accept binding terms on Florian’s behalf.
- jevlike (Vincent Wang-Maścianica) — Only Doom and chess vision checkpoints are released; there is no general text-decision checkpoint for this suite.
- AlexWortega/openjev (Alex Wortega) — Its released NLI and task-specific heads do not define a distribution over an arbitrary supplied label set. Inventing that mapping would measure our assumption.
- Needle 3 (Cactus Compute) — Its native response is a chosen label plus one accept/refuse confidence, not a categorical distribution over the supplied labels. The options-as-tools adaptation remains published as a partial run.
- Succinct Router 14M (Pedro Marques) — A router over three fixed GPT settings, not a general typed-decision model.
- jev-model-router, Director, Loki (various) — Applications built on decision models, not decision models themselves.
- ProgramAsWeights (ProgramAsWeights) — The compiler still requires GitHub authentication and the available path would expose held-out rubrics to a third party. No public weights or anonymous endpoint are available.
- EigenJev (EigenJev) — The endpoint requires authentication and no public weights or runnable implementation are published.
- NanoJev (NanoJev) — Public weights exist, but the server exposes a different schema (including boolean rather than Noul) and lacks the full structured/null contract. It needs substantive compatibility work before a fair full-suite run.
- Werr (pCwOrM) — Its documented server imports a module (scratch.jevbench_eval.optimize_werr_jevbench) that is not in the public repository, so the submitted configuration cannot be started; its engine also sends telemetry about each request to an outside server by default.
- DIY Jev (VakeDomen) — The repository named in the request (github.com/VakeDomen/DIY-Jev) answers 404, so there is nothing to run.
- SimpleJev RWKV variants (SimpleJev) — The public demo exposes RWKV IDs, but it does not identify their exact checkpoints or licences. Without reproducible model provenance, we do not publish benchmark rows for them.
Method and tiers
Built and run by us, not collected from someone else's leaderboard; the results describe the tested configurations, not every application.
JevBench Score. Equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost (power mean p=-1). If Intelligence <50 multiply by (I/50)^2. For Speed and Cost separately, if below 50 multiply by (axis/50)^2.
Intelligence. 0.8 × v1.3 chance-corrected Intelligence on the frozen v1.2 items + 0.2 × 100 × max(0, (sealed accuracy − 0.293)/(1 − 0.293)); then multiply by 1 − max(0, public-minus-sealed accuracy gap in percentage points − 25)/100. Original tier weights: easy .14, standard .28, judge .28, hard .30.
Revision v1.4.1. v1.4.1 adds 6 systems omitted from the v1.4.0 freeze. The v1.4 scoring formulas and prior system measurements are unchanged.
- easy: 72 clear-cut decisions (intent, explicit yes/no fact, enum extraction, one-obvious-tool selection); new in v1.1
- standard: 96 authored decisions from v1.0 (policy, intent, extraction, ordinal, adequacy, routing), unchanged
- judge: 146 imported decisions from v1.0 (routing real task prompts into 9 categories; judging whether a saved math answer is correct), unchanged
- hard: 220 new decisions (111 public, 109 held out): long multi-condition policy documents (2-6k tokens), priority trade-offs, deliberately ambiguous cases with a 'no clear answer' label, traps, multi-hop lookups, date/number reasoning, adversarial distractors, subtle answer-judging, overlapping routing, and probability items with an exact gold distribution. Half written by Claude Opus 5, half by GPT-5.6 Sol; each item reviewed blind and then against its gold by the other model; one discussion round; frozen and hashed before any benchmarked system saw an item. No item was selected on any system's answers.
- sealed: 308 fresh private decisions across ten families, run once per system. Only system-level aggregates — overall and per-family accuracy, calibration — are published; the item text, answers and per-item results stay private and rotate between versions.
Every system sees the same state, instructions, rubric and exact label set; only the transport differs. Requests go out one at a time with no retries, so latency includes the network. Estimated costs are hosted-provider prices for the same weights or size class and are marked “est.” — hover one for its basis, or see how costs are estimated. Every system has a price; none gets a free 100.
A service running another entrant's model is listed, but not ranked against the models. Ranking it would rank the same model twice, once at the model's own price and once at the service's. The row keeps every number, axis, cost basis and per-task outcome; it carries no rank number. Which rows this applies to, and why: Honorable mentions — services built on another entrant's model.
v1.4 scores are not comparable with v1.3 or earlier (sealed blend, gap penalty and harmonic mean). The v1.3.0 board stays below as history; the v1.0 page keeps its own numbers, calibration plots and per-family tables.
Limits
- 842 decisions (534 public, 308 sealed) is a pilot, not a census, and it is English-only. The v1.4 sealed set is very hard: most systems score close to chance on it, so sealed accuracy separates the field less than public accuracy does.
- The weights are a choice. The JevBench Score weights the four axes equally and uses a harmonic mean, so the weakest axis dominates; if a wrong decision costs you more than a slow or expensive one, read the Intelligence column and the accuracy radars rather than the score alone.
- The latency adjustment (×2, +0.15 s on our own servers) is an assumption, not a measurement. We ran the self-hosted and demo endpoints one request at a time (parallelism 1, no other load), so their latency is likely better than the same model on a busy production server. The official Jev API is presumably under high load, given the public interest. Serving under load trades per-user speed for throughput: in the NVIDIA chart shown by SemiAnalysis, moving to the throughput-maximising setting cuts per-user tokens per second by far more than 2×. That chart is a 1.8T mixture-of-experts model on GPU clusters, not a 4B model on one GPU, so it supports the direction and size of the effect, not our exact factor. The +0.15 s stands for infrastructure our self-hosted tests lacked: authentication, load balancing, logging, billing and an API gateway. Both numbers are assumptions; raw p50/p95 latencies are in the table and the repo, and a measurement under load is planned.
- Held-out decisions are sent to the evaluated services to get predictions. Not public is not the same as not seen.
- Latency is one origin at one time of day; hosted endpoints, public demos and a local CPU are different kinds of latency. Public demo endpoints are shared with everyone else using them.
- Estimated costs describe what a large inference provider would charge for a model of that size, not what the author pays; a system on a tariff pays its tariff.
Credit
Harness, public tasks and every scoring rule: github.com/fstandhartinger/jevbench (MIT). Each project links its author's repository or vendor page.
- Alibaba GTE Reranker ModernBERT-base — Alibaba-NLP, Apache-2.0 — huggingface.co/Alibaba-NLP/gte-reranker-modernbert-base
- BAAI bge-reranker-v2-m3 — BAAI, Apache-2.0 — huggingface.co/BAAI/bge-reranker-v2-m3
- Bespoke Nimble 9B (Bespoke Labs) — Bespoke Labs, Apache-2.0 (weights); repository without a licence file as of 19 Sep — github.com/bespokelabsai/nimble
- Certo v1 (AltSlate Labs) — AltSlate Labs, MIT — huggingface.co/altslate/certo-decision-model
- classifier.dev (fast tier) — mrmps (@michael_chomsky), MIT (code); hosted service — classifier.dev
- decider-2b (Mapika) — Mapika, Apache-2.0 — huggingface.co/Mapika/decider-2b
- decider-35b-a3b (Mapika) — Mapika, Apache-2.0 — huggingface.co/Mapika/decider-35b-a3b
- Decision 2B (FlyMy.AI, v59) — FlyMy.AI (@denti), Apache-2.0 notices on the included code and the pinned base; the weights are an evaluation preview under EVALUATION-PERMISSION.md, not a cleared commercial release — huggingface.co/flymy-ai/decision-2b-preview
- Decision Fast (FlyMy.AI, v53a) — FlyMy.AI (@denti), Apache-2.0 notices on the included code and the pinned base; the weights are an evaluation preview under EVALUATION-PERMISSION.md, not a cleared commercial release — huggingface.co/flymy-ai/decision-fast-preview
- decision-machine-1 (milliseconds.ai) — milliseconds.ai (Baptiste Laget), proprietary API, closed weights — www.milliseconds.ai
- DeepSeek V4.1 Flash (thinking default) — DeepSeek, open weights, proprietary API route
- djev (Maisa, diffusion-gemma) — Maisa (David Villalón), Apache-2.0 code; Google DiffusionGemma Apache-2.0 weights; no djev-specific weights — github.com/Davipar/djev-dev
- djev (thinking) — David Villalon / Maisa, Apache-2.0 — github.com/Davipar/djev-dev
- Gemini 3.1 Flash-Lite — Google, proprietary API
- GLiNER2 (Fastino, gliner2.5-base) — Fastino, Apache-2.0 — github.com/fastino-ai/GLiNER2
- GLiNER2 large (Fastino) — Fastino, Apache-2.0 — huggingface.co/fastino/gliner2-large-v1
- GLiNER2.5 multi (Fastino, 287M) — Fastino, Apache-2.0 — huggingface.co/fastino/gliner2.5-multi-v1
- GLiNER2.5 small (Fastino, 74M) — Fastino, Apache-2.0 — huggingface.co/fastino/gliner2.5-small-v1
- GPT-5.6 Luna (low reasoning effort) — OpenAI, proprietary API
- GPT-6 Luna (default medium reasoning effort) — OpenAI, proprietary API
- GPT-6 Luna (low reasoning effort) — OpenAI, proprietary API
- Hopper — HopitAI, Component-specific terms recorded in RESULT.md; submitted adapter release and Qwen base retain their respective terms — huggingface.co/HopitAI/hopper
- jeff (Logan Markewich, GLiFormer 400M) — Logan Markewich, MIT (code); GLiFormer weights per their model card — github.com/logan-markewich/jeff
- Jev 1.13.0 (TypeSafe AI) — TypeSafe AI, proprietary API — docs.typesafe.ai
- JEV Qwen3.5-9B Base NVFP4 — WilfLin, Apache-2.0 entrant and upstream checkpoint — huggingface.co/WIlfLin/JEV-Qwen3.5-9B-Base-NVFP4
- jev-local (Qwen3.5-9B) — us (GitHub), no licence stated in the repository (public code); Apache-2.0 base weights — github.com/us/jev-local
- Jev-Omni (akhilaaa3, Gemma-4-12B merged) — akhilaaa3, Apache-2.0, following Gemma 4; dataset rights stated separately by the author — huggingface.co/akhilaaa3/Jev-Omni
- JevK5 v0.2.0 — allebee, Apache-2.0 (code/adapter); Apache-2.0 (Qwen3.5-4B base) — github.com/allebee/jevk5
- JevOne — Juspay, Component-specific JevOne/Qwen/SGLang terms recorded in RESULT.md — huggingface.co/juspay/jev-one
- Jobe Qwen3.5-4B (frozen) — MantisShrimpdev, MIT code; Apache-2.0 weights — github.com/MantisShrimpdev/jobe
- jqv (Qwen3-32B zero-shot) — hjmurmur (Octalab), Apache-2.0 (Qwen3-32B weights); serving code public — github.com/Octalab-Inc/jqv
- kev 0.5B — Jared Palmer, Apache-2.0 — github.com/jaredpalmer/kev
- kev 0.6B (research preview) — Jared Palmer, Apache-2.0 — github.com/jaredpalmer/kev
- kev 4B (research preview) — Jared Palmer, Apache-2.0 — github.com/jaredpalmer/kev
- kev 8B (research preview) — Jared Palmer, Apache-2.0 — github.com/jaredpalmer/kev
- Laya (Convai Innovations, ModernBERT-large 421M) — Convai Innovations, Apache-2.0 — huggingface.co/convaiinnovations/laya
- lev-350m (Franck Verrot, LFM2.5-350M) — Franck Verrot (franckverrot), Apache-2.0 (code); weights under LiquidAI's LFM1.0 licence, following the LFM2.5-350M base — github.com/franckverrot/lev
- LitJev (Qwen3.8-27B) — Zhengxu Yu, Apache-2.0 (code); Apache-2.0 base weights — github.com/zhengxuyu/litjev
- local-jev Qwen3.5-4B — Amith Chandrappa (amithgc), MIT code; Apache-2.0 Qwen weights — github.com/amithgc/local-jev
- metask-jev-4b — Wayfind (metask-ai), Apache-2.0 — github.com/metask-ai/metask-jev
- Mirror — Bluusun, Apache-2.0 wrapper and Mirror release; upstream DeBERTa/model assets retain their own terms — github.com/Bluusun/Decision-API
- Mixedbread mxbai-rerank-base-v2 — Mixedbread, Apache-2.0 — huggingface.co/mixedbread-ai/mxbai-rerank-base-v2
- Needle 3 (Cactus, 2-bit, local CPU) — Cactus Compute, Apache-2.0 (model and package) — github.com/cactus-compute/needle
- NInfer Qwen3.8-27B NVFP4 — Igor L. / NInfer contributors, Apache-2.0 (engine and submitted 27B artifact/base) — github.com/igorls/ninfer
- NInfer Qwen3.8-27B NVFP4 (T=1.5) — Igor L. / NInfer contributors, Apache-2.0 (engine and submitted 27B artifact/base) — github.com/igorls/ninfer
- NInfer Qwen3.8-Flash-Next mixed — Igor L. / NInfer contributors, Qwen Community License 1.0 (model); Apache-2.0 (engine) — github.com/igorls/ninfer
- Open Jev JSON Canvas (JoshuaSP) — JoshuaSP, MIT code; Apache-2.0 DiffusionGemma weights — github.com/JoshuaSP/open-jev
- open-alternative-jev (Qwen3.5-4B, IkerMoel) — IkerMoel, Apache-2.0 (code and weights) — github.com/ikermoel/open-alternative-jev
- Open-Jev 2B (Zefan Cai) — Zefan Cai (@Zefan_Cai), MIT (loader); Apache-2.0 (adapter and pinned Qwen base); CC0-1.0 public training projection — github.com/Zefan-Cai/Open-Jev
- Open-Jev 9B (Zefan Cai) — Zefan Cai (@Zefan_Cai), MIT (loader); Apache-2.0 (adapter and pinned Qwen base); CC0-1.0 public training projection — github.com/Zefan-Cai/Open-Jev
- open-jev-deberta-v3-large (local CPU) — Kotoba Labs, Apache-2.0 (model card); DeBERTa-v3 keeps its own terms — github.com/kotoba-lang/typed-decisions
- OpenDecision (ModernBERT-large zero-shot) — Deepan Wadhwa, Apache-2.0 — github.com/deepanwadhwa/OpenDecision
- OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16) — razorback16 / Codiv, Apache-2.0 (repo and weights) — github.com/razorback16/openjev
- OpenJev (thinking, BF16) — razorback16, Apache-2.0 — github.com/razorback16/openjev
- openJev Verdict (heman10x, ModernBERT-base 151M) — Hemant (heman10x), Apache-2.0 — github.com/Heman10x-NGU/openJev-verdict-2.0
- openJev Verdict 1.4 — Hemant (heman10x), Apache-2.0 — huggingface.co/heman10x/rlcd-modernbert-151m
- openjev-sglang (Qwen3.6-35B-A3B on SGLang) — ekzhang, no licence file in the repository as of 2026-09-19; Qwen3.6 weights keep their own terms — github.com/ekzhang/openjev-sglang
- OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp) — sabeel111, MIT (repository code); Apache-2.0 (Qwen/Qwen3.5-4B base and unsloth/Qwen3.5-4B-GGUF Q4_K_M conversion) — github.com/sabeel111/OpenSourceJev
- Qwen3-Reranker-4B — Qwen, Apache-2.0 — huggingface.co/Qwen/Qwen3-Reranker-4B
- Qwen3.5-0.8B Decision Model (Mourad Ghafiri) — Mourad Ghafiri, MIT code and training data; Apache-2.0 model weights — huggingface.co/mghafiri/qwen3.5-0.8B-decision-model
- Qwen3.5-9B Jev-like data-mix v2 — jsaurabh, Apache-2.0 (adapter code and weights; Qwen3.5-9B base is Apache-2.0) — huggingface.co/jsaurabh/qwen3.5-9b-jev-data-mix-v2
- Qwen3.8 27B (Chutes TEE) — Qwen / Chutes, open weights
- Raw Phi-4 mini direct logits — Microsoft / neutral reproduction, MIT — huggingface.co/microsoft/Phi-4-mini-instruct
- Raw Qwen3 0.6B direct logits — Alibaba Qwen / neutral reproduction, Apache-2.0 — huggingface.co/Qwen/Qwen3-0.6B
- Raw Qwen3 1.7B direct logits — Alibaba Qwen / neutral reproduction, Apache-2.0 — huggingface.co/Qwen/Qwen3-1.7B
- Raw Qwen3 4B Instruct 2507 direct logits — Alibaba Qwen / neutral reproduction, Apache-2.0 — huggingface.co/Qwen/Qwen3-4B-Instruct-2507
- Raw Qwen3 8B direct logits — Alibaba Qwen / neutral reproduction, Apache-2.0 — huggingface.co/Qwen/Qwen3-8B
- reflex 4B (kshetrajna12) — kshetrajna12, MIT (code, adapter); Apache-2.0 (base) — github.com/kshetrajna12/reflex
- reflex-27b (Qwen3.8-27B) — kshetrajna12, MIT code; Apache-2.0 Qwen weights — github.com/kshetrajna12/reflex
- SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ) — Theodore Lee (TheoLeeCJ), MIT (code); Qwen3.5 weights Apache-2.0 — github.com/TheoLeeCJ/openjev
- SimpleJev (Qwen3.5-0.8B, CPU) — sabeel111 / Featherless AI, Apache-2.0 server; Apache-2.0 Qwen3.5-0.8B checkpoint — github.com/featherless-ai/simple-jev
- SimpleJev Qwen3.6-35B-A3B — Featherless AI, Apache-2.0 (Qwen weights); repository licence not stated — github.com/featherless-ai/simple-jev
- SimpleJev Qwen3.8-27B — Featherless AI, Apache-2.0 (Qwen weights); repository licence not stated — github.com/featherless-ai/simple-jev
- smalljev semantic-v9 — Aditya (isHeSatoshi), Apache-2.0 — github.com/isHeSatoshi/smalljev
- spark-s1-4b-v6 (Open Spark Jev, abhishek085) — Abhishek Rai (abhishek085), Apache-2.0 (code and weights); base Qwen/Qwen3.5-4B Apache-2.0 — github.com/abhishek085/open-spark-jev
- swanOne — swanOne submitter, Patch/repository and Qwen/Mia component-specific terms recorded in RESULT.md — github.com/swanOne
- system-one (Qwen3-8B, Sean Goedecke) — Sean Goedecke, no licence file in the repository as of 19 Sep; Qwen3 weights Apache-2.0 — github.com/sgoedecke/system-one
- system-one-open (Gemma 4 E2B LoRA on an L4) — mithalouni, MIT (repository LICENSE; Gemma weights keep Google’s terms) — github.com/mithalouni/system-one-open
- Von (wfzyx, Option-Marker 395M) — wfzyx (Victor Hugo), Apache-2.0 code; Apache-2.0 weights (answerdotai/ModernBERT-large base) — github.com/wfzyx/von
- Winnow-12B Q8 — Eldan Ring, Apache-2.0, including the applicable Gemma 4 base/derivative licence terms — huggingface.co/EldanRing/Winnow-12B
- ZeroEntropy zerank-2 — ZeroEntropy, Apache-2.0 — huggingface.co/zeroentropy/zerank-2-reranker
Authors: if we tested the wrong configuration, tell us and we will rerun it. New entrants become a new version rather than silently changing this one.
JevBench v1.4.1 · additional views
Capability, cost and speed
Capability is the arithmetic mean of Intelligence and Calibration: (Intelligence + Calibration) / 2, on a 0–100 scale. Cost is USD per 1,000 decisions; its axis is logarithmic, and lower is better. Speed uses the JevBench Speed axis, where higher is faster. Estimated costs are marked.
3 placeholder rows have no published Intelligence or Calibration values and are omitted.
Top 20 by Capability
Capability with cost alongside
Each system has a wide Capability bar and a narrower cost bar. The cost scale is logarithmic: longer bars mean higher cost, so shorter is cheaper.
- 1GPT-6 Luna (medium)API95.4Cost $0.14 · I 97.4 · C 93.5
- 2DeepSeek V4.1 FlashAPI94.7Cost $0.59 · I 94.0 · C 95.5
- 3GPT-6 Luna (low)API93.9Cost $0.13 · I 95.8 · C 92.0
- 4GPT-5.6 LunaAPI90.3Cost $0.24 · I 93.1 · C 87.4
- 5djev79.7Cost $0.27 est. · I 71.6 · C 87.8
- 6Qwen3.8 27BAPI67.0Cost $2.67 est. · I 40.4 · C 93.6
- 7Jev 1.13.0API64.7Cost $0.040 · I 53.1 · C 76.3
- 8NInfer Qwen3.8-Flash-Next mixed64.1Cost $0.11 est. · I 49.5 · C 78.6
- 9NInfer Qwen3.8-27B NVFP463.7Cost $0.14 est. · I 51.5 · C 76.0
- 10Hopper63.5Cost $0.024 est. · I 48.0 · C 79.1
- 11SimpleJev Qwen3.8-27BAPI63.0Cost $0.10 est. · I 51.6 · C 74.5
- 12JevOne62.6Cost $0.14 est. · I 47.6 · C 77.6
- 13classifier.devAPI62.0Cost $0.0033 est. · I 51.6 · C 72.4
- 14reflex-27b61.8Cost $0.18 est. · I 46.4 · C 77.2
- 15JevK5 v0.2.061.7Cost $0.022 est. · I 48.9 · C 74.5
- 16LitJev61.4Cost $0.16 est. · I 46.3 · C 76.6
- 17openjev-sglangAPI59.3Cost $0.13 est. · I 49.4 · C 69.2
- 18NInfer Qwen3.8-27B NVFP459.3Cost $0.14 est. · I 51.5 · C 67.2
- 19jqv59.0Cost $0.056 est. · I 46.4 · C 71.6
- 20reflex 4B58.9Cost $0.022 est. · I 47.5 · C 70.4
- 21ZeroEntropy zerank-258.9Cost $0.047 · I 42.1 · C 75.8
- 22local-jev Qwen3.5-4B58.9Cost $0.030 est. · I 44.4 · C 73.3
- 23OpenJev58.1Cost $0.25 est. · I 58.1 · C 58.1
- 24JEV Qwen3.5-9B Base NVFP457.3Cost $0.077 est. · I 46.8 · C 67.7
- 25Gemini 3.1 Flash-LiteAPI56.9Cost $0.26 · I 54.5 · C 59.3
- 26Winnow-12B Q856.6Cost $0.037 est. · I 48.3 · C 64.8
- 27Decision 2B56.4Cost $0.018 est. · I 38.8 · C 74.1
- 28decider-35b-a3b56.2Cost $0.067 est. · I 47.2 · C 65.3
- 29metask-jev-4b55.8Cost $0.033 est. · I 44.7 · C 66.9
- 30SemIf55.6Cost $0.022 est. · I 44.4 · C 66.8
- 31Jev-Omni55.4Cost $0.037 est. · I 46.8 · C 64.1
- 32Von55.1Cost $0.0055 est. · I 34.5 · C 75.7
- 33Jobe Qwen3.5-4B55.1Cost $0.022 est. · I 44.1 · C 66.1
- 34Qwen3-Reranker-4B54.9Cost $0.050 · I 44.6 · C 65.2
- 35decision-machine-1API54.8Cost $0.035 · I 41.3 · C 68.3
- 36jev-local54.7Cost $0.077 est. · I 45.2 · C 64.2
- 37Qwen3.5-9B Jev-like data-mix v254.4Cost $0.083 est. · I 47.4 · C 61.3
- 38Open-Jev 9B53.0Cost $0.25 est. · I 44.2 · C 61.8
- 39SimpleJev Qwen3.6-35B-A3BAPI52.8Cost $0.12 est. · I 45.7 · C 59.8
- 40lev-350m52.7Cost $0.0063 est. · I 34.8 · C 70.6
- 41jeff52.3Cost $0.0060 est. · I 36.8 · C 67.9
- 42Bespoke Nimble 9B51.4Cost $0.17 est. · I 46.3 · C 56.4
- 43Decision Fast51.2Cost $0.0063 est. · I 37.1 · C 65.3
- 44djev51.2Cost $0.026 · I 47.0 · C 55.4
- 45OpenSourceJev51.1Cost $0.016 est. · I 41.8 · C 60.3
- 46openJev Verdict 1.450.7Cost $0.0039 est. · I 29.4 · C 72.0
- 47Raw Phi-4 mini direct logits50.3Cost $0.048 est. · I 41.8 · C 58.8
- 48OpenJev50.2Cost $0.066 est. · I 45.4 · C 55.0
- 49Laya49.9Cost $0.0029 est. · I 36.1 · C 63.7
- 50system-one-openAPI49.5Cost $0.015 est. · I 44.2 · C 54.9
- 51Open-Jev 2B48.8Cost $0.25 est. · I 42.3 · C 55.3
- 52open-alternative-jev48.6Cost $0.022 est. · I 38.6 · C 58.7
- 53Qwen3.5-0.8B Decision Model48.2Cost $0.0065 est. · I 28.1 · C 68.2
- 54spark-s1-4b-v646.3Cost $0.025 est. · I 45.1 · C 47.6
- 55open-jev-deberta-v3-large46.1Cost $0.0073 est. · I 25.6 · C 66.6
- 56Mixedbread mxbai-rerank-base-v245.5Cost $0.012 · I 6.8 · C 84.1
- 57BAAI bge-reranker-v2-m344.6Cost $0.0077 · I 5.0 · C 84.2
- 58OpenDecision44.4Cost $0.0066 est. · I 31.8 · C 57.1
- 59smalljev semantic-v942.4Cost $0.025 est. · I 25.7 · C 59.2
- 60kev 0.6B42.1Cost $0.0063 est. · I 34.2 · C 50.0
- 61Alibaba GTE Reranker ModernBERT-base41.9Cost $0.010 · I 4.8 · C 78.9
- 62Certo v141.5Cost $0.00097 est. · I 0.1 · C 83.0
- 63decider-2b41.0Cost $0.020 est. · I 38.5 · C 43.5
- 64kev 8B41.0Cost $0.073 est. · I 41.8 · C 40.2
- 65kev 4B40.9Cost $0.019 est. · I 42.1 · C 39.6
- 66GLiNER2.5 multi40.1Cost $0.0039 est. · I 23.1 · C 57.2
- 67kev 0.5B40.1Cost $0.0063 est. · I 30.5 · C 49.7
- 68openJev Verdict38.5Cost $0.0037 est. · I 30.0 · C 47.0
- 69system-one38.2Cost $0.089 est. · I 43.6 · C 32.8
- 70Raw Qwen3 4B Instruct 2507 direct logits37.7Cost $0.022 est. · I 46.4 · C 29.1
- 71GLiNER2.5 small35.6Cost $0.0039 est. · I 20.5 · C 50.7
- 72SimpleJev35.3Cost $0.011 est. · I 21.5 · C 49.1
- 73Raw Qwen3 8B direct logits34.9Cost $0.087 est. · I 45.7 · C 24.1
- 74Raw Qwen3 1.7B direct logits28.8Cost $0.015 est. · I 33.2 · C 24.3
- 75GLiNER2 large28.0Cost $0.0077 est. · I 31.1 · C 24.8
- 76GLiNER226.3Cost $0.0037 est. · I 27.4 · C 25.2
- 77Open Jev JSON Canvas24.1Cost $0.065 est. · I 48.2 · C 0.0
- 78Raw Qwen3 0.6B direct logits21.8Cost $0.0074 est. · I 22.8 · C 20.7
- 79Mirror19.8Cost $0.0077 est. · I 13.6 · C 26.0
- Jev (TypeSafe, closed)
- Jev rebuild
- Instruction model, JSON schema
- Service built on Jev
- Zero-shot classifier
- Closed decision API
- Reranker (neutral adapter)
- Raw-logit control (base model)
- Native-logit decision engine
- Cost per 1,000 decisions · log scale
Capability vs cost
The upper-left is the more attractive area: higher Capability and lower cost.
- Capability #1 · GPT-6 Luna (medium)
Capability 95.4 · Cost $0.14 / 1,000 · Speed 72.6 · JevBench #30 - Capability #2 · DeepSeek V4.1 Flash
Capability 94.7 · Cost $0.59 / 1,000 · Speed 71.6 · JevBench #71 - Capability #3 · GPT-6 Luna (low)
Capability 93.9 · Cost $0.13 / 1,000 · Speed 73.7 · JevBench #27 - Capability #4 · GPT-5.6 Luna
Capability 90.3 · Cost $0.24 / 1,000 · Speed 77.5 · JevBench #53 - Capability #5 · djev
Capability 79.7 · Cost $0.27 estimated / 1,000 · Speed 75.2 · JevBench #57
Capability vs speed
The upper-right is the more attractive area: higher Capability and higher Speed.
- Capability #1 · GPT-6 Luna (medium)
Capability 95.4 · Cost $0.14 / 1,000 · Speed 72.6 · JevBench #30 - Capability #2 · DeepSeek V4.1 Flash
Capability 94.7 · Cost $0.59 / 1,000 · Speed 71.6 · JevBench #71 - Capability #3 · GPT-6 Luna (low)
Capability 93.9 · Cost $0.13 / 1,000 · Speed 73.7 · JevBench #27 - Capability #4 · GPT-5.6 Luna
Capability 90.3 · Cost $0.24 / 1,000 · Speed 77.5 · JevBench #53 - Capability #5 · djev
Capability 79.7 · Cost $0.27 estimated / 1,000 · Speed 75.2 · JevBench #57
All three at once
The 3D view plots Capability vertically, lower cost to the right, and higher Speed toward you. Sphere size follows the JevBench score. Drag to rotate; pinch or scroll to zoom. The view loads when it scrolls into view.
Scroll here to load the interactive 3D view.
- Capability #1 · GPT-6 Luna (medium)
Capability 95.4 · Cost $0.14 / 1,000 decisions · Speed 72.6 · JevBench #30 - Capability #2 · DeepSeek V4.1 Flash
Capability 94.7 · Cost $0.59 / 1,000 decisions · Speed 71.6 · JevBench #71 - Capability #3 · GPT-6 Luna (low)
Capability 93.9 · Cost $0.13 / 1,000 decisions · Speed 73.7 · JevBench #27 - Capability #4 · GPT-5.6 Luna
Capability 90.3 · Cost $0.24 / 1,000 decisions · Speed 77.5 · JevBench #53 - Capability #5 · djev
Capability 79.7 · Cost $0.27 / 1,000 decisions · Speed 75.2 · JevBench #57
Vertical: Capability · Right: cheaper · Toward you: faster
The interactive 3D view loads when this panel scrolls into view.
79 systems plotted; systems missing cost or Speed are omitted. three.js r128 is included under its MIT license.
JevBench v1.4.1 public split · input capacity and long inputs
Context length
Context length is the amount of input a model or service can accept in one request. It matters when an app sends a long conversation state, policy set, or document: a smaller window can force truncation or chunking. A larger window is a capacity ceiling, not a promise that the system will use every token well.
Across 82 v1.4.1 rows (77 ranked systems and 5 unranked additions), supported published values range from 512 tokens to 1,050,000 tokens; 8 have no published maximum we could verify. For Jev-class rebuilds, the chart and table separate a model's trained length from any published serving cap and show training truncation limits where available.
Public accuracy by actual input length
Each line is one of the 13 top-15 systems with reconciled public outcomes and input-token counts. A point's tooltip shows the system, accuracy, correct answers and bucket size.
- Top five · #1 Jev 1.13.0 (TypeSafe AI)
- Top five · #3 Hopper
- Top five · #4 Winnow-12B Q8
- Top five · #5 reflex 4B (kshetrajna12)
- Top five · #6 djev (Maisa, diffusion-gemma)
- #7 Jev-Omni (akhilaaa3, Gemma-4-12B merged)
- #8 metask-jev-4b
- #9 SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)
- #10 Jobe Qwen3.5-4B (frozen)
- #11 local-jev Qwen3.5-4B
- #13 spark-s1-4b-v6 (Open Spark Jev, abhishek085)
- #14 jqv (Qwen3-32B zero-shot)
- #15 Qwen3-Reranker-4B
Exact correct counts and denominators by bucket
| System | Mean input tokens | Length n | <2k | 2–8k | 8–16k | 16–64k | 64–256k | 256k–1M | ≥1M |
|---|---|---|---|---|---|---|---|---|---|
| #1 Jev 1.13.0 (TypeSafe AI) | 1,057.8 | 183/231 | 125/146 · 85.6% | 27/37 · 73.0% | No data | No data | No data | No data | No data |
| #3 Hopper | 739.2 | 231/231 | 167/195 · 85.6% | 23/36 · 63.9% | No data | No data | No data | No data | No data |
| #4 Winnow-12B Q8 | 692.3 | 231/231 | 170/194 · 87.6% | 28/37 · 75.7% | No data | No data | No data | No data | No data |
| #5 reflex 4B (kshetrajna12) | 696.8 | 231/231 | 162/195 · 83.1% | 21/36 · 58.3% | No data | No data | No data | No data | No data |
| #6 djev (Maisa, diffusion-gemma) | 692.3 | 231/231 | 169/194 · 87.1% | 25/37 · 67.6% | No data | No data | No data | No data | No data |
| #7 Jev-Omni (akhilaaa3, Gemma-4-12B merged) | 686.5 | 231/231 | 175/194 · 90.2% | 30/37 · 81.1% | No data | No data | No data | No data | No data |
| #8 metask-jev-4b | 761.2 | 230/231 | 165/194 · 85.1% | 19/36 · 52.8% | No data | No data | No data | No data | No data |
| #9 SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ) | 698.6 | 231/231 | 166/195 · 85.1% | 21/36 · 58.3% | No data | No data | No data | No data | No data |
| #10 Jobe Qwen3.5-4B (frozen) | 698.6 | 231/231 | 166/195 · 85.1% | 21/36 · 58.3% | No data | No data | No data | No data | No data |
| #11 local-jev Qwen3.5-4B | 696.7 | 231/231 | 163/195 · 83.6% | 23/36 · 63.9% | No data | No data | No data | No data | No data |
| #13 spark-s1-4b-v6 (Open Spark Jev, abhishek085) | 794.6 | 231/231 | 161/194 · 83.0% | 22/37 · 59.5% | No data | No data | No data | No data | No data |
| #14 jqv (Qwen3-32B zero-shot) | 662.7 | 231/231 | 167/195 · 85.6% | 18/36 · 50.0% | No data | No data | No data | No data | No data |
| #15 Qwen3-Reranker-4B | 2,401.9 | 231/231 | 134/177 · 75.7% | 7/24 · 29.2% | 13/27 · 48.1% | 3/3 · 100.0% | No data | No data | No data |
Coverage: 13 of the top 15 systems are shown. Exclusions: JevK5 v0.2.0 (excluded: per-item accuracy 199/231 does not match published v1.4.1); system-one-open (Gemma 4 E2B LoRA on an L4) (excluded: no per-item token counts). All 13 included systems exactly reproduce their published public accuracy; stored lengths cover 183–231 decisions per system.
Long-policy tasks show a separate stress point
Across 19 public items in the long_policy family, several systems scored well below their full public-set accuracy. The comparison uses the family label, not only the token buckets.
- metask-jev-4b
Long policy 6/19 (31.6%) vs 184/231 overall (79.7%): −48.1 pp. - spark-s1-4b-v6 (Open Spark Jev, abhishek085)
Long policy 7/19 (36.8%) vs 183/231 overall (79.2%): −42.4 pp. - system-one-open (Gemma 4 E2B LoRA on an L4)
Long policy 6/19 (31.6%) vs 169/231 overall (73.2%): −41.6 pp. - Winnow-12B Q8
Long policy 15/19 (78.9%) vs 198/231 overall (85.7%): −6.8 pp. - Jev-Omni (akhilaaa3, Gemma-4-12B merged)
Long policy 15/19 (78.9%) vs 205/231 overall (88.7%): −9.8 pp.
Show long_policy results for all 14 matched systems
| System | Overall | Long policy (19 items) | Change |
|---|---|---|---|
| #8 metask-jev-4b | 184/231 · 79.7% | 6/19 · 31.6% | −48.1 pp |
| #13 spark-s1-4b-v6 (Open Spark Jev, abhishek085) | 183/231 · 79.2% | 7/19 · 36.8% | −42.4 pp |
| #12 system-one-open (Gemma 4 E2B LoRA on an L4) | 169/231 · 73.2% | 6/19 · 31.6% | −41.6 pp |
| #14 jqv (Qwen3-32B zero-shot) | 185/231 · 80.1% | 9/19 · 47.4% | −32.7 pp |
| #6 djev (Maisa, diffusion-gemma) | 194/231 · 84.0% | 10/19 · 52.6% | −31.4 pp |
| #9 SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ) | 187/231 · 81.0% | 10/19 · 52.6% | −28.3 pp |
| #10 Jobe Qwen3.5-4B (frozen) | 187/231 · 81.0% | 10/19 · 52.6% | −28.3 pp |
| #5 reflex 4B (kshetrajna12) | 183/231 · 79.2% | 10/19 · 52.6% | −26.6 pp |
| #3 Hopper | 190/231 · 82.3% | 11/19 · 57.9% | −24.4 pp |
| #1 Jev 1.13.0 (TypeSafe AI) | 200/231 · 86.6% | 12/19 · 63.2% | −23.4 pp |
| #11 local-jev Qwen3.5-4B | 186/231 · 80.5% | 12/19 · 63.2% | −17.4 pp |
| #15 Qwen3-Reranker-4B | 157/231 · 68.0% | 10/19 · 52.6% | −15.3 pp |
| #7 Jev-Omni (akhilaaa3, Gemma-4-12B merged) | 205/231 · 88.7% | 15/19 · 78.9% | −9.8 pp |
| #4 Winnow-12B Q8 | 198/231 · 85.7% | 15/19 · 78.9% | −6.8 pp |
For example, metask-jev-4b scored 31.6% on long_policy versus 79.7% overall (change −48.1 pp), while Winnow-12B Q8 scored 78.9% versus 85.7% (change −6.8 pp). These are descriptive public-set comparisons. Prompt wrappers and tokenizers differ by system, and the 19-item family is small, so the results do not isolate context length as the cause. Only public item results were used; sealed-set item rows were not used.
Published context limits · logarithmic scale
Each row shows the system's exact published maximum input context. API/serving caps, hard limits and trained lengths use different bar colors; training configuration limits appear as separate markers and values where published.
- #1 Jev 1.13.0 (TypeSafe AI)64,000 · API cap
- #2 JevK5 v0.2.0262,144 · Trained lengthTraining configuration: max_seq_len 2,048 tokens.
- #3 Hopper262,144 · Trained length
- #4 Winnow-12B Q865,536 · API cap
- #5 reflex 4B (kshetrajna12)262,144 · Trained length
- #6 djev (Maisa, diffusion-gemma)32,768 · API cap
- #7 Jev-Omni (akhilaaa3, Gemma-4-12B merged)262,144 · Trained length
- #8 metask-jev-4b262,144 · Trained length
- #9 SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)262,144 · Trained length
- #10 Jobe Qwen3.5-4B (frozen)262,144 · Trained lengthTraining configuration: max_seq_len 4,096 tokens.
- #11 local-jev Qwen3.5-4B262,144 · Trained length
- #12 system-one-open (Gemma 4 E2B LoRA on an L4)131,072 · Trained lengthTraining configuration: state limit 2,048 tokens.
- #13 spark-s1-4b-v6 (Open Spark Jev, abhishek085)262,144 · Trained lengthTraining configuration: max_seq_len 2,048 tokens.
- #14 jqv (Qwen3-32B zero-shot)32,768 · Trained length
- #15 Qwen3-Reranker-4B32,768 · Trained length
- #16 decider-35b-a3b (Mapika)262,144 · Hard limit
- #17 Raw Qwen3 4B Instruct 2507 direct logits262,144 · Trained length
- #18 OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp)32,768 · Trained length
- #19 ZeroEntropy zerank-232,768 · Trained length
- #20 decision-machine-1 (milliseconds.ai)Unknown
- #21 Raw Phi-4 mini direct logits131,072 · Trained length
- #22 JEV Qwen3.5-9B Base NVFP4262,144 · Trained length
- #23 OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)65,536 · API cap
- #24 kev 4B (research preview)32,768 · Trained length
- #25 Decision 2B (FlyMy.AI, v59)131,072 · Trained length
- #26 Qwen3.5-9B Jev-like data-mix v2262,144 · Trained length
- #27 GPT-6 Luna (low reasoning effort)1,050,000 · API cap
- #28 SimpleJev Qwen3.8-27B2,000 · API cap
- #29 NInfer Qwen3.8-Flash-Next mixed262,144 · Trained length
- #30 GPT-6 Luna (default medium reasoning effort)1,050,000 · API cap
- #31 open-alternative-jev (Qwen3.5-4B, IkerMoel)262,144 · Trained length
- #32 jev-local (Qwen3.5-9B)262,144 · Trained length
- #33 Decision Fast (FlyMy.AI, v53a)32,768 · Trained length
- #34 decider-2b (Mapika)262,144 · Hard limit
- #35 jeff (Logan Markewich, GLiFormer 400M)8,192 · Hard limit
- #36 Laya (Convai Innovations, ModernBERT-large 421M)512 · Hard limit
- #37 lev-350m (Franck Verrot, LFM2.5-350M)32,768 · Trained length
- #38 openjev-sglang (Qwen3.6-35B-A3B on SGLang)32,768 · API cap
- #39 Von (wfzyx, Option-Marker 395M)8,192 · Trained length
- #40 NInfer Qwen3.8-27B NVFP4 (T=1.5)262,144 · Trained length
- #41 NInfer Qwen3.8-27B NVFP4262,144 · Trained length
- #42 kev 8B (research preview)32,768 · Trained length
- #43 JevOne262,144 · Trained length
- #44 SimpleJev Qwen3.6-35B-A3B262,144 · Trained length
- #45 kev 0.6B (research preview)32,768 · Trained length
- #46 Raw Qwen3 8B direct logits32,768 · Trained length
- #47 system-one (Qwen3-8B, Sean Goedecke)32,768 · Trained length
- #48 OpenDecision (ModernBERT-large zero-shot)8,192 · Hard limit
- #49 LitJev (Qwen3.8-27B)262,144 · Trained length
- #50 openJev Verdict 1.4512 · API cap
- #51 kev 0.5B8,192 · API cap
- #52 Bespoke Nimble 9B (Bespoke Labs)8,192 · API capTraining configuration: max_seq_len 2,048 tokens.
- #53 GPT-5.6 Luna (low reasoning effort)1,050,000 · API cap
- #54 openJev Verdict (heman10x, ModernBERT-base 151M)8,192 · Hard limit
- #55 Raw Qwen3 1.7B direct logits32,768 · Trained length
- #56 reflex-27b (Qwen3.8-27B)262,144 · Trained length
- #57 djev (thinking)262,144 · Trained length
- #58 GLiNER2 large (Fastino)Unknown
- #59 OpenJev (thinking, BF16)65,536 · API cap
- #60 Qwen3.5-0.8B Decision Model (Mourad Ghafiri)262,144 · Hard limit
- #61 Gemini 3.1 Flash-Lite1,048,576 · API cap
- #62 open-jev-deberta-v3-large (local CPU)512 · Hard limit
- #63 smalljev semantic-v9131,072 · Trained length
- #64 GLiNER2 (Fastino, gliner2.5-base)Unknown
- #65 Open-Jev 9B (Zefan Cai)262,144 · Trained length
- #66 Open-Jev 2B (Zefan Cai)262,144 · Trained length
- #67 GLiNER2.5 multi (Fastino, 287M)Unknown
- #68 SimpleJev (Qwen3.5-0.8B, CPU)262,144 · Trained length
- #69 GLiNER2.5 small (Fastino, 74M)Unknown
- #70 Raw Qwen3 0.6B direct logits32,768 · Trained length
- #71 DeepSeek V4.1 Flash (thinking default)1,000,000 · API cap
- #72 Mirror512 · Hard limit
- #73 Mixedbread mxbai-rerank-base-v232,768 · Hard limit
- #74 BAAI bge-reranker-v2-m38,192 · Hard limit
- #75 Alibaba GTE Reranker ModernBERT-base8,192 · Hard limit
- #76 Certo v1 (AltSlate Labs)8,192 · Hard limit
- #77 Open Jev JSON Canvas (JoshuaSP)262,144 · Trained length
- Unranked classifier.dev (fast tier)Unknown
- Unranked Needle 3 (Cactus, 2-bit, local CPU)Unknown
- Unranked Needle 3, options as tools (post-hoc adapter mode)Unknown
- Unranked Qwen3.8 27B (Chutes TEE)262,144 · API cap
- Unranked swanOne262,144 · Trained length
- API / serving cap
- Hard limit
- Trained length
- Training max_seq_len
- Training state limit
Context limits by system
82 systems · sources checked 24 Sept 2026
Sort by selecting a column heading. Source links open the primary model card, vendor documentation, or API documentation.
| #1Jev 1.13.0 (TypeSafe AI)Basis, training and serving notes Evidence: TypeSafe's official Jev 1.13 docs specify a 64K request budget and a 32K state-plus-longest-question budget. | 64,000 total request; 32,000 state + longest question | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | API cap |
|---|---|---|---|
| #2JevK5 v0.2.0Basis, training and serving notes Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. Training max_seq_len: 2,048 tokens. Training utility defaults --max-len to 2,048 and skips longer rows. Runtime has no smaller total context cap documented; native Qwen3.5-4B window is 262,144. | 262,144 | Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026 | Trained length |
| #3HopperBasis, training and serving notes Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. JevBench registry identifies this as a LoRA on Qwen3.5-4B. Its public model card does not specify a shorter max_seq_len or serving truncation, so the base model's native 262,144 window is listed; adapter training length is unknown. | 262,144 | Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026 | Trained length |
| #4Winnow-12B Q8Basis, training and serving notes Base model: google/gemma-4-12B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: The Winnow model card documents a 65,536-position configured Q8 test profile; Gemma 4 12B base supports 262,144. | 65,536 configured/tested; base 262,144 | Open primary sourceSource date: 21 Sept 2026 · checked: 23 Sept 2026 | API cap |
| #5reflex 4B (kshetrajna12)Basis, training and serving notes Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. Server refuses an input above the base model's context window rather than truncating it. The repo's 8,192 max-pack-tokens is a batching budget, not the per-request context cap. | 262,144 | Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026 | Trained length |
| #6djev (Maisa, diffusion-gemma)Basis, training and serving notes Base model: google/diffusiongemma-26B-A4B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Runtime setting is bounded to 1,024–32,768; DiffusionGemma base is 262,144. | 32,768 (prompt + reserved canvas) | Open primary sourceSource date: 19 Sept 2026 · checked: 23 Sept 2026 | API cap |
| #7Jev-Omni (akhilaaa3, Gemma-4-12B merged)Basis, training and serving notes Base model: google/gemma-4-12B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Gemma 4 12B card states a 256K context window. | 262,144 | Open primary sourceSource date: 20 Jul 2026 · checked: 23 Sept 2026 | Trained length |
| #8metask-jev-4bBasis, training and serving notes Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. Reported validation point: 4,096 tokens. This is not automatically the maximum accepted input. Model card reports 4,096 as the validated evaluation point, not an architectural limit; the same card/config says native 262,144. Published validated evaluation point: 4,096 tokens. | 262,144 | Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026 | Trained length |
| #9SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)Basis, training and serving notes Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. | 262,144 | Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026 | Trained length |
| #10Jobe Qwen3.5-4B (frozen)Basis, training and serving notes Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. Training max_seq_len: 4,096 tokens. The optional adapter-training helper uses max_tokens=4,096 and rejects longer training examples. The ranked v1.4.1 entry is the frozen Qwen3.5-4B backbone, so this optional training helper does not set its inference window. | 262,144 | Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026 | Trained length |
| #11local-jev Qwen3.5-4BBasis, training and serving notes Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. The README documents long-state shortening. Its 32,768-token context_tokens value is only an illustrative custom-model card; the effective Qwen3.5 deployment cap is not published. The listed 262,144 is the base model window. | 262,144 | Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026 | Trained length |
| #12system-one-open (Gemma 4 E2B LoRA on an L4)Basis, training and serving notes Base model: google/gemma-4-E2B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Gemma 4 E2B card/config state 131,072 (128K) positions. Training state limit: 2,048 tokens. Training batch token budget: 24,576 tokens. Full training profile caps state at 2,048 tokens and total batch tokens at 24,576; this is a training profile, not an inference limit. | 131,072 | Open primary sourceSource date: 20 Jul 2026 · checked: 23 Sept 2026 | Trained length |
| #13spark-s1-4b-v6 (Open Spark Jev, abhishek085)Basis, training and serving notes Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. Training max_seq_len: 2,048 tokens. Published RLCD configs use max_len=2,048 for training; Qwen3.5-4B native inference window is 262,144. | 262,144 | Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026 | Trained length |
| #14jqv (Qwen3-32B zero-shot)Basis, training and serving notes Base model: Qwen/Qwen3-32B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3 card specifies 32,768 native. The config's 40,960 positions reserve output space; 131,072 requires YaRN. | 32,768 | Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026 | Trained length |
| #15Qwen3-Reranker-4BBasis, training and serving notes Base model: Qwen/Qwen3-Reranker-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Official reranker card states 32K context; config has extra positions reserved for prompt/output. Official card's stated context is 32K; do not substitute the larger config allocation because the card is explicit. | 32,768 | Open primary sourceSource date: 16 Apr 2026 · checked: 23 Sept 2026 | Trained length |
| #16decider-35b-a3b (Mapika)Basis, training and serving notes Base model: Qwen/Qwen3.5-35B-A3B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: The submitted model's config.json sets 262,144 positions; it is a decision readout on Qwen3.5-35B-A3B-Base, whose native context is also 262,144. | 262,144 | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | Hard limit |
| #17Raw Qwen3 4B Instruct 2507 direct logitsBasis, training and serving notes Base model: Qwen/Qwen3-4B-Instruct-2507. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Official model card states 262,144 natively. | 262,144 | Open primary sourceSource date: 17 Sept 2025 · checked: 23 Sept 2026 | Trained length |
| #18OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp)Basis, training and serving notes Base model: Qwen/Qwen3-1.7B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3-1.7B card specifies 32,768 context. | 32,768 | Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026 | Trained length |
| #19ZeroEntropy zerank-2Basis, training and serving notes Base model: zeroentropy/zerank-2-reranker. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Official model card states 32,768 context. | 32,768 | Open primary sourceSource date: 24 Jul 2026 · checked: 23 Sept 2026 | Trained length |
| #20decision-machine-1 (milliseconds.ai)Basis, training and serving notes Evidence: No primary public model-card or API context limit was found. | Unknown | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | Unknown |
| #21Raw Phi-4 mini direct logitsBasis, training and serving notes Base model: microsoft/Phi-4-mini-instruct. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Microsoft card states 128K context; config uses long-RoPE scaling. | 131,072 | Open primary sourceSource date: 10 Dec 2025 · checked: 23 Sept 2026 | Trained length |
| #22JEV Qwen3.5-9B Base NVFP4Basis, training and serving notes Base model: Qwen/Qwen3.5-9B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.5-9B: native context length 262,144; optional YaRN extension is not enabled by default. | 262,144 | Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026 | Trained length |
| #23OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)Basis, training and serving notes Base model: google/diffusiongemma-26B-A4B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: OPENJEV_MAX_MODEL_LEN defaults to 65,536 in the submitted runner; DiffusionGemma base supports 262,144. | 65,536; base 262,144 | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | API cap |
| #24kev 4B (research preview)Basis, training and serving notes Base model: Qwen/Qwen3-4B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3-4B-Base is in the Qwen3 family with a 32,768-token native window. | 32,768 | Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026 | Trained length |
| #25Decision 2B (FlyMy.AI, v59)Basis, training and serving notes Base model: openbmb/MiniCPM5-2B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: MiniCPM5-2B card/config states 131,072 context. | 131,072 | Open primary sourceSource date: 12 Sept 2026 · checked: 23 Sept 2026 | Trained length |
| #26Qwen3.5-9B Jev-like data-mix v2Basis, training and serving notes Base model: Qwen/Qwen3.5-9B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.5-9B: native context length 262,144; optional YaRN extension is not enabled by default. | 262,144 | Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026 | Trained length |
| #27GPT-6 Luna (low reasoning effort)Basis, training and serving notes Evidence: OpenAI model docs: 1,050,000 context window and 128,000 maximum output. | 1,050,000 context window; 128,000 max output | Open primary sourceSource date: 22 Sept 2026 · checked: 23 Sept 2026 | API cap |
| #28SimpleJev Qwen3.8-27BBasis, training and serving notes Base model: Qwen/Qwen3.8-27B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: The tested public Simple Jev demo API documents a 2,000-token context limit; the Qwen3.8-27B base window is 262,144. | 2,000 demo API context; base 262,144 | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | API cap |
| #29NInfer Qwen3.8-Flash-Next mixedBasis, training and serving notes Base model: Qwen/Qwen3.8-Flash-Next. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.8-Flash-Next card: 262,144 native; the 1M extension requires YaRN. | 262,144 | Open primary sourceSource date: 27 Aug 2026 · checked: 23 Sept 2026 | Trained length |
| #30GPT-6 Luna (default medium reasoning effort)Basis, training and serving notes Evidence: OpenAI model docs: 1,050,000 context window and 128,000 maximum output. | 1,050,000 context window; 128,000 max output | Open primary sourceSource date: 22 Sept 2026 · checked: 23 Sept 2026 | API cap |
| #31open-alternative-jev (Qwen3.5-4B, IkerMoel)Basis, training and serving notes Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default. | 262,144 | Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026 | Trained length |
| #32jev-local (Qwen3.5-9B)Basis, training and serving notes Base model: Qwen/Qwen3.5-9B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.5-9B: native context length 262,144; optional YaRN extension is not enabled by default. | 262,144 | Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026 | Trained length |
| #33Decision Fast (FlyMy.AI, v53a)Basis, training and serving notes Base model: Qwen/Qwen3-0.6B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3-0.6B-Base config specifies 32,768 positions. | 32,768 | Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026 | Trained length |
| #34decider-2b (Mapika)Basis, training and serving notes Base model: Qwen/Qwen3.5-2B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: The submitted model's config.json sets 262,144 positions; it is a decision readout on Qwen3.5-2B-Base, whose native context is also 262,144. | 262,144 | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | Hard limit |
| #35jeff (Logan Markewich, GLiFormer 400M)Basis, training and serving notes Base model: knowledgator/gliformer-large-v1. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: GLiFormer card states configured max_len=8,192. | 8,192 | Open primary sourceSource date: 18 Sept 2026 · checked: 23 Sept 2026 | Hard limit |
| #36Laya (Convai Innovations, ModernBERT-large 421M)Basis, training and serving notes Base model: convaiinnovations/laya. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Laya model card documents a 512-token base context; head_max_len=192 is its answer-candidate budget. | 512 | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | Hard limit |
| #37lev-350m (Franck Verrot, LFM2.5-350M)Basis, training and serving notes Base model: LiquidAI/LFM2.5-350M. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: LiquidAI model card states 32,768 context; config has a larger positional allocation. LiquidAI card states a 32,768 context length. Config has a larger positional allocation; no larger trained/evaluated sequence is claimed. | 32,768 | Open primary sourceSource date: 5 Aug 2026 · checked: 23 Sept 2026 | Trained length |
| #38openjev-sglang (Qwen3.6-35B-A3B on SGLang)Basis, training and serving notes Base model: Qwen/Qwen3.6-35B-A3B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Runtime defaults max_input_tokens=32,768 and max_total_input_tokens=262,144. | 32,768 per question; 262,144 total across questions | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | API cap |
| #39Von (wfzyx, Option-Marker 395M)Basis, training and serving notes Evidence: The Von model card states an 8,192-token context for its ModernBERT-large scoring model and describes accurate premise reading to about 2,048 tokens. | 8,192 model context; reads well to about 2,048 | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | Trained length |
| #40NInfer Qwen3.8-27B NVFP4 (T=1.5)Basis, training and serving notes Base model: Qwen/Qwen3.8-27B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.8-27B card: 262,144 native; one-million-token extension requires YaRN. | 262,144 | Open primary sourceSource date: 14 Aug 2026 · checked: 23 Sept 2026 | Trained length |
| #41NInfer Qwen3.8-27B NVFP4Basis, training and serving notes Base model: Qwen/Qwen3.8-27B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.8-27B card: 262,144 native; one-million-token extension requires YaRN. | 262,144 | Open primary sourceSource date: 14 Aug 2026 · checked: 23 Sept 2026 | Trained length |
| #42kev 8B (research preview)Basis, training and serving notes Base model: Qwen/Qwen3-8B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3 card specifies 32,768 native; 131,072 requires YaRN. | 32,768 | Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026 | Trained length |
| #43JevOneBasis, training and serving notes Base model: Qwen/Qwen3.6-35B-A3B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.6-35B-A3B card: 262,144 native; optional YaRN extension is not assumed. JevOne is published as Qwen3.6-35B-A3B BF16 with a bidirectional option-logit mapping; no smaller serving or training sequence cap is documented. | 262,144 | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | Trained length |
| #44SimpleJev Qwen3.6-35B-A3BBasis, training and serving notes Base model: Qwen/Qwen3.6-35B-A3B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.6-35B-A3B card: 262,144 native; optional YaRN extension is not assumed. | 262,144 | Open primary sourceSource date: 24 Apr 2026 · checked: 23 Sept 2026 | Trained length |
| #45kev 0.6B (research preview)Basis, training and serving notes Base model: Qwen/Qwen3-0.6B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3-0.6B-Base config specifies 32,768 positions. | 32,768 | Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026 | Trained length |
| #46Raw Qwen3 8B direct logitsBasis, training and serving notes Base model: Qwen/Qwen3-8B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3 card specifies 32,768 native; 131,072 requires YaRN. | 32,768 | Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026 | Trained length |
| #47system-one (Qwen3-8B, Sean Goedecke)Basis, training and serving notes Base model: Qwen/Qwen3-8B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3 card specifies 32,768 native; 131,072 requires YaRN. | 32,768 | Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026 | Trained length |
| #48OpenDecision (ModernBERT-large zero-shot)Basis, training and serving notes Base model: MoritzLaurer/ModernBERT-large-zeroshot-v2.0. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Config has 8,192 positions. Card says v2.0 may not fully use the 8K window; exact trained length is not stated. ModernBERT config allows 8,192 positions. The v2.0 card says the older zero-shot checkpoint may not fully use the long window; it does not state a smaller exact trained limit. | 8,192 | Open primary sourceSource date: 16 Jan 2025 · checked: 23 Sept 2026 | Hard limit |
| #49LitJev (Qwen3.8-27B)Basis, training and serving notes Base model: Qwen/Qwen3.8-27B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.8-27B card: 262,144 native; one-million-token extension requires YaRN. | 262,144 | Open primary sourceSource date: 14 Aug 2026 · checked: 23 Sept 2026 | Trained length |
| #50openJev Verdict 1.4Basis, training and serving notes Base model: heman10x/rlcd-modernbert-151m. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: The fixed v1.4 inference path sets its context budget to 512; underlying model config is 8,192. | 512 service budget; model config 8,192 | Open primary sourceSource date: 20 Sept 2026 · checked: 23 Sept 2026 | API cap |
| #51kev 0.5BBasis, training and serving notes Base model: Qwen/Qwen2.5-0.5B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Kev 0.5B model card specifies an 8,192-token serving allowance per branch and a 32K backbone window. | 8,192 per branch; backbone 32,768 | Open primary sourceSource date: 20 Sept 2026 · checked: 23 Sept 2026 | API cap |
| #52Bespoke Nimble 9B (Bespoke Labs)Basis, training and serving notes Base model: Qwen/Qwen3.5-9B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Serving code defaults NIMBLE_MAX_PROMPT_TOKENS to 8,192 and launches the backend with that request cap. Training max_seq_len: 2,048 tokens. The published adapter-training recipe uses --max-length=2,048; the ranked serving profile has an 8,192-token prompt cap. System repository · 23 Sept 2026 | 8,192 prompt cap; base 262,144 | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | API cap |
| #53GPT-5.6 Luna (low reasoning effort)Basis, training and serving notes Evidence: OpenAI model docs: 1,050,000 context window and 128,000 maximum output. | 1,050,000 context window; 128,000 max output | Open primary sourceSource date: 9 Jul 2026 · checked: 23 Sept 2026 | API cap |
| #54openJev Verdict (heman10x, ModernBERT-base 151M)Basis, training and serving notes Base model: knowledgator/gliclass-modern-base-v2.0. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Model tokenizer config sets an 8,192-token maximum. Model config/tokenizer sets 8,192; its current separate v1.4 inference-engine cap is not documented in the available primary sources. | 8,192 | Open primary sourceSource date: 12 Aug 2025 · checked: 23 Sept 2026 | Hard limit |
| #55Raw Qwen3 1.7B direct logitsBasis, training and serving notes Base model: Qwen/Qwen3-1.7B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3-1.7B card specifies 32,768 context. | 32,768 | Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026 | Trained length |
| #56reflex-27b (Qwen3.8-27B)Basis, training and serving notes Base model: Qwen/Qwen3.8-27B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.8-27B card: 262,144 native; one-million-token extension requires YaRN. | 262,144 | Open primary sourceSource date: 14 Aug 2026 · checked: 23 Sept 2026 | Trained length |
| #57djev (thinking)Basis, training and serving notes Base model: google/diffusiongemma-26B-A4B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: DiffusionGemma card/config states 256K context. The standard djev runtime caps requests at 32,768, but the separately measured full-generation thinking run has no matching runtime cap published. The listed 262,144 is the DiffusionGemma base window, not a verified cap for this run. | 262,144 | Open primary sourceSource date: 15 Jul 2026 · checked: 23 Sept 2026 | Trained length |
| #58GLiNER2 large (Fastino)Basis, training and serving notes Base model: microsoft/deberta-v3-large. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Official docs provide chunked long-document helpers but no maximum total document size; the named DeBERTa-v3-large encoder has 512 positions. | Unknown | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | Unknown |
| #59OpenJev (thinking, BF16)Basis, training and serving notes Base model: google/diffusiongemma-26B-A4B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: OpenJev runtime defaults OPENJEV_MAX_MODEL_LEN to 65,536; the 512-token thinking allowance is generated output, not input context. | 65,536; base 262,144 | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | API cap |
| #60Qwen3.5-0.8B Decision Model (Mourad Ghafiri)Basis, training and serving notes Base model: Qwen/Qwen3.5-0.8B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: The submitted decision model's config.json sets 262,144 positions; it is based on Qwen3.5-0.8B-Base, whose native context is 262,144. | 262,144 | Open primary sourceSource date: 22 Sept 2026 · checked: 23 Sept 2026 | Hard limit |
| #61Gemini 3.1 Flash-LiteBasis, training and serving notes Evidence: Google model docs explicitly list a 1,048,576 input-token limit and 65,536 output-token limit. | 1,048,576 input token limit | Open primary sourceSource date: 21 Jul 2026 · checked: 23 Sept 2026 | API cap |
| #62open-jev-deberta-v3-large (local CPU)Basis, training and serving notes Base model: microsoft/deberta-v3-large. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Microsoft config max_position_embeddings=512. | 512 | Open primary sourceSource date: 19 Mar 2023 · checked: 23 Sept 2026 | Hard limit |
| #63smalljev semantic-v9Basis, training and serving notes Base model: openbmb/MiniCPM5-2B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: MiniCPM5-2B card/config states 131,072 context. | 131,072 | Open primary sourceSource date: 12 Sept 2026 · checked: 23 Sept 2026 | Trained length |
| #64GLiNER2 (Fastino, gliner2.5-base)Basis, training and serving notes Base model: microsoft/deberta-v3-large. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Official docs provide chunked long-document helpers but no maximum total document size; underlying DeBERTa-v3-base encoder has 512 positions. | Unknown | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | Unknown |
| #65Open-Jev 9B (Zefan Cai)Basis, training and serving notes Base model: Qwen/Qwen3.5-9B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.5-9B: native context length 262,144; optional YaRN extension is not enabled by default. | 262,144 | Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026 | Trained length |
| #66Open-Jev 2B (Zefan Cai)Basis, training and serving notes Base model: Qwen/Qwen3.5-2B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Mapika model config and Qwen3.5 family card give 262,144 positions. | 262,144 | Open primary sourceSource date: 23 Apr 2026 · checked: 23 Sept 2026 | Trained length |
| #67GLiNER2.5 multi (Fastino, 287M)Basis, training and serving notes Evidence: Model docs say max_len truncates and long-context helpers chunk documents; no fixed total input ceiling is published. | Unknown | Open primary sourceSource date: 20 Sept 2026 · checked: 23 Sept 2026 | Unknown |
| #68SimpleJev (Qwen3.5-0.8B, CPU)Basis, training and serving notes Base model: Qwen/Qwen3.5-0.8B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Official Qwen3.5-0.8B card/config reports 262,144 natively; optional YaRN extends the window. | 262,144 | Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026 | Trained length |
| #69GLiNER2.5 small (Fastino, 74M)Basis, training and serving notes Evidence: Model docs say max_len truncates and long-context helpers chunk documents; no fixed total input ceiling is published. | Unknown | Open primary sourceSource date: 20 Sept 2026 · checked: 23 Sept 2026 | Unknown |
| #70Raw Qwen3 0.6B direct logitsBasis, training and serving notes Base model: Qwen/Qwen3-0.6B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3-0.6B card specifies 32,768 context. | 32,768 | Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026 | Trained length |
| #71DeepSeek V4.1 Flash (thinking default)Basis, training and serving notes Evidence: DeepSeek API model/pricing docs list DeepSeek V4.1 Flash at 1M context and 384K maximum output. | 1,000,000 context; 384,000 max output | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | API cap |
| #72MirrorBasis, training and serving notes Base model: microsoft/deberta-v3-large. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Microsoft config max_position_embeddings=512. | 512 | Open primary sourceSource date: 19 Mar 2023 · checked: 23 Sept 2026 | Hard limit |
| #73Mixedbread mxbai-rerank-base-v2Basis, training and serving notes Base model: mixedbread-ai/mxbai-rerank-base-v2. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Official config max_position_embeddings=32,768. | 32,768 | Open primary sourceSource date: 8 Apr 2026 · checked: 23 Sept 2026 | Hard limit |
| #74BAAI bge-reranker-v2-m3Basis, training and serving notes Base model: BAAI/bge-reranker-v2-m3. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Official tokenizer limit is 8,192; config has 8,194 positions. | 8,192 | Open primary sourceSource date: 24 Jun 2024 · checked: 23 Sept 2026 | Hard limit |
| #75Alibaba GTE Reranker ModernBERT-baseBasis, training and serving notes Base model: Alibaba-NLP/gte-reranker-modernbert-base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Official config max_position_embeddings=8,192. | 8,192 | Open primary sourceSource date: 4 Jul 2025 · checked: 23 Sept 2026 | Hard limit |
| #76Certo v1 (AltSlate Labs)Basis, training and serving notes Base model: altslate/certo-decision-model. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Submitted model config/tokenizer sets 8,192. | 8,192 | Open primary sourceSource date: 21 Sept 2026 · checked: 23 Sept 2026 | Hard limit |
| #77Open Jev JSON Canvas (JoshuaSP)Basis, training and serving notes Base model: google/diffusiongemma-26B-A4B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: DiffusionGemma card/config states 256K context. | 262,144 | Open primary sourceSource date: 15 Jul 2026 · checked: 23 Sept 2026 | Trained length |
| Unrankedclassifier.dev (fast tier)Basis, training and serving notes Evidence: The Jev model behind the service has a documented limit, but no primary source documents a narrower or matching classifier.dev endpoint cap. | Unknown | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | Unknown |
| UnrankedNeedle 3 (Cactus, 2-bit, local CPU)Basis, training and serving notes Evidence: Official Needle 3 docs describe text input but publish no maximum context/token limit. | Unknown | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | Unknown |
| UnrankedNeedle 3, options as tools (post-hoc adapter mode)Basis, training and serving notes Evidence: Official Needle 3 docs describe tool inputs but publish no maximum context/token limit. | Unknown | Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026 | Unknown |
| UnrankedQwen3.8 27B (Chutes TEE)Basis, training and serving notes Evidence: Chutes model catalog lists Qwen3.8-27B-TEE at 262K context. | 262,144 context | Open primary sourceSource date: 17 Aug 2026 · checked: 23 Sept 2026 | API cap |
| UnrankedswanOneBasis, training and serving notes Base model: Qwen/Qwen3.8-Flash-Next. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input. Evidence: Qwen3.8-Flash-Next card: 262,144 native; the 1M extension requires YaRN. The submitted NVFP4 runner is based on Qwen3.8-Flash-Next; 262,144 is its native window. No larger runtime setting is documented for the submitted patch. | 262,144 | Open primary sourceSource date: 27 Aug 2026 · checked: 23 Sept 2026 | Trained length |
“Hard limit” is an explicit model or tokenizer ceiling; “Trained length” is a published base-model or training length; “API cap” is a published service limit. These are different kinds of evidence. A base-model window does not prove that a particular hosted endpoint accepts the same length; row notes identify cases where its serving cap is unpublished. Some API docs publish a combined context window and a separate output ceiling, so the usable input can be lower when output tokens share that window. “Unknown” means no supported maximum was found.
The input-length chart uses each system's existing usage.input_tokens telemetry and public item outcomes; no new model runs were made. Source dates are listed beside each primary-source link, and every source was checked 24 Sept 2026. Training limits and inference or API caps are shown separately where published.
Historical v1.3.0 board: weightings, per-task grid, topic radars and held-out diagnostics
The following public-only tables and diagnostics preserve the earlier JevBench v1.3.0 view. The ranking above is the current v1.4.1 result.
JevBench v1.3.0 · 534 decisions per system
JevBench Score (Intelligence, Calibration, Speed, Cost — 25 % each)
OfficialIntelligence above chance, Calibration, Speed, Cost — 25 % each, geometric mean; below 50 Intelligence receives a growing near-chance penalty. Change the weighting ↓
- 1Jev 1.13.074.4I 86 · C 83 · S 83 · K 52 · $0.040
- 2SemIf (Qwen3.5-4B)73.1I 79 · C 73 · S 84 · K 59 · ~$0.022 est.
- 3djev (Maisa, diffusion-gemma)†73.0I 83 · C 65 · S 91 · K 58 · $0.026 ann.
- 4Winnow-12B Q8†71.2I 82 · C 72 · S 82 · K 53 · ~$0.037 est.
- 5reflex 4B†70.3I 80 · C 75 · S 68 · K 60 · ~$0.022 est.
- 6jqv†68.6I 79 · C 79 · S 75 · K 47 · ~$0.056 est.
- 7decision-machine-1†68.3I 62 · C 70 · S 93 · K 54 · $0.035
- 8decider-35b-a3b†67.6I 80 · C 72 · S 81 · K 45 · ~$0.067 est.
- 9open-alternative-jev (Qwen3.5-4B)†67.0I 64 · C 63 · S 83 · K 60 · ~$0.022 est.
- 10system-one-open66.6I 70 · C 57 · S 77 · K 65 · ~$0.015 est.
- 11OpenJev (razorback16)66.4I 79 · C 65 · S 83 · K 45 · ~$0.066 est.
- 12SimpleJev Qwen3.8-27B†66.3I 85 · C 81 · S 71 · K 39 · ~$0.104 est.
- 13ZeroEntropy zerank-2†66.0I 63 · C 76 · S 79 · K 50 · $0.047
- 14GPT-5.6 Luna (low)65.9I 95 · C 90 · S 78 · K 28 · $0.242
- 15openjev-sglang65.3I 83 · C 77 · S 77 · K 36 · ~$0.131 est.
- 16Qwen3-Reranker-4B†63.8I 64 · C 67 · S 79 · K 49 · $0.050
- 17reflex-27b†63.3I 86 · C 86 · S 67 · K 32 · ~$0.181 est.
- 18LitJev†62.7I 82 · C 84 · S 67 · K 34 · ~$0.163 est.
- 19kev 0.6B†62.5I 52 · C 51 · S 76 · K 76 · ~$0.0063 est.
- 20SimpleJev Qwen3.6-35B-A3B†62.5I 80 · C 67 · S 75 · K 38 · ~$0.116 est.
- 21djev†62.4I 81 · C 93 · S 75 · K 27 · ~$0.274 est.
- 22jev-local†61.8I 71 · C 69 · S 69 · K 43 · ~$0.077 est.
- 23decider-2b†61.7I 61 · C 47 · S 83 · K 61 · ~$0.020 est.
- 24Bespoke Nimble 9B†60.5I 78 · C 65 · S 79 · K 33 · ~$0.166 est.
- 25Gemini 3.1 Flash-Lite60.1I 86 · C 68 · S 82 · K 27 · $0.264
- 26OpenJev†60.0I 88 · C 70 · S 76 · K 28 · ~$0.255 est.
- 27kev 4B†59.7I 65 · C 42 · S 76 · K 62 · ~$0.019 est.
- 28DeepSeek V4.1 Flash57.5I 94 · C 97 · S 72 · K 17 · $0.594
- 29kev 8B†56.4I 69 · C 44 · S 75 · K 44 · ~$0.073 est.
- 30Open-Jev 9B†55.0I 71 · C 63 · S 72 · K 28 · ~$0.249 est.
- 31system-one54.8I 70 · C 37 · S 84 · K 41 · ~$0.089 est.
- 32jeff†54.4I 47 · C 65 · S 63 · K 77 · ~$0.0060 est.
- 33Laya†54.4I 46 · C 62 · S 71 · K 86 · ~$0.0029 est.
- 34Open-Jev 2B†51.3I 61 · C 55 · S 73 · K 28 · ~$0.249 est.
- 35OpenDecision†40.6I 41 · C 56 · S 80 · K 75 · ~$0.0066 est.
- 36openJev Verdict 1.4†38.9I 39 · C 74 · S 78 · K 82 · ~$0.0039 est.
- 37openJev Verdict†38.1I 40 · C 51 · S 77 · K 83 · ~$0.0037 est.
- 38kev 0.5B†33.2I 38 · C 47 · S 77 · K 76 · ~$0.0063 est.
- 39GLiNER2 large†29.6I 40 · C 24 · S 62 · K 73 · ~$0.0077 est.
- 40smalljev semantic-v9†27.4I 35 · C 59 · S 80 · K 58 · ~$0.025 est.
- 41GLiNER2†24.0I 36 · C 24 · S 72 · K 83 · ~$0.0037 est.
- 42open-jev-deberta-v3-large23.1I 32 · C 66 · S 66 · K 74 · ~$0.0073 est.
- 43GLiNER2.5 multi†16.6I 28 · C 56 · S 68 · K 82 · ~$0.0039 est.
- 44GLiNER2.5 small†13.8I 26 · C 47 · S 78 · K 82 · ~$0.0039 est.
- 45Mixedbread mxbai-rerank-base-v2†0.8I 7 · C 83 · S 88 · K 68 · $0.012
- 46BAAI bge-reranker-v2-m3†0.7I 6 · C 84 · S 90 · K 73 · $0.0077
- 47Alibaba GTE Reranker ModernBERT-base†0.3I 5 · C 77 · S 91 · K 70 · $0.010
- 48Certo v1†0.0I 0 · C 82 · S 94 · K 100 · ~$0.0010 est.
- classifier.dev (fast tier)† (honorable mention)83.6I 85 · C 78 · S 88 · K 84 · ~$0.0033 est.
- Qwen3.8 27B (partial run)24.8I 67 · C 92 · S 61 · K 0 · ~$2.669 est.
- Needle 3, options as tools (partial run)1.1I 14 · C – · S 53 · K 65 · ~$0.014 est.
- Needle 3 (partial run)0.1I 5 · C – · S 60 · K 59 · ~$0.024 est.
Score = Intelligence0.25 × Calibration0.25 × Speed0.25 × Cost0.25 (each 0–100; geometric mean; below 50 Intelligence, × (I / 50)²)
- Jev (TypeSafe, closed)
- Jev rebuild (open, or open source planned)
- Instruction model, JSON schema
- Small tool-calling model
- Service built on Jev
- Zero-shot classifier (not a Jev rebuild)
- Closed decision model (API only, not Jev)
- Shown, not ranked — honorable mention (runs another entrant's model) · partial run
- ~ est. = no measured bill; priced like a large inference provider (how costs are estimated).
- ann. = the provider’s announced price, not yet charged.
- Names link to each project.
- A label-only system has no calibration (–, counted as 0).
- djev (Maisa, diffusion-gemma): The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated).
- Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
- reflex 4B: The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
- jqv: A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
- decision-machine-1: A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
- decider-35b-a3b: The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
- open-alternative-jev (Qwen3.5-4B): With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
- SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
- ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- reflex-27b: The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
- LitJev: The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
- kev 0.6B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
- SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
- djev: Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill.
- jev-local: The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
- decider-2b: The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
- Bespoke Nimble 9B: Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
- OpenJev: OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
- kev 4B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
- kev 8B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
- Open-Jev 9B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
- jeff: Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
- Laya: The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
- Open-Jev 2B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
- OpenDecision: A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
- openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
- openJev Verdict: The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
- kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release.
- GLiNER2 large: The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
- smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
- GLiNER2: A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
- GLiNER2.5 multi: The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
- GLiNER2.5 small: The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
- Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- Certo v1: The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
- classifier.dev (fast tier): Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.
Weighting: Intelligence : Calibration : Speed : Cost
Official default
Custom
The official JevBench Score weights the four axes 25 % each and takes their geometric mean. The other buttons are the earlier views (Balanced 33:33:33 and the three “Emphasis on” weightings, which leave Calibration out), recomputed the same way. Any of them is your view, recomputed in your browser from the published axis scores — not the published score.
Explore by task difficulty
All tasks is the published default. Choose a scope to see how the ranking changes by difficulty. Hard only uses all 220 hard-tier decisions and their measured Intelligence, Calibration, Speed and Cost.
Tier mapping: Easy = easy; Medium = standard. Easy scopes change Intelligence only. Hard only measures all four axes on the same hard-tier subset; systems without a hard-tier run are shown as partial and are not ranked.
What changed in the score
A system that is cheap and fast but barely better than guessing could rank high; intelligence is now measured above chance, and systems below half-way get a growing penalty. The tasks, Calibration, Speed, Cost and ranking eligibility are unchanged.
What the run says (JevBench Score)
- Jev 1.13.0 (TypeSafe AI) leads with 74.4: Intelligence 85.7, Calibration 82.7, Speed 83.3, Cost 52.0 ($0.040 per 1,000 decisions).
- classifier.dev scores 83.6 — higher than anything in the ranking — but is not ranked: it runs Jev (TypeSafe), so ranking it would put the same model in the list twice, once at the model's own price and once at the service's. It keeps every number it earned under Honorable mentions — services built on another entrant's model.
- Open rebuilds of Jev appeared within days. The best of them, SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ), is #2 at 73.1 — 1.3 points behind: more speed and a lower (estimated) price, less intelligence and calibration.
- GPT-5.6 Luna (low reasoning effort) has the highest Intelligence (95.3) but places #14: its cost score is 28.5 ($0.242 per 1,000 decisions), and the geometric mean does not let accuracy buy that back.
- Qwen3.8 27B, Needle 3, options as tools, Needle 3 did not answer every tier — each for the reason in its † note; they are shown below the ranking as partial runs, without a rank.
Axes, tiers, latency and cost
Sort by any column; values the run could not produce always sort last. Hover a cost for how it was priced, a latency for the endpoint. Names link to each project.
⏱ Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.
💲
$ per 1,000 decisions
, not per 1,000 tokens — one decision ≈ 950 input tokens.
| Rank# | System | Endpoint | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | by TypeSafe AIJev 1.13.0 | 74.4 | 85.7 | 82.7 | 83.3 | 52.0 | $0.040 | 100.0% | 99.0% | 94.5% | 74.1% | 0.65 s rawp95 0.72 s raw | production API | |
| 2 | by Theodore Lee (TheoLeeCJ)SemIfformerly OpenJev (Qwen3.5-4B, TheoLeeCJ | 73.1 | 79.0 | 72.6 | 83.7 | 59.5 | ~$0.022 est. | 100.0% | 97.9% | 95.2% | 59.5% | 0.20 s raw→ 0.55 s adjustedp95 0.32 s raw → 0.78 s | our RunPod GPU | |
| 3 | by Maisa (David Villalón)djev†Maisa, diffusion-gemma | 73.0 | 82.7 | 65.4 | 91.4 | 57.6 | $0.026 announced | 100.0% | 97.9% | 93.2% | 69.5% | 0.24 s rawp95 0.31 s raw | production API | |
| 4 | by Eldan RingWinnow-12B Q8† | 71.2 | 82.0 | 72.0 | 82.3 | 52.9 | ~$0.037 est. | 100.0% | 96.9% | 91.1% | 70.9% | 0.23 s raw→ 0.60 s adjustedp95 0.41 s raw → 0.98 s | our RunPod GPU | |
| 5 | by kshetrajna12reflex 4B† | 70.3 | 80.1 | 75.2 | 68.0 | 59.7 | ~$0.022 est. | 100.0% | 94.8% | 97.3% | 63.2% | 1.80 s raw→ 3.75 s adjustedp95 2.05 s raw → 4.26 s | our RunPod GPU | |
| 6 | by hjmurmur (Octalab)jqv†Qwen3-32B zero-shot | 68.6 | 79.3 | 79.0 | 74.6 | 47.5 | ~$0.056 est. | 100.0% | 95.8% | 92.5% | 64.5% | 0.75 s raw→ 1.64 s adjustedp95 0.97 s raw → 2.10 s | our RunPod GPU | |
| 7 | by milliseconds.ai (Baptiste Laget)decision-machine-1†milliseconds.ai | 68.3 | 62.1 | 70.4 | 92.9 | 53.7 | $0.035 | 100.0% | 76.0% | 89.7% | 46.8% | 0.17 s rawp95 0.30 s raw | production API | |
| 8 | by Mapikadecider-35b-a3b† | 67.6 | 79.6 | 71.5 | 80.8 | 45.3 | ~$0.067 est. | 100.0% | 96.9% | 91.1% | 65.5% | 0.29 s raw→ 0.73 s adjustedp95 0.49 s raw → 1.14 s | our RunPod GPU | |
| 9 | by IkerMoelopen-alternative-jev†Qwen3.5-4B, IkerMoel | 67.0 | 64.0 | 63.2 | 83.5 | 59.6 | ~$0.022 est. | 100.0% | 84.4% | 74.7% | 56.8% | 0.21 s raw→ 0.56 s adjustedp95 0.32 s raw → 0.80 s | our RunPod GPU | |
| 10 | by mithalounisystem-one-openGemma 4 E2B LoRA on an L4 | 66.6 | 69.5 | 56.7 | 77.0 | 64.8 | ~$0.015 est. | 100.0% | 93.8% | 87.7% | 49.1% | 0.65 s raw→ 1.30 s adjustedp95 0.77 s raw → 1.54 s | author's demo server | |
| 11 | by razorback16 / CodivOpenJevDiffusionGemma 26B-A4B NVFP4, razorback16 | 66.4 | 79.2 | 64.8 | 83.2 | 45.5 | ~$0.066 est. | 100.0% | 95.8% | 91.1% | 65.5% | 0.24 s raw→ 0.63 s adjustedp95 0.31 s raw → 0.76 s | our RunPod GPU | |
| 12 | by Featherless AISimpleJev Qwen3.8-27B† | 66.3 | 84.7 | 81.1 | 71.2 | 39.5 | ~$0.104 est. | 100.0% | 96.9% | 93.2% | 75.0% | 1.01 s raw→ 2.03 s adjustedp95 1.88 s raw → 3.76 s | author's demo server | |
| 13 | by ZeroEntropyZeroEntropy zerank-2† | 66.0 | 63.0 | 76.5 | 79.0 | 49.8 | $0.047 | 100.0% | 79.2% | 88.4% | 47.3% | 0.13 s raw→ 0.40 s adjustedp95 1.50 s raw → 3.15 s | our RunPod GPU | |
| 14 | by OpenAIGPT-5.6 Lunalow reasoning effort | 65.9 | 95.3 | 89.8 | 77.5 | 28.5 | $0.242 | 100.0% | 97.9% | 96.6% | 94.5% | 0.97 s rawp95 1.82 s raw | production API | |
| 15 | by ekzhangopenjev-sglangQwen3.6-35B-A3B on SGLang | 65.3 | 83.4 | 77.4 | 77.1 | 36.5 | ~$0.131 est. | 100.0% | 95.8% | 95.2% | 71.4% | 0.68 s raw→ 1.36 s adjustedp95 0.73 s raw → 1.45 s | author's demo server | |
| 16 | by QwenQwen3-Reranker-4B† | 63.8 | 64.0 | 67.0 | 78.7 | 49.2 | $0.050 | 100.0% | 79.2% | 87.7% | 50.0% | 0.13 s raw→ 0.41 s adjustedp95 1.56 s raw → 3.27 s | our RunPod GPU | |
| 17 | by kshetrajna12reflex-27b†Qwen3.8-27B | 63.3 | 85.8 | 86.2 | 67.5 | 32.3 | ~$0.181 est. | 100.0% | 95.8% | 95.9% | 75.9% | 1.89 s raw→ 3.93 s adjustedp95 2.21 s raw → 4.57 s | our RunPod GPU | |
| 18 | by Zhengxu YuLitJev†Qwen3.8-27B | 62.7 | 82.4 | 83.5 | 66.7 | 33.6 | ~$0.163 est. | 100.0% | 97.9% | 88.4% | 73.2% | 2.03 s raw→ 4.20 s adjustedp95 2.46 s raw → 5.06 s | our RunPod GPU | |
| 19 | by Jared Palmerkev 0.6B†research preview | 62.5 | 51.9 | 51.1 | 75.6 | 76.1 | ~$0.0063 est. | 100.0% | 81.3% | 66.4% | 40.0% | 0.59 s raw→ 1.33 s adjustedp95 0.97 s raw → 2.09 s | our RunPod GPU | |
| 20 | by Featherless AISimpleJev Qwen3.6-35B-A3B† | 62.5 | 79.5 | 67.1 | 75.0 | 38.1 | ~$0.116 est. | 100.0% | 93.8% | 93.2% | 66.4% | 0.85 s raw→ 1.70 s adjustedp95 0.93 s raw → 1.86 s | author's demo server | |
| 21 | by David Villalon / Maisadjev†thinking | 62.4 | 80.8 | 92.7 | 75.2 | 26.9 | ~$0.274 est. | 95.8% | 99.0% | 80.1% | 77.7% | 0.43 s raw→ 1.00 s adjustedp95 1.45 s raw → 3.05 s | our RunPod GPU | |
| 22 | by us (GitHub)jev-local†Qwen3.5-9B | 61.8 | 70.8 | 68.7 | 69.2 | 43.3 | ~$0.077 est. | 100.0% | 84.4% | 89.0% | 59.1% | 1.05 s raw→ 2.24 s adjustedp95 2.62 s raw → 5.38 s | our RunPod GPU | |
| 23 | by Mapikadecider-2b† | 61.7 | 61.2 | 46.6 | 83.2 | 61.0 | ~$0.020 est. | 100.0% | 85.4% | 77.4% | 47.3% | 0.26 s raw→ 0.67 s adjustedp95 0.28 s raw → 0.72 s | our RunPod GPU | |
| 24 | by Bespoke LabsBespoke Nimble 9B† | 60.5 | 77.9 | 65.3 | 78.7 | 33.4 | ~$0.166 est. | 100.0% | 94.8% | 89.0% | 65.5% | 0.39 s raw→ 0.93 s adjustedp95 0.65 s raw → 1.46 s | our RunPod GPU | |
| 25 | by GoogleGemini 3.1 Flash-Lite | 60.1 | 85.6 | 68.1 | 81.8 | 27.4 | $0.264 | 100.0% | 99.0% | 93.2% | 75.0% | 0.76 s rawp95 0.88 s raw | production API | |
| 26 | by razorback16OpenJev†thinking, BF16 | 60.0 | 88.0 | 69.6 | 76.1 | 27.8 | ~$0.255 est. | 100.0% | 100.0% | 94.5% | 78.2% | 0.46 s raw→ 1.08 s adjustedp95 1.08 s raw → 2.31 s | our RunPod GPU | |
| 27 | by Jared Palmerkev 4B†research preview | 59.7 | 64.8 | 42.0 | 75.7 | 61.8 | ~$0.019 est. | 100.0% | 91.7% | 85.6% | 42.3% | 0.55 s raw→ 1.25 s adjustedp95 0.99 s raw → 2.13 s | our RunPod GPU | |
| 28 | by DeepSeekDeepSeek V4.1 Flashthinking default | 57.5 | 94.3 | 96.7 | 71.6 | 16.8 | $0.594 | 98.6% | 99.0% | 93.2% | 95.0% | 1.42 s rawp95 4.89 s raw | production API | |
| 29 | by Jared Palmerkev 8B†research preview | 56.4 | 69.4 | 44.2 | 74.9 | 44.0 | ~$0.073 est. | 100.0% | 92.7% | 90.4% | 47.3% | 0.59 s raw→ 1.33 s adjustedp95 1.15 s raw → 2.45 s | our RunPod GPU | |
| 30 | by Zefan Cai (@Zefan_Cai)Open-Jev 9B†Zefan Cai | 55.0 | 71.2 | 63.3 | 72.0 | 28.1 | ~$0.249 est. | 100.0% | 90.6% | 81.5% | 60.9% | 0.75 s raw→ 1.66 s adjustedp95 1.81 s raw → 3.77 s | our RunPod GPU | |
| 31 | by Sean Goedeckesystem-oneQwen3-8B, Sean Goedecke | 54.8 | 70.3 | 36.8 | 84.4 | 41.5 | ~$0.089 est. | 100.0% | 90.6% | 91.8% | 50.0% | 0.17 s raw→ 0.48 s adjustedp95 0.30 s raw → 0.76 s | our RunPod GPU | |
| 32 | by Logan Markewichjeff†Logan Markewich, GLiFormer 400M | 54.4 | 46.9 | 64.6 | 63.5 | 76.6 | ~$0.0060 est. | 100.0% | 76.0% | 61.6% | 37.7% | 0.94 s raw→ 2.03 s adjustedp95 10.97 s raw → 22.09 s | our CPU | |
| 33 | by Convai InnovationsLaya†Convai Innovations, ModernBERT-large 421M | 54.4 | 45.8 | 62.5 | 71.1 | 86.2 | ~$0.0029 est. | 94.4% | 72.9% | 69.2% | 34.1% | 0.79 s raw→ 1.72 s adjustedp95 2.20 s raw → 4.54 s | our CPU | |
| 34 | by Zefan Cai (@Zefan_Cai)Open-Jev 2B†Zefan Cai | 51.3 | 61.0 | 55.1 | 73.5 | 28.1 | ~$0.249 est. | 100.0% | 79.2% | 88.4% | 42.7% | 0.66 s raw→ 1.48 s adjustedp95 1.45 s raw → 3.05 s | our RunPod GPU | |
| 35 | by Deepan WadhwaOpenDecision†ModernBERT-large zero-shot | 40.6 | 40.8 | 56.1 | 79.9 | 75.3 | ~$0.0066 est. | 87.5% | 62.5% | 71.2% | 33.2% | 0.34 s raw→ 0.83 s adjustedp95 0.54 s raw → 1.24 s | our RunPod GPU | |
| 36 | by Hemant (heman10x)openJev Verdict 1.4† | 38.9 | 38.6 | 74.1 | 78.1 | 82.4 | ~$0.0039 est. | 86.1% | 67.7% | 56.2% | 37.7% | 0.31 s raw→ 0.78 s adjustedp95 0.92 s raw → 2.00 s | our CPU | |
| 37 | by Hemant (heman10x)openJev Verdict†heman10x, ModernBERT-base 151M | 38.1 | 39.8 | 51.3 | 76.7 | 83.1 | ~$0.0037 est. | 86.1% | 65.6% | 61.0% | 38.2% | 0.28 s raw→ 0.71 s adjustedp95 1.45 s raw → 3.04 s | our CPU | |
| 38 | by Jared Palmerkev 0.5B† | 33.2 | 38.2 | 47.4 | 77.0 | 76.1 | ~$0.0063 est. | 95.8% | 52.1% | 71.2% | 30.9% | 0.43 s raw→ 1.01 s adjustedp95 0.92 s raw → 1.99 s | our RunPod GPU | |
| 39 | by FastinoGLiNER2 large† | 29.6 | 40.1 | 24.3 | 61.7 | 73.3 | ~$0.0077 est. | 98.6% | 62.5% | 61.0% | 36.4% | 1.10 s raw→ 2.34 s adjustedp95 14.49 s raw → 29.13 s | our CPU | |
| 40 | by Aditya (isHeSatoshi)smalljev semantic-v9† | 27.4 | 35.1 | 58.9 | 79.8 | 57.9 | ~$0.025 est. | 97.2% | 68.8% | 40.4% | 38.2% | 0.41 s raw→ 0.98 s adjustedp95 0.46 s raw → 1.07 s | our RunPod GPU | |
| 41 | by FastinoGLiNER2†Fastino, gliner2.5-base | 24.0 | 35.6 | 23.7 | 71.8 | 83.1 | ~$0.0037 est. | 97.2% | 66.7% | 45.9% | 36.4% | 0.31 s raw→ 0.78 s adjustedp95 4.15 s raw → 8.46 s | our CPU | |
| 42 | by Kotoba Labsopen-jev-deberta-v3-largelocal CPU | 23.1 | 31.9 | 66.4 | 66.0 | 74.0 | ~$0.0073 est. | 100.0% | 49.0% | 53.4% | 36.4% | 1.77 s raw→ 3.69 s adjustedp95 3.35 s raw → 6.85 s | our CPU | |
| 43 | by FastinoGLiNER2.5 multi†Fastino, 287M | 16.6 | 27.7 | 56.1 | 67.8 | 82.4 | ~$0.0039 est. | 90.3% | 51.0% | 43.8% | 37.7% | 0.43 s raw→ 1.01 s adjustedp95 8.18 s raw → 16.50 s | our CPU | |
| 44 | by FastinoGLiNER2.5 small†Fastino, 74M | 13.8 | 25.6 | 47.2 | 77.8 | 82.4 | ~$0.0039 est. | 83.3% | 47.9% | 50.0% | 33.2% | 0.11 s raw→ 0.38 s adjustedp95 2.10 s raw → 4.35 s | our CPU | |
| 45 | by MixedbreadMixedbread mxbai-rerank-base-v2† | 0.8 | 6.7 | 83.1 | 87.5 | 67.9 | $0.012 | 44.4% | 33.3% | 26.7% | 40.0% | 0.07 s raw→ 0.29 s adjustedp95 0.23 s raw → 0.62 s | our RunPod GPU | |
| 46 | by BAAIBAAI bge-reranker-v2-m3† | 0.7 | 6.3 | 83.8 | 89.5 | 73.4 | $0.0077 | 43.1% | 36.5% | 8.9% | 36.8% | 0.03 s raw→ 0.22 s adjustedp95 0.18 s raw → 0.51 s | our RunPod GPU | |
| 47 | by Alibaba-NLPAlibaba GTE Reranker ModernBERT-base† | 0.3 | 4.6 | 76.8 | 90.6 | 69.6 | $0.010 | 33.3% | 39.6% | 30.1% | 33.6% | 0.05 s raw→ 0.25 s adjustedp95 0.10 s raw → 0.35 s | our RunPod GPU | |
| 48 | by AltSlate LabsCerto v1† | 0.0 | 0.0 | 82.0 | 94.0 | 100.0 | ~$0.0010 est. | 27.8% | 30.2% | 21.9% | 31.8% | 0.02 s raw→ 0.19 s adjustedp95 0.03 s raw → 0.21 s | our RunPod GPU | |
| Honorable mentions — services built on another entrant's model — shown, not ranked: A service that runs another entrant's model is listed with all of its scores and axes, but is not ranked against the models. | ||||||||||||||
| by mrmps (@michael_chomsky)classifier.dev†fast tierhonorable mention · not ranked | 83.6 | 85.1 | 77.9 | 87.6 | 84.3 | ~$0.0033 est. | 100.0% | 99.0% | 97.3% | 70.5% | 0.39 s rawp95 0.45 s raw | production API | ||
| Partial runs — shown, not ranked: a tier attempted for fewer than 95 % of its decisions. | ||||||||||||||
| by Qwen / ChutesQwen3.8 27BChutes TEEpartial run · not ranked | 24.8 | 67.4 | 92.1 | 61.3 | 0.0 | ~$2.669 est. | 98.6% | 99.0% | 95.3% | 21.4% | 5.75 s rawp95 12.97 s raw | production API | ||
| by Cactus ComputeNeedle 3, options as toolspost-hoc adapter modepartial run · not ranked | 1.1 | 13.5 | none (label only) | 52.8 | 65.3 | ~$0.014 est. | 66.7% | 31.3% | 34.2% | — | 3.78 s raw→ 7.71 s adjustedp95 33.64 s raw → 67.42 s | our CPU | ||
| by Cactus ComputeNeedle 3Cactus, 2-bit, local CPUpartial run · not ranked | 0.1 | 4.6 | none (label only) | 59.9 | 58.7 | ~$0.024 est. | 47.2% | 16.7% | 31.5% | 7.7% | 1.69 s raw→ 3.52 s adjustedp95 14.36 s raw → 28.88 s | our CPU | ||
† Notes on 39 marked systems — how each was run
- † djev: The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated).
- † Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
- † reflex 4B: The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
- † jqv: A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
- † decision-machine-1: A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
- † decider-35b-a3b: The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
- † open-alternative-jev: With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
- † SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
- † ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- † Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- † reflex-27b: The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
- † LitJev: The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
- † kev 0.6B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
- † SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
- † djev: Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill.
- † jev-local: The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
- † decider-2b: The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
- † Bespoke Nimble 9B: Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
- † OpenJev: OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
- † kev 4B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
- † kev 8B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
- † Open-Jev 9B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
- † jeff: Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
- † Laya: The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
- † Open-Jev 2B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
- † OpenDecision: A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
- † openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
- † openJev Verdict: The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
- † kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release.
- † GLiNER2 large: The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
- † smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
- † GLiNER2: A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
- † GLiNER2.5 multi: The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
- † GLiNER2.5 small: The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
- † Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- † BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- † Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
- † Certo v1: The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
- † classifier.dev: Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.
Honorable mentions — services built on another entrant's model
A service that runs another entrant's model is listed with all of its scores and axes, but is not ranked against the models.
classifier.devno rank
83.6 JevBench Score · Jev 1.13.0 (#1) scores 74.4
Runs on Jev (TypeSafe).
- Intelligence85.1
- Calibration77.9
- Speed87.6
- Cost84.3
- $ per 1,000 decisions~$0.0033 est.
Runs on Jev (TypeSafe) — listed, not ranked. Ranking it against Jev would rank Jev's model against Jev's model, so from v1.2.4 it is an honorable mention instead of #1.
Why it is not ranked, what its price assumes, and what we found
classifier.dev is not its own model. Its own pages say so: "The fast tier is Jev, TypeSafe's decision model" (https://classifier.dev/benchmark, read 2026-09-20), and the API answers with "model": "jev-1.13.0" — the same model version this benchmark measures directly as Jev 1.13.0. What it adds is a price and, on its smart tier, an orchestration layer: "The smart tier is Jev plus a reasoning model re-asking only the answers Jev put under 0.7 confidence" — escalation on low confidence (a model cascade), not best-of-N, not self-consistency and not a committee. Its published escalation model is gemini-3.8-flash. Ranking it against Jev would rank Jev's model against Jev's model, so from v1.2.4 it is an honorable mention instead of #1.
Only the fast tier was measured. The smart tier's escalation was never run, so nothing here scores it.
Price. $0.0033 per 1,000 decisions is an estimate from the published flat-rate plan at full use: classifier.dev Pro is $20/month for 200,000 fast classifications a day (https://classifier.dev/pricing, read 2026-09-20), and one classification is one decision. Lower use costs more per decision — at a tenth of that allowance it is $0.033 per 1,000 — and the free tier (20,000 fast classifications a day), which is what our run used, costs nothing. Their pages do not say how the flat rate is funded, so we do not know their cost basis; the only figure they publish is what the model costs a caller: "The model behind the fast tier costs about $0.005 per thousand classifications and needs a TypeSafe key" (https://classifier.dev/pricing) — for their short single-sentence inputs, not for JevBench's whole questions.
Not a pass-through. On our set the fast tier scored 97.3 % on the judge tier against Jev's 94.5 %, and 70.5 % against 74.1 % on the hard tier. classifier.dev's own explanation for differences of this kind is batching ("The fast tier is Jev, packed a thousand to a request"); on their own two test sets they measured the same difference as noise.
A legitimate, well-documented product: free without an account, open source (https://github.com/mrmps/classifier-dev), by Michael Ryaboy (@michael_chomsky). Read 2026-09-20: classifier.dev · classifier.dev/benchmark · classifier.dev/pricing · classifier.dev/about
Which public tasks did each system get right?
This view shows public task outcomes only: 231 of 231 public tasks in the selected scope. Held-out and imported task text is not shipped.
Show 231 public task outcomes across 52 systems
| Task | Jev 1.13.0 | SemIf | djev | Winnow-12B Q8 | reflex 4B | jqv | decision-machine-1 | decider-35b-a3b | open-alternative-jev | system-one-open | OpenJev | SimpleJev Qwen3.8-27B | ZeroEntropy zerank-2 | GPT-5.6 Luna | openjev-sglang | Qwen3-Reranker-4B | reflex-27b | LitJev | kev 0.6B | SimpleJev Qwen3.6-35B-A3B | djev | jev-local | decider-2b | Bespoke Nimble 9B | Gemini 3.1 Flash-Lite | OpenJev | kev 4B | DeepSeek V4.1 Flash | kev 8B | Open-Jev 9B | system-one | jeff | Laya | Open-Jev 2B | OpenDecision | openJev Verdict 1.4 | openJev Verdict | kev 0.5B | GLiNER2 large | smalljev semantic-v9 | GLiNER2 | open-jev-deberta-v3-large | GLiNER2.5 multi | GLiNER2.5 small | Mixedbread mxbai-rerank-base-v2 | BAAI bge-reranker-v2-m3 | Alibaba GTE Reranker ModernBERT-base | Certo v1 | classifier.dev | Qwen3.8 27B | Needle 3, options as tools | Needle 3 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Easy · 48 of 72 decisions publicEasy · 48 of 72 public | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 46/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 48/48 | 46/48 | 48/48 | 42/48 | 42/48 | 41/48 | 46/48 | 48/48 | 47/48 | 47/48 | 48/48 | 44/48 | 41/48 | 22/48 | 23/48 | 17/48 | 12/48 | 48/48 | 48/48 | 32/48 | 23/48 |
| easy-intent-00intent- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | × | × | ✓ | ✓ | ✓ | × |
| easy-intent-01intent- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | × | ✓ | ✓ | ✓ | ✓ |
| easy-intent-02intent- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | × | ✓ | ✓ | × | × |
| easy-intent-03intent- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | × |
| easy-intent-04intent- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | × | × |
| easy-intent-05intent- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | × | ✓ | ✓ | ✓ | × |
| easy-intent-06intent- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | × |
| easy-intent-07intent- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | × | ✓ | ✓ | ✓ | ✓ |
| easy-intent-08intent- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | × | × | ✓ | ✓ | ✓ | × |
| easy-intent-09intent- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | × |
| easy-intent-10intent- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | × |
| easy-intent-11intent- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ |
| easy-fact-00fact- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-fact-01fact- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | × | × | ✓ | ✓ | ✓ | × | ✓ | × | × | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ |
| easy-fact-02fact- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | × | ✓ | ✓ | ✓ | ✓ |
| easy-fact-03fact- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | × |
| easy-fact-04fact- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | × | ✓ | ✓ | ✓ | × | ✓ |
| easy-fact-05fact- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | × | × |
| easy-fact-06fact- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-fact-07fact- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | × | × |
| easy-fact-08fact- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-fact-09fact- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | × | × |
| easy-fact-10fact- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | × | × |
| easy-fact-11fact- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | × | × |
| easy-extraction-00extraction- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | × | ✓ | ✓ | × | ✓ |
| easy-extraction-01extraction- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | ✓ | × | ✓ |
| easy-extraction-02extraction- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | × | ✓ |
| easy-extraction-03extraction- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | ✓ | ✓ | ✓ |
| easy-extraction-04extraction- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | × | ✓ | ✓ | × | ✓ |
| easy-extraction-05extraction- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ |
| easy-extraction-06extraction- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | ✓ | ✓ |
| easy-extraction-07extraction- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ |
| easy-extraction-08extraction- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ |
| easy-extraction-09extraction- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | × | ✓ |
| easy-extraction-10extraction- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × | ✓ | ✓ | × | ✓ |
| easy-extraction-11extraction- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ |
| easy-tool_selection-00tool_ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ |
| easy-tool_selection-01tool_ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | ✓ | ✓ | × |
| easy-tool_selection-02tool_ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | × |
| easy-tool_selection-03tool_ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ! | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | ✓ | × |
| easy-tool_selection-04tool_ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | × | ✓ | ✓ | ✓ | × |
| easy-tool_selection-05tool_ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | ✓ | × |
| easy-tool_selection-06tool_ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | × | ✓ | ✓ | ✓ | × |
| easy-tool_selection-07tool_ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ! | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | × | ✓ | ✓ | ✓ | ✓ |
| easy-tool_selection-08tool_ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | × | × |
| easy-tool_selection-09tool_ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | ✓ | × |
| easy-tool_selection-10tool_ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × |
| easy-tool_selection-11tool_ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | × |
| Medium (standard) · 72 of 96 decisions publicMedium (standard) · 72 of 96 public | 71/72 | 71/72 | 71/72 | 69/72 | 68/72 | 69/72 | 54/72 | 70/72 | 60/72 | 67/72 | 70/72 | 70/72 | 57/72 | 70/72 | 68/72 | 54/72 | 69/72 | 71/72 | 58/72 | 67/72 | 71/72 | 60/72 | 61/72 | 67/72 | 71/72 | 72/72 | 64/72 | 71/72 | 67/72 | 65/72 | 64/72 | 54/72 | 50/72 | 55/72 | 43/72 | 50/72 | 45/72 | 35/72 | 42/72 | 49/72 | 46/72 | 31/72 | 32/72 | 30/72 | 24/72 | 26/72 | 26/72 | 24/72 | 71/72 | 71/72 | 19/72 | 12/72 |
| original-policy-01-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | × | × |
| original-policy-01-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ! | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | × | ✓ | × | ✓ | ✓ | × | × |
| original-policy-02-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | × | ✓ |
| original-policy-02-1original- | ✓ | × | ✓ | ✓ | × | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × | ✓ | × | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | × | × |
| original-policy-03-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | × | × | ✓ | × | × | × | ✓ | × | ✓ | × | × | ✓ | ✓ | × | × |
| original-policy-03-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × | × | × | × | × | × | × | ✓ | × | × | ✓ | ✓ | × | × |
| original-policy-04-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | × | ✓ | × | ✓ | ✓ | × | × |
| original-policy-04-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | × | ✓ | × | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | × | × |
| original-policy-05-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | × | × |
| original-policy-05-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | × | × |
| original-policy-06-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × | ✓ | × | × | ✓ | ✓ | × | ✓ | × | × | ✓ | ✓ | × | × |
| original-policy-06-1original- | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | × | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | × | × | × | ✓ | × | × |
| original-intent-01-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | × | × | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | × |
| original-intent-01-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | × | × | ✓ | ✓ | × | ✓ |
| original-intent-02-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | × | × | ✓ | × | × | × | × | × | × | × | × | × | × | × | × | × | ✓ | ✓ | ✓ | × |
| original-intent-02-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | × | × | × | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | × | × | × | ✓ | ✓ | ✓ | × |
| original-intent-03-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | × | × | × | × | ✓ | ✓ | × | × | × | × | × | ✓ | ✓ | × | × |
| original-intent-03-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | × | ✓ | ✓ | ✓ | × |
| original-intent-04-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | × | ✓ | ✓ | ✓ | × |
| original-intent-04-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | × | × | × | × | ✓ | ✓ | ✓ | × |
| original-intent-05-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | × | ✓ | ✓ | ✓ | × | × | ✓ | × | ✓ | × | ✓ | × | ✓ | × | ✓ | ✓ | × | × |
| original-intent-05-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | × | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | × | ✓ | × | ✓ | × | ✓ | ✓ | × | × |
| original-intent-06-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | × | ✓ | ✓ | × | ✓ | × | ✓ | × | × | ✓ | × | × | × | × | ✓ | ✓ | ✓ | × |
| original-intent-06-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | × | ✓ | × | × | ✓ | × | ✓ | × | ✓ | ✓ | × | × | × | × | ✓ | ✓ | ✓ | × |
| original-ordinal-01-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | × | ✓ | × | × | ✓ | ✓ | × | × |
| original-ordinal-01-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | × | ✓ | × | ✓ | ✓ | ✓ | × | × |
| original-ordinal-02-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | × | × | × | × | ✓ | ✓ | × | × |
| original-ordinal-02-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | × | × | × | × | ✓ | ✓ | × | × |
| original-ordinal-03-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | × | × | ✓ | ✓ | ✓ | × | × |
| original-ordinal-03-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | × | × | × | ✓ | ✓ | ✓ | × | × |
| original-ordinal-04-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | × | × |
| original-ordinal-04-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | × | × |
| original-ordinal-05-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ! | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | × | ✓ | ✓ | ✓ | × | × | × | × | ✓ | × | ✓ | ✓ | ✓ | × | × |
| original-ordinal-05-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | × | × | ✓ | × | ✓ | ✓ | ✓ | × | × |
| original-ordinal-06-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | ✓ | × | × |
| original-ordinal-06-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | × | × |
| original-extraction-01-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | × | × | × | × | × | ✓ | × | × | × | × | × | × | × | ✓ | ✓ | ✓ | × | × |
| original-extraction-01-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | × | × | ✓ | × | ✓ | × | × | × | × | × | ✓ | ✓ | ✓ | × | ✓ |
| original-extraction-02-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × |
| original-extraction-02-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × |
| original-extraction-03-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | × | ✓ | × | × | × | × | × | × | ✓ | ✓ | × | ✓ |
| original-extraction-03-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | × | ✓ | × | × | × | × | ✓ | ✓ | × | × |
| original-extraction-04-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | × | × | ✓ | ✓ | ✓ | × | × | × | × | × | ✓ | ✓ | × | × |
| original-extraction-04-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | × | ✓ | × | ✓ | × | × | × | × | ✓ | ✓ | ✓ | ✓ |
| original-extraction-05-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × |
| original-extraction-05-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × |
| original-extraction-06-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | × | ✓ | ✓ | ✓ | ✓ |
| original-extraction-06-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | × | ✓ | × | × | × | × | × | × | ✓ | ✓ | × | ✓ |
| original-adequacy-01-0original- | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | × | × | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ |
| original-adequacy-01-1original- | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | × | × | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ |
| original-adequacy-02-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | × | × | ✓ | × | × | × | ✓ | × | ✓ | ✓ | ✓ | × | × | × | × | ✓ | ✓ | ✓ | ✓ | × | × |
| original-adequacy-02-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| original-adequacy-03-0original- | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | × | × | ✓ | × | × | × | × | ✓ | × | ✓ | × | ✓ | × | × | × | × | × | × | × | × | × | × | ✓ | × | × | × | × | × | ✓ | × | ✓ | ✓ | × | × |
| original-adequacy-03-1original- | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | × | ✓ | × | ✓ | × | × | × | ✓ | × | × | ✓ | × | × | × | ✓ | ✓ | × | ✓ | × | ✓ | × | × | ✓ | × | × | × | × | × | × | × | ✓ | × | × | × | × | × | ✓ | × | ✓ | ✓ | × | × |
| original-adequacy-04-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | × | × | × | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | × |
| original-adequacy-04-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | × | × | ✓ | × | × | × | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | × |
| original-adequacy-05-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | × | ✓ | × | × | ✓ | ✓ | × | × | ✓ | × | × | × | ✓ | ✓ | × | ✓ | × | × | × | × | ✓ | × | × | × | × | ✓ | × | ✓ | ✓ | ✓ | × | × | × | × | ✓ | × | ✓ | ✓ | × | × |
| original-adequacy-05-1original- | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × | × | ✓ | ✓ | × | × | × | × | × | ✓ | ✓ | × | × | × | × | × | ✓ | × | ✓ | ✓ | × | × |
| original-adequacy-06-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | × | × | × | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ |
| original-adequacy-06-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | ✓ | ✓ | × | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ |
| original-routing-01-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | × | ✓ | × | × | ✓ | ✓ | ✓ | × |
| original-routing-01-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | × | × | × | ✓ | × | × | × | × | × | × | × | × | ✓ | ✓ | × | × |
| original-routing-02-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | × | × | × | × | ✓ | × | ✓ | ✓ | × | × |
| original-routing-02-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | × | × | × | × | ✓ | × | ✓ | ✓ | × | × |
| original-routing-03-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × |
| original-routing-03-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × |
| original-routing-04-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | × | × | × | × | × | × | ✓ | ✓ | ✓ | × |
| original-routing-04-1original- | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | × | × | × | ✓ | ✓ | × | ✓ | × | × | × | × | × | × | × | × | × | ✓ | ✓ | ✓ | × |
| original-routing-05-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | ✓ | × | ✓ | ✓ | × | × | ✓ | × | × | × | × | × | × | × | × | ✓ | ✓ | × | × |
| original-routing-05-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | × | × | ✓ | × | × | × | × | × | × | × | × | × | × | × | × | × | ✓ | ✓ | × | × |
| original-routing-06-0original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × | × | ✓ | × | ✓ | × | × | × | × | × | × | × | ✓ | ✓ | ! | ✓ | × |
| original-routing-06-1original- | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | × |