JevBench v1.4.1 发布:面向类型化决策模型的可重复基准测试

Hacker News 热门(buzzing.cc 中文翻译)·2026-09-23 16:00·1天前·florianstandhar
AI 导读

JevBench v1.4.1 发布,这是 Benchmark Heaven 面向 Jev 类决策模型的可重复基准,输入状态与有界评分标准、输出类型化答案,基于 534 个公开加 308 个密封决策、协议 jevbench::v1.4 评分。榜单共 77 个系统,Jev 1.13.0 以 77 分居首,JevK5 v0.2.0 与 Hopper 分列二、三。

Hacker News 热门(buzzing.cc 中文翻译)
44AI 编辑部评分,满分 100

JevBench v1.4.1 发布:面向类型化决策模型的可重复基准测试

2026-09-23 16:00· 1天前· florianstandhar
AI 导读

JevBench v1.4.1 发布,这是 Benchmark Heaven 面向 Jev 类决策模型的可重复基准,输入状态与有界评分标准、输出类型化答案,基于 534 个公开加 308 个密封决策、协议 jevbench::v1.4 评分。榜单共 77 个系统,Jev 1.13.0 以 77 分居首,JevK5 v0.2.0 与 Hopper 分列二、三。

JevBench is Benchmark Heaven's own benchmark for Jev-class decision models: state and a bounded rubric in, a typed answer out.

Compare Jev alternativesChoose a Jev-class model by use case

Scored 23 Sept 2026 · protocol jevbench::v1.4 · 534 public + 308 sealed aggregate decisions · one request at a time from a server in Germany · harness, public tasks & scoring rules (MIT) · results JSON sha256 808909c4ecc8… · v1.0 results

Share this version

JevBench v1.4.1 ranking

媒体内容 · 前往原文查看

JevBench v1.4.1

JevBench Score: 77 ranked systems

Official· four axes 0–100, equal-weight harmonic mean · What changed in v1.4 ↓

Intel.Calib.SpeedCost$/1k dec.
  1. 1Jev 1.13.0API63.3I 53 · C 76 · S 83 · K 52 · $0.040
  2. 2JevK5 v0.2.062.0I 49 · C 75 · S 91 · K 60 · ~$0.022 est.
  3. 3Hopper59.4I 48 · C 79 · S 87 · K 59 · ~$0.024 est.
  4. 4Winnow-12B Q855.6I 48 · C 65 · S 82 · K 53 · ~$0.037 est.
  5. 5reflex 4B54.0I 47 · C 70 · S 68 · K 60 · ~$0.022 est.
  6. 6djev52.2I 47 · C 55 · S 91 · K 58 · $0.026 ann.
  7. 7Jev-Omni51.3I 47 · C 64 · S 82 · K 53 · ~$0.037 est.
  8. 8metask-jev-4b47.8I 45 · C 67 · S 89 · K 55 · ~$0.033 est.
  9. 9SemIf47.7I 44 · C 67 · S 84 · K 59 · ~$0.022 est.
  10. 10Jobe Qwen3.5-4B46.9I 44 · C 66 · S 86 · K 60 · ~$0.022 est.
  11. 11local-jev Qwen3.5-4B46.8I 44 · C 73 · S 75 · K 56 · ~$0.030 est.
  12. 12system-one-openAPI45.1I 44 · C 55 · S 77 · K 65 · ~$0.015 est.
  13. 13spark-s1-4b-v644.6I 45 · C 48 · S 81 · K 58 · ~$0.025 est.
  14. 14jqv44.4I 46 · C 72 · S 75 · K 47 · ~$0.056 est.
  15. 15Qwen3-Reranker-4B43.5I 45 · C 65 · S 79 · K 49 · $0.050
  16. 16decider-35b-a3b41.2I 47 · C 65 · S 81 · K 45 · ~$0.067 est.
  17. 17Raw Qwen3 4B Instruct 2507 direct logits41.0I 46 · C 29 · S 88 · K 60 · ~$0.022 est.
  18. 18OpenSourceJev40.9I 42 · C 60 · S 82 · K 64 · ~$0.016 est.
  19. 19ZeroEntropy zerank-240.2I 42 · C 76 · S 79 · K 50 · $0.047
  20. 20decision-machine-1API39.9I 41 · C 68 · S 93 · K 54 · $0.035
Show all 82 systems (57 more ranked, 5 not ranked)
  1. 21Raw Phi-4 mini direct logits38.0I 42 · C 59 · S 89 · K 50 · ~$0.048 est.
  2. 22JEV Qwen3.5-9B Base NVFP437.7I 47 · C 68 · S 93 · K 43 · ~$0.077 est.
  3. 23OpenJev36.9I 45 · C 55 · S 83 · K 45 · ~$0.066 est.
  4. 24kev 4B36.1I 42 · C 40 · S 76 · K 62 · ~$0.019 est.
  5. 25Decision 2B35.8I 39 · C 74 · S 84 · K 63 · ~$0.018 est.
  6. 26Qwen3.5-9B Jev-like data-mix v235.2I 47 · C 61 · S 82 · K 42 · ~$0.083 est.
  7. 27GPT-6 LunaAPI35.1I 96 · C 92 · S 74 · K 37 · $0.127
  8. 28SimpleJev Qwen3.8-27BAPI34.6I 52 · C 74 · S 71 · K 39 · ~$0.104 est.
  9. 29NInfer Qwen3.8-Flash-Next mixed34.0I 50 · C 79 · S 88 · K 39 · ~$0.109 est.
  10. 30GPT-6 LunaAPI33.3I 97 · C 93 · S 73 · K 36 · $0.135
  11. 31open-alternative-jev33.2I 39 · C 59 · S 83 · K 60 · ~$0.022 est.
  12. 32jev-local32.5I 45 · C 64 · S 69 · K 43 · ~$0.077 est.
  13. 33Decision Fast32.5I 37 · C 65 · S 82 · K 76 · ~$0.0063 est.
  14. 34decider-2b30.7I 39 · C 43 · S 83 · K 61 · ~$0.020 est.
  15. 35jeff30.6I 37 · C 68 · S 63 · K 77 · ~$0.0060 est.
  16. 36Laya30.3I 36 · C 64 · S 71 · K 86 · ~$0.0029 est.
  17. 37lev-350m28.5I 35 · C 71 · S 85 · K 76 · ~$0.0063 est.
  18. 38openjev-sglangAPI27.7I 49 · C 69 · S 77 · K 36 · ~$0.131 est.
  19. 39Von27.5I 34 · C 76 · S 70 · K 78 · ~$0.0055 est.
  20. 40NInfer Qwen3.8-27B NVFP426.9I 51 · C 76 · S 80 · K 35 · ~$0.145 est.
  21. 41NInfer Qwen3.8-27B NVFP426.3I 51 · C 67 · S 80 · K 35 · ~$0.145 est.
  22. 42kev 8B25.6I 42 · C 40 · S 75 · K 44 · ~$0.073 est.
  23. 43JevOne25.5I 48 · C 78 · S 88 · K 36 · ~$0.137 est.
  24. 44SimpleJev Qwen3.6-35B-A3BAPI24.9I 46 · C 60 · S 75 · K 38 · ~$0.116 est.
  25. 45kev 0.6B24.8I 34 · C 50 · S 76 · K 76 · ~$0.0063 est.
  26. 46Raw Qwen3 8B direct logits23.7I 46 · C 24 · S 86 · K 42 · ~$0.087 est.
  27. 47system-one23.4I 44 · C 33 · S 84 · K 41 · ~$0.089 est.
  28. 48OpenDecision21.6I 32 · C 57 · S 80 · K 75 · ~$0.0066 est.
  29. 49LitJev19.5I 46 · C 77 · S 67 · K 34 · ~$0.163 est.
  30. 50openJev Verdict 1.419.0I 29 · C 72 · S 78 · K 82 · ~$0.0039 est.
  31. 51kev 0.5B18.9I 31 · C 50 · S 77 · K 76 · ~$0.0063 est.
  32. 52Bespoke Nimble 9B18.7I 46 · C 56 · S 79 · K 33 · ~$0.166 est.
  33. 53GPT-5.6 LunaAPI18.5I 93 · C 87 · S 78 · K 28 · $0.242
  34. 54openJev Verdict18.1I 30 · C 47 · S 77 · K 83 · ~$0.0037 est.
  35. 55Raw Qwen3 1.7B direct logits18.1I 33 · C 24 · S 90 · K 65 · ~$0.015 est.
  36. 56reflex-27b17.8I 46 · C 77 · S 67 · K 32 · ~$0.181 est.
  37. 57djev15.2I 72 · C 88 · S 75 · K 27 · ~$0.274 est.
  38. 58GLiNER2 large15.1I 31 · C 25 · S 62 · K 73 · ~$0.0077 est.
  39. 59OpenJev14.8I 58 · C 58 · S 76 · K 28 · ~$0.255 est.
  40. 60Qwen3.5-0.8B Decision Model14.5I 28 · C 68 · S 49 · K 76 · ~$0.0065 est.
  41. 61Gemini 3.1 Flash-LiteAPI14.3I 54 · C 59 · S 82 · K 27 · $0.264
  42. 62open-jev-deberta-v3-large12.6I 26 · C 67 · S 66 · K 74 · ~$0.0073 est.
  43. 63smalljev semantic-v912.3I 26 · C 59 · S 80 · K 58 · ~$0.025 est.
  44. 64GLiNER211.8I 27 · C 25 · S 72 · K 83 · ~$0.0037 est.
  45. 65Open-Jev 9B11.2I 44 · C 62 · S 72 · K 28 · ~$0.249 est.
  46. 66Open-Jev 2B10.0I 42 · C 55 · S 73 · K 28 · ~$0.249 est.
  47. 67GLiNER2.5 multi9.8I 23 · C 57 · S 68 · K 82 · ~$0.0039 est.
  48. 68SimpleJev7.5I 21 · C 49 · S 57 · K 68 · ~$0.011 est.
  49. 69GLiNER2.5 small7.2I 20 · C 51 · S 78 · K 82 · ~$0.0039 est.
  50. 70Raw Qwen3 0.6B direct logits7.1I 23 · C 21 · S 90 · K 74 · ~$0.0074 est.
  51. 71DeepSeek V4.1 FlashAPI4.8I 94 · C 96 · S 72 · K 17 · $0.594
  52. 72Mirror2.1I 14 · C 26 · S 71 · K 73 · ~$0.0077 est.
  53. 73Mixedbread mxbai-rerank-base-v20.4I 7 · C 84 · S 88 · K 68 · $0.012
  54. 74BAAI bge-reranker-v2-m30.2I 5 · C 84 · S 90 · K 73 · $0.0077
  55. 75Alibaba GTE Reranker ModernBERT-base0.2I 5 · C 79 · S 91 · K 70 · $0.010
  56. 76Certo v10.0I 0 · C 83 · S 94 · K 100 · ~$0.0010 est.
  57. 77Open Jev JSON Canvas0.0I 48 · C 0 · S 84 · K 46 · ~$0.065 est.
  58. classifier.dev (honorable mention)API70.8I 52 · C 72 · S 88 · K 84 · ~$0.0033 est.
  59. Qwen3.8 27B (partial run)API0.0I 40 · C 94 · S 61 · K 0 · ~$2.669 est.
  60. swanOne (partial run)—I – · C – · S – · K – · ~$0.111 est.
  61. Needle 3, options as tools (partial run)—I – · C – · S – · K – · ~$0.014 est.
  62. Needle 3 (partial run)—I – · C – · S – · K – · ~$0.024 est.
020406080100

Score = 4 / (1/I + 1/C + 1/S + 1/K) (each 0–100; × (axis / 50)² for Intelligence, Speed or Cost below 50)

  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Instruction model, JSON schema
  • Small tool-calling model
  • Service built on Jev
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • Shown, not ranked
I, C, S, K = Intelligence, Calibration, Speed, Cost; ~ est. = estimated cost; ann. = announced price; API = the operator's endpoint saw sealed item text, without answers. Names link to each project.

Axes, accuracy, latency and cost

Every system with its four axes, public and sealed accuracy and the gap between them. On a phone the name column stays put while the table scrolls sideways. † = a note on that system — tap it to read.

#SystemJevBench ScoreIntelligenceCalibrationSpeedCost axisPublic accuracy
534
Sealed accuracy
308
Public − sealed gapCost / 1,000p50 latencyEndpoint
163.353.176.383.352.086.6%36.7%+49.9 pp$0.0400.65 sAPI
2
JevK5 v0.2.0†Author says no JevBench items or outputs were used for training, tuning, or selection; public results are reported. Scan found only one generic instruction shared by 8 public hard items; unreleased teacher/replay corpora were unavailable.by allebee
62.048.974.591.159.585.3%33.1%+52.2 pp~$0.022est.—unknown
3
Hopper†Author discloses heavy public-benchmark-directed development (26 model/prompt configs, 20+ calibration-map variants observed against the public half); scan of 17 released files vs 231 public tasks found 0 state/instruction matches; training corpus not released so overlap not independently verifiable.by HopitAI
59.448.079.186.858.782.3%34.1%+48.2 pp~$0.024est.0.13 sRunPod GPU
4
Winnow-12B Q8†The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.by Eldan Ring
55.648.364.882.352.985.7%33.1%+52.6 pp~$0.037est.0.23 sRunPod GPU
5
reflex 4B†The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.by kshetrajna12
54.047.570.468.059.779.2%28.2%+51.0 pp~$0.022est.1.80 sRunPod GPU
6
djev†The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated). v1.4: hosted api.djev.dev was paused by its operator ("Serving is paused by the administrator"); sealed tier measured on the public djev runtime (Davipar/djev-dev 3ce907e, same weights, default mode) self-hosted on an H100; Speed/Cost kept from v1.3.by Maisa (David Villalón) · Maisa, diffusion-gemma
52.247.055.491.457.684.0%29.9%+54.1 pp$0.026announced0.24 sAPI
7
Jev-Omni†Akhilaaa3/Jev-Omni revision c050d51354147985d13286cf4acf90f562f2c631, the author's own load_model.py (merged text decision model + 256-way head) and his own predict(), transformers 5.17.0 / torch 2.8.0 from the pod image; built on the CPU and moved to CUDA with every nn.Linear weight cast to bfloat16 first - the same cast his reference loader jev_omni.py applies - because our 46 GB GPU cannot hold his fp32 copy; on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.by akhilaaa3 · akhilaaa3, Gemma-4-12B merged
51.346.864.181.553.088.7%32.1%+56.6 pp~$0.037est.0.22 sRunPod GPU
8
metask-jev-4b†Model card discloses 44.8k+16.1k+390 training rows incl. synthetic families intentionally mirroring JevBench hard families, plus repeated evaluation on all 231 public items (benchmark-directed development disclosed). Scan of 61 files: 0 exact matches.by Wayfind (metask-ai)
47.844.766.989.154.579.7%27.6%+52.1 pp~$0.033est.0.07 sRunPod GPU
9
SemIfby Theodore Lee (TheoLeeCJ) · formerly OpenJev (Qwen3.5-4B, TheoLeeCJ
47.744.466.883.759.581.0%26.3%+54.7 pp~$0.022est.0.20 sRunPod GPU
10
Jobe Qwen3.5-4B†No trained weights/LoRA/calibration fit; release explicitly rejects fitted temperature/order averaging. Scan of 37 files vs 231 public tasks: 0 matches.by MantisShrimpdev · frozen
46.944.166.185.659.581.0%25.6%+55.3 pp~$0.022est.0.13 sRunPod GPU
11
local-jev Qwen3.5-4B†Scan of 72 files: 0 matches. Author discloses choosing JSON layout after observing results on the 231 public items (benchmark-directed choice, disclosed).by Amith Chandrappa (amithgc)
46.844.473.375.055.880.5%26.0%+54.5 pp~$0.030est.0.71 sRunPod GPU
12
system-one-openAPIby mithalouni · Gemma 4 E2B LoRA on an L4
45.144.254.977.064.873.2%27.6%+45.6 pp~$0.015est.0.65 sauthor demo
13
spark-s1-4b-v6†Abhishek085/spark-s1-4b-v6 revision 93d49ddbfb29212e3296635a75a3e80cf69da027, code github.com/abhishek085/open-spark-jev 30ac6d89b7fa36c644cf86aac68f35c1d276a919, the author's own MenuScorer.decide with his fitted calibration.json temperature, bf16, base Qwen/Qwen3.5-4B, transformers 5.17.0 / torch 2.8.0 from the pod image, flash-linear-attention 0.5.2 installed, causal_conv1d not installable here (no wheel builds against this toolchain), on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.by Abhishek Rai (abhishek085) · Open Spark Jev, abhishek085
44.645.147.681.057.979.2%26.6%+52.6 pp~$0.025est.0.31 sRunPod GPU
14
jqv†A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.by hjmurmur (Octalab) · Qwen3-32B zero-shot
44.446.471.674.647.580.1%28.2%+51.8 pp~$0.056est.0.75 sRunPod GPU
15
Qwen3-Reranker-4B†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.by Qwen
43.544.665.278.749.268.0%29.9%+38.1 pp$0.0500.13 sRunPod GPU
16
decider-35b-a3b†The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.by Mapika
41.247.265.380.845.383.1%31.5%+51.6 pp~$0.067est.0.29 sRunPod GPU
1741.046.429.187.659.769.7%27.3%+42.4 pp~$0.022est.0.08 sRunPod GPU
18
OpenSourceJev†DM submission, measure-only (JevBench publishing HOLD in force). Same author as the existing simplejev-qwen3.5-0.8b row (sabeel111/Featherless AI). Round-4 audit of an earlier commit could not be measured (no public GGUF, Windows-only DLL loader, unpinned llama.cpp build); this round the author published the exact unsloth Q4_K_M GGUF (hash/size independently verified) and we built llama.cpp CUDA from current upstream master on Linux ourselves -- its ABI matched the ctypes bindings exactly, so only a loader file-naming fix was needed (documented diff), no code/scoring/calibration change. Our public-231 subset exactly reproduced the author-reported table: easy 48/48, standard 67/72, hard 66/111, schema 231/231. Calibration (noul temperature) fit only on Google BoolQ, not JevBench. Repo docs name 3 public task IDs while describing benchmark-directed algorithm fixes on the public half (disclosed); 0 exact state/instruction text matches in a released-file scan.by sabeel111 · Qwen3.5-4B Q4_K_M, native llama.cpp
40.941.860.382.064.078.4%26.3%+52.1 pp~$0.016est.—unknown
19
ZeroEntropy zerank-2†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.by ZeroEntropy
40.242.175.879.049.870.1%28.6%+41.6 pp$0.0470.13 sRunPod GPU
20
decision-machine-1†A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.APIby milliseconds.ai (Baptiste Laget)
39.941.368.392.953.767.5%25.6%+41.9 pp$0.0350.17 sAPI
21
Raw Phi-4 mini direct logits†Neutral raw-logit control, not JevBench-directed.by Microsoft / neutral reproduction
38.041.858.888.849.665.8%29.2%+36.6 pp~$0.048est.0.06 sRunPod GPU
22
JEV Qwen3.5-9B Base NVFP4†Byte-identical to upstream March-2026 NVFP4 checkpoint, predates JevBench v1.2, no task-specific training added. Had 3 preliminary scope/policy scoring errors corrected during review (pooled-ECE, full-run latency, cost estimate); score above is final corrected value.by WilfLin
37.746.867.793.343.375.3%29.5%+45.8 pp~$0.077est.0.02 sRunPod GPU
23
OpenJevby razorback16 / Codiv · DiffusionGemma 26B-A4B NVFP4, razorback16
36.945.455.083.245.581.8%28.6%+53.2 pp~$0.066est.0.24 sRunPod GPU
24
kev 4B†Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 306/308 sealed items answered validly (failures count as wrong)by Jared Palmer · research preview
36.142.139.675.761.866.2%22.4%+43.8 pp~$0.019est.0.55 sRunPod GPU
25
Decision 2B†Flymy-ai/decision-2b-preview revision df57b75db927acc9ad91ec8115508c1e487086eb (checkpoint minicpm5_reduced_v16_4k_v59), base openbmb/MiniCPM5-2B revision 12a3808a956f869c767195e9266b59c4d21d92e2, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.by FlyMy.AI (@denti) · FlyMy.AI, v59
35.838.874.184.362.575.3%26.0%+49.4 pp~$0.018est.0.19 sRunPod GPU
26
Qwen3.5-9B Jev-like data-mix v2†The author disclosed development on the public JevBench set and public-result comparisons; released-data overlap scan found no matches, but 764 gap and 382 replay training rows are unreleased.by jsaurabh
35.247.461.382.042.478.4%29.2%+49.1 pp~$0.083est.—unknown
27
GPT-6 Luna†OpenAI direct API baseline; reasoning effort low; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.APIby OpenAI · low reasoning effort
35.195.892.073.736.999.1%92.9%+6.3 pp$0.1271.44 sAPI
28
SimpleJev Qwen3.8-27B†Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.APIby Featherless AI
34.651.674.571.239.586.6%35.7%+50.9 pp~$0.104est.1.01 sauthor demo
29
NInfer Qwen3.8-Flash-Next mixed†Engine scan (1,758 files) vs 231 public tasks: 0 exact matches. Submitter discloses no training for Flash-Next but repeated consultation of public items and a public-hard temperature sweep.by Igor L. / NInfer contributors
34.049.578.688.238.989.6%34.1%+55.5 pp~$0.109est.0.08 sRunPod GPU
30
GPT-6 Luna†OpenAI direct API baseline; reasoning effort default medium; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.APIby OpenAI · default medium reasoning effort
33.397.493.572.636.099.6%95.5%+4.1 pp$0.1351.48 sAPI
31
open-alternative-jev†With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.by IkerMoel · Qwen3.5-4B, IkerMoel
33.238.658.783.559.674.0%24.4%+49.7 pp~$0.022est.0.21 sRunPod GPU
32
jev-local†The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.by us (GitHub) · Qwen3.5-9B
32.545.264.269.243.374.9%29.5%+45.3 pp~$0.077est.1.05 sRunPod GPU
33
Decision Fast†Flymy-ai/decision-fast-preview revision 4225d41c66119fe28e95a2631bb0103decae6d56 (checkpoint qwen3_06b_headfirst_ep2a_v53), base Qwen/Qwen3-0.6B-Base revision da87bfb608c14b7cf20ba1ce41287e8de496c0cd, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.by FlyMy.AI (@denti) · FlyMy.AI, v53a
32.537.165.381.676.163.2%25.6%+37.6 pp~$0.0063est.0.24 sRunPod GPU
34
decider-2b†The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.by Mapika
30.738.543.583.261.071.0%24.7%+46.3 pp~$0.020est.0.26 sRunPod GPU
35
jeff†Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.by Logan Markewich · Logan Markewich, GLiFormer 400M
30.636.867.963.576.662.8%33.1%+29.7 pp~$0.0060est.0.94 sCPU
36
Laya†The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.by Convai Innovations · Convai Innovations, ModernBERT-large 421M
30.336.163.771.186.258.4%30.8%+27.6 pp~$0.0029est.0.79 sCPU
37
lev-350m†Weights franckverrot/lev-350m revision ab08ad8b8f346994d983152917e114224f6adac7, code github.com/franckverrot/lev c48a945dbf629998d7458dcc5c16f58df964db94, the author's own lev.serve /v1/systemone endpoint with its shipped calibration temperature, base LiquidAI/LFM2.5-350M, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.by Franck Verrot (franckverrot) · Franck Verrot, LFM2.5-350M
28.534.870.685.376.158.4%25.0%+33.4 pp~$0.0063est.0.17 sRunPod GPU
38
openjev-sglangAPIby ekzhang · Qwen3.6-35B-A3B on SGLang
27.749.469.277.136.585.3%33.1%+52.2 pp~$0.131est.0.68 sauthor demo
39
Von†The author disclosed that its temperature calibration map used the 231 public JevBench items; monotonic scaling does not change accuracy. Estimated cost is USD 0.00551 per 1,000 decisions.by wfzyx (Victor Hugo) · wfzyx, Option-Marker 395M
27.534.575.770.577.857.1%27.9%+29.2 pp~$0.0055est.—unknown
40
NInfer Qwen3.8-27B NVFP4†T=1.5 is an offline recomputation from raw logits of the same run, not a second execution. T=1.5 was chosen by sweeping public hard items (development-set calibrated, disclosed).by Igor L. / NInfer contributors · T=1.5
26.951.576.080.135.283.1%33.1%+50.0 pp~$0.145est.0.37 sRunPod GPU
41
NInfer Qwen3.8-27B NVFP4†Raw T=1.0 row. Author discloses repeated public-item consultation and public-hard tuning (applies to both NInfer 27B rows).by Igor L. / NInfer contributors
26.351.567.280.135.283.1%33.1%+50.0 pp~$0.145est.0.37 sRunPod GPU
42
kev 8B†Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.by Jared Palmer · research preview
25.641.840.274.944.071.4%21.8%+49.7 pp~$0.073est.0.59 sRunPod GPU
43
JevOne†Scan vs 231 public tasks: 0 matches. Training corpus/provenance not disclosed — overlap unknown.by Juspay
25.547.677.688.535.989.6%33.8%+55.8 pp~$0.137est.0.09 sRunPod GPU
44
SimpleJev Qwen3.6-35B-A3B†Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.APIby Featherless AI
24.945.759.875.038.181.4%28.2%+53.1 pp~$0.116est.0.85 sauthor demo
45
kev 0.6B†Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 307/308 sealed items answered validly (failures count as wrong)by Jared Palmer · research preview
24.834.250.075.676.166.7%24.0%+42.6 pp~$0.0063est.0.59 sRunPod GPU
4623.745.724.186.341.968.4%26.3%+42.1 pp~$0.087est.0.08 sRunPod GPU
47
system-oneby Sean Goedecke · Qwen3-8B, Sean Goedecke
23.443.632.884.441.571.9%24.4%+47.5 pp~$0.089est.0.17 sRunPod GPU
48
OpenDecision†A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.by Deepan Wadhwa · ModernBERT-large zero-shot
21.631.857.179.975.353.2%25.6%+27.6 pp~$0.0066est.0.34 sRunPod GPU
49
LitJev†The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.by Zhengxu Yu · Qwen3.8-27B
19.546.376.666.733.686.1%30.8%+55.3 pp~$0.163est.2.03 sRunPod GPU
50
openJev Verdict 1.4†Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.by Hemant (heman10x)
19.029.472.078.182.457.6%27.9%+29.7 pp~$0.0039est.0.31 sCPU
51
kev 0.5B†Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release. 307/308 sealed items answered validly (failures count as wrong)by Jared Palmer
18.930.549.777.076.149.4%27.3%+22.1 pp~$0.0063est.0.43 sRunPod GPU
52
Bespoke Nimble 9B†Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.by Bespoke Labs
18.746.356.478.733.479.7%28.9%+50.8 pp~$0.166est.0.39 sRunPod GPU
53
GPT-5.6 LunaAPIby OpenAI · low reasoning effort
18.593.187.477.528.597.4%89.0%+8.4 pp$0.2420.97 sAPI
54
openJev Verdict†The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.by Hemant (heman10x) · heman10x, ModernBERT-base 151M
18.130.047.076.783.155.4%24.7%+30.7 pp~$0.0037est.0.28 sCPU
5518.133.224.389.764.954.1%26.0%+28.1 pp~$0.015est.0.07 sRunPod GPU
56
reflex-27b†The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.by kshetrajna12 · Qwen3.8-27B
17.846.477.267.532.387.0%29.5%+57.5 pp~$0.181est.1.89 sRunPod GPU
57
djev†Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill. 220/308 sealed items answered validly (failures count as wrong)by David Villalon / Maisa · thinking
15.271.687.875.226.987.4%60.1%+27.4 pp~$0.274est.0.43 sRunPod GPU
58
GLiNER2 large†The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.by Fastino
15.131.124.861.773.356.7%28.6%+28.1 pp~$0.0077est.1.10 sCPU
59
OpenJev†OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.by razorback16 · thinking, BF16
14.858.158.176.127.888.7%42.2%+46.5 pp~$0.255est.0.46 sRunPod GPU
60
Qwen3.5-0.8B Decision Model†JevLite SystemOne on CPU; bundled per-question calibration; no operator endpoint or network access. Local CPU measurement on the existing 534-decision v1.3 set plus 308 sealed v1.4 decisions; no operator endpoint received sealed text.by Mourad Ghafiri
14.528.168.249.275.759.3%34.7%+24.6 pp~$0.0065est.7.15 sCPU
6114.354.559.381.827.487.0%38.6%+48.4 pp$0.2640.76 sAPI
62
open-jev-deberta-v3-large†297/308 sealed items answered validly (failures count as wrong)by Kotoba Labs · local CPU
12.625.666.666.074.052.4%29.5%+22.8 pp~$0.0073est.1.77 sCPU
63
smalljev semantic-v9†The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.by Aditya (isHeSatoshi)
12.325.759.279.857.960.6%26.9%+33.7 pp~$0.025est.0.41 sRunPod GPU
64
GLiNER2†A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).by Fastino · Fastino, gliner2.5-base
11.827.425.271.883.158.0%29.2%+28.8 pp~$0.0037est.0.31 sCPU
65
Open-Jev 9B†The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.by Zefan Cai (@Zefan_Cai)
11.244.261.872.028.177.5%29.9%+47.6 pp~$0.249est.0.75 sRunPod GPU
66
Open-Jev 2B†The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.by Zefan Cai (@Zefan_Cai)
10.042.355.373.528.164.5%26.3%+38.2 pp~$0.249est.0.66 sRunPod GPU
67
GLiNER2.5 multi†The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.by Fastino · Fastino, 287M
9.823.157.267.882.448.9%32.8%+16.1 pp~$0.0039est.0.43 sCPU
68
SimpleJev†143 pinned files scanned: 0 exact matches. No JevBench-specific fine-tuning.by sabeel111 / Featherless AI · Qwen3.5-0.8B, CPU
7.521.549.157.568.354.5%34.7%+19.8 pp~$0.011est.3.99 sCPU
69
GLiNER2.5 small†The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.by Fastino · Fastino, 74M
7.220.550.777.882.445.9%28.6%+17.3 pp~$0.0039est.0.11 sCPU
707.122.820.789.973.948.1%25.6%+22.4 pp~$0.0074est.0.07 sRunPod GPU
71
DeepSeek V4.1 FlashAPIby DeepSeek · thinking default
4.894.095.571.616.897.8%94.8%+3.0 pp$0.5941.42 sAPI
72
Mirror†171 intended HTTP 422 context rejections (over 512-token limit) counted once each as misses; only 363/534 valid distributions returned. Found via Gmail submission (Lewis). 102/308 sealed items answered validly (failures count as wrong)by Bluusun
2.113.626.070.873.341.1%9.1%+32.0 pp~$0.0077est.0.90 sauthor demo
73
Mixedbread mxbai-rerank-base-v2†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.by Mixedbread
0.46.884.187.567.937.2%34.4%+2.8 pp$0.0120.07 sRunPod GPU
74
BAAI bge-reranker-v2-m3†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.by BAAI
0.25.084.289.573.439.4%27.9%+11.5 pp$0.00770.03 sRunPod GPU
75
Alibaba GTE Reranker ModernBERT-base†Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.by Alibaba-NLP
0.24.878.990.669.633.8%33.4%+0.3 pp$0.0100.05 sRunPod GPU
76
Certo v1†The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.by AltSlate Labs
0.00.183.094.0100.031.6%29.5%+2.1 pp~$0.0010est.0.02 sRunPod GPU
77
Open Jev JSON Canvas†Returns only a final label, not a probability distribution — calibration counts as 0 in the composite. Scan of 39 files: 0 matches.by JoshuaSP
0.048.20.084.145.684.4%31.2%+53.2 pp~$0.065est.0.22 sRunPod GPU
—
classifier.dev†Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.APIby mrmps (@michael_chomsky) · fast tier
honorable mention · not ranked
70.851.672.487.684.385.3%34.4%+50.9 pp~$0.0033est.0.39 sAPI
—
Qwen3.8 27BAPIby Qwen / Chutes · Chutes TEE
partial · not ranked
0.040.493.661.30.071.9%21.8%+50.1 pp~$2.669est.5.75 sAPI
—
swanOne†Draft vocabulary derived via AGPL-3.0 generator (provenance recorded separately). Submitter consulted all 231 public tasks and swept temperature over 111 public hard tasks. Score corrected during review from pooled-534 ECE/latency to hard-tier/242-item block. v1.4: Not re-measured: needs its own H100 NVL pod (99 GiB weights); RunPod balance ran low during this job. Ready-to-run recipe kept.by swanOne submitter
partial · not ranked
—————88.7%——~$0.111est.0.34 sRunPod GPU
—
Needle 3, options as tools†V1.4: Not re-measured: same as needle-3.by Cactus Compute · post-hoc adapter mode
partial · not ranked
—————22.1%——~$0.014est.3.78 sCPU
—
Needle 3†V1.4: Not re-measured: ~100-250 s per item on a rented CPU pod (19 s on Sandy); needs a dedicated CPU host.by Cactus Compute · Cactus, 2-bit, local CPU
partial · not ranked
—————22.5%——~$0.024est.1.69 sCPU

API = the operator's endpoint received sealed item text during evaluation; the answers and item-level results are not published. The sealed text and answers remain private; only system-level aggregates appear here. Cost is per 1,000 decisions. Hover endpoint, cost and API labels for their recorded details.

All 72 system notes and disclosures

  • † JevK5 v0.2.0: Author says no JevBench items or outputs were used for training, tuning, or selection; public results are reported. Scan found only one generic instruction shared by 8 public hard items; unreleased teacher/replay corpora were unavailable.
  • † Hopper: Author discloses heavy public-benchmark-directed development (26 model/prompt configs, 20+ calibration-map variants observed against the public half); scan of 17 released files vs 231 public tasks found 0 state/instruction matches; training corpus not released so overlap not independently verifiable.
  • † Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
  • † reflex 4B (kshetrajna12): The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
  • † djev (Maisa, diffusion-gemma): The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated). v1.4: hosted api.djev.dev was paused by its operator ("Serving is paused by the administrator"); sealed tier measured on the public djev runtime (Davipar/djev-dev 3ce907e, same weights, default mode) self-hosted on an H100; Speed/Cost kept from v1.3.
  • † Jev-Omni (akhilaaa3, Gemma-4-12B merged): Akhilaaa3/Jev-Omni revision c050d51354147985d13286cf4acf90f562f2c631, the author's own load_model.py (merged text decision model + 256-way head) and his own predict(), transformers 5.17.0 / torch 2.8.0 from the pod image; built on the CPU and moved to CUDA with every nn.Linear weight cast to bfloat16 first - the same cast his reference loader jev_omni.py applies - because our 46 GB GPU cannot hold his fp32 copy; on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • † metask-jev-4b: Model card discloses 44.8k+16.1k+390 training rows incl. synthetic families intentionally mirroring JevBench hard families, plus repeated evaluation on all 231 public items (benchmark-directed development disclosed). Scan of 61 files: 0 exact matches.
  • † Jobe Qwen3.5-4B (frozen): No trained weights/LoRA/calibration fit; release explicitly rejects fitted temperature/order averaging. Scan of 37 files vs 231 public tasks: 0 matches.
  • † local-jev Qwen3.5-4B: Scan of 72 files: 0 matches. Author discloses choosing JSON layout after observing results on the 231 public items (benchmark-directed choice, disclosed).
  • † spark-s1-4b-v6 (Open Spark Jev, abhishek085): Abhishek085/spark-s1-4b-v6 revision 93d49ddbfb29212e3296635a75a3e80cf69da027, code github.com/abhishek085/open-spark-jev 30ac6d89b7fa36c644cf86aac68f35c1d276a919, the author's own MenuScorer.decide with his fitted calibration.json temperature, bf16, base Qwen/Qwen3.5-4B, transformers 5.17.0 / torch 2.8.0 from the pod image, flash-linear-attention 0.5.2 installed, causal_conv1d not installable here (no wheel builds against this toolchain), on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • † jqv (Qwen3-32B zero-shot): A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
  • † Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † decider-35b-a3b (Mapika): The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
  • † Raw Qwen3 4B Instruct 2507 direct logits: Neutral raw-logit control.
  • † OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp): DM submission, measure-only (JevBench publishing HOLD in force). Same author as the existing simplejev-qwen3.5-0.8b row (sabeel111/Featherless AI). Round-4 audit of an earlier commit could not be measured (no public GGUF, Windows-only DLL loader, unpinned llama.cpp build); this round the author published the exact unsloth Q4_K_M GGUF (hash/size independently verified) and we built llama.cpp CUDA from current upstream master on Linux ourselves -- its ABI matched the ctypes bindings exactly, so only a loader file-naming fix was needed (documented diff), no code/scoring/calibration change. Our public-231 subset exactly reproduced the author-reported table: easy 48/48, standard 67/72, hard 66/111, schema 231/231. Calibration (noul temperature) fit only on Google BoolQ, not JevBench. Repo docs name 3 public task IDs while describing benchmark-directed algorithm fixes on the public half (disclosed); 0 exact state/instruction text matches in a released-file scan.
  • † ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † decision-machine-1 (milliseconds.ai): A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
  • † Raw Phi-4 mini direct logits: Neutral raw-logit control, not JevBench-directed.
  • † JEV Qwen3.5-9B Base NVFP4: Byte-identical to upstream March-2026 NVFP4 checkpoint, predates JevBench v1.2, no task-specific training added. Had 3 preliminary scope/policy scoring errors corrected during review (pooled-ECE, full-run latency, cost estimate); score above is final corrected value.
  • † kev 4B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 306/308 sealed items answered validly (failures count as wrong)
  • † Decision 2B (FlyMy.AI, v59): Flymy-ai/decision-2b-preview revision df57b75db927acc9ad91ec8115508c1e487086eb (checkpoint minicpm5_reduced_v16_4k_v59), base openbmb/MiniCPM5-2B revision 12a3808a956f869c767195e9266b59c4d21d92e2, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • † Qwen3.5-9B Jev-like data-mix v2: The author disclosed development on the public JevBench set and public-result comparisons; released-data overlap scan found no matches, but 764 gap and 382 replay training rows are unreleased.
  • † GPT-6 Luna (low reasoning effort): OpenAI direct API baseline; reasoning effort low; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
  • † SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • † NInfer Qwen3.8-Flash-Next mixed: Engine scan (1,758 files) vs 231 public tasks: 0 exact matches. Submitter discloses no training for Flash-Next but repeated consultation of public items and a public-hard temperature sweep.
  • † GPT-6 Luna (default medium reasoning effort): OpenAI direct API baseline; reasoning effort default medium; strict JSON-schema probability response; temperature unset; max_completion_tokens=4096; price cost from returned usage at official standard list rates.
  • † open-alternative-jev (Qwen3.5-4B, IkerMoel): With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
  • † jev-local (Qwen3.5-9B): The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
  • † Decision Fast (FlyMy.AI, v53a): Flymy-ai/decision-fast-preview revision 4225d41c66119fe28e95a2631bb0103decae6d56 (checkpoint qwen3_06b_headfirst_ep2a_v53), base Qwen/Qwen3-0.6B-Base revision da87bfb608c14b7cf20ba1ce41287e8de496c0cd, the submitter's own FlyMyJevPackageAdapter and frozen calibrator, bf16, unmerged adapter, 4096-token packing, transformers 4.57.6 / peft 0.15.2 as pinned, torch 2.8.0 from the pod image, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • † decider-2b (Mapika): The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
  • † jeff (Logan Markewich, GLiFormer 400M): Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
  • † Laya (Convai Innovations, ModernBERT-large 421M): The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
  • † lev-350m (Franck Verrot, LFM2.5-350M): Weights franckverrot/lev-350m revision ab08ad8b8f346994d983152917e114224f6adac7, code github.com/franckverrot/lev c48a945dbf629998d7458dcc5c16f58df964db94, the author's own lev.serve /v1/systemone endpoint with its shipped calibration temperature, base LiquidAI/LFM2.5-350M, on our RunPod L40 in Czechia Offline local open-weight inference on the frozen 308-item v1.4 set in a network-disabled, read-only container; no operator endpoint received sealed text.
  • † Von (wfzyx, Option-Marker 395M): The author disclosed that its temperature calibration map used the 231 public JevBench items; monotonic scaling does not change accuracy. Estimated cost is USD 0.00551 per 1,000 decisions.
  • † NInfer Qwen3.8-27B NVFP4 (T=1.5): T=1.5 is an offline recomputation from raw logits of the same run, not a second execution. T=1.5 was chosen by sweeping public hard items (development-set calibrated, disclosed).
  • † NInfer Qwen3.8-27B NVFP4: Raw T=1.0 row. Author discloses repeated public-item consultation and public-hard tuning (applies to both NInfer 27B rows).
  • † kev 8B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • † JevOne: Scan vs 231 public tasks: 0 matches. Training corpus/provenance not disclosed — overlap unknown.
  • † SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • † kev 0.6B (research preview): Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview. 307/308 sealed items answered validly (failures count as wrong)
  • † Raw Qwen3 8B direct logits: Neutral raw-logit control.
  • † OpenDecision (ModernBERT-large zero-shot): A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
  • † LitJev (Qwen3.8-27B): The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
  • † openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
  • † kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release. 307/308 sealed items answered validly (failures count as wrong)
  • † Bespoke Nimble 9B (Bespoke Labs): Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
  • † openJev Verdict (heman10x, ModernBERT-base 151M): The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
  • † Raw Qwen3 1.7B direct logits: Neutral raw-logit control.
  • † reflex-27b (Qwen3.8-27B): The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
  • † djev (thinking): Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill. 220/308 sealed items answered validly (failures count as wrong)
  • † GLiNER2 large (Fastino): The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • † OpenJev (thinking, BF16): OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
  • † Qwen3.5-0.8B Decision Model (Mourad Ghafiri): JevLite SystemOne on CPU; bundled per-question calibration; no operator endpoint or network access. Local CPU measurement on the existing 534-decision v1.3 set plus 308 sealed v1.4 decisions; no operator endpoint received sealed text.
  • † open-jev-deberta-v3-large (local CPU): 297/308 sealed items answered validly (failures count as wrong)
  • † smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
  • † GLiNER2 (Fastino, gliner2.5-base): A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
  • † Open-Jev 9B (Zefan Cai): The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • † Open-Jev 2B (Zefan Cai): The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • † GLiNER2.5 multi (Fastino, 287M): The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
  • † SimpleJev (Qwen3.5-0.8B, CPU): 143 pinned files scanned: 0 exact matches. No JevBench-specific fine-tuning.
  • † GLiNER2.5 small (Fastino, 74M): The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • † Raw Qwen3 0.6B direct logits: Neutral raw-logit control.
  • † Mirror: 171 intended HTTP 422 context rejections (over 512-token limit) counted once each as misses; only 363/534 valid distributions returned. Found via Gmail submission (Lewis). 102/308 sealed items answered validly (failures count as wrong)
  • † Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † Certo v1 (AltSlate Labs): The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
  • † Open Jev JSON Canvas (JoshuaSP): Returns only a final label, not a probability distribution — calibration counts as 0 in the composite. Scan of 39 files: 0 matches.
  • † classifier.dev (fast tier): Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.
  • † swanOne: Draft vocabulary derived via AGPL-3.0 generator (provenance recorded separately). Submitter consulted all 231 public tasks and swept temperature over 111 public hard tasks. Score corrected during review from pooled-534 ECE/latency to hard-tier/242-item block. v1.4: Not re-measured: needs its own H100 NVL pod (99 GiB weights); RunPod balance ran low during this job. Ready-to-run recipe kept.
  • † Needle 3, options as tools (post-hoc adapter mode): V1.4: Not re-measured: same as needle-3.
  • † Needle 3 (Cactus, 2-bit, local CPU): V1.4: Not re-measured: ~100-250 s per item on a rented CPU pod (19 s on Sandy); needs a dedicated CPU host.

Rows without a † have no note beyond the shared provenance: every row was measured or re-run with its recorded recipe, and deviations are in its run manifest.

Artifact: v1.4.1 results JSON · SHA-256 808909c4ecc8… · JevBench v1.4.1 release and method

Compare two systems

Pick any two. Four radars: the score axes, accuracy per tier including the sealed set, and accuracy by family on the v1.2 hard tier and on the sealed set. Further out is better on every spoke; the link keeps the pair.

  • A: Jev 1.13.0 — Jev (TypeSafe, closed) · Score 63.3 (#1)
  • B: JevK5 v0.2.0 — Jev rebuild · Score 62.0 (#2)
媒体内容 · 前往原文查看

The four score axes

0–100, the values in the table. A label-only system has no calibration (counted as 0).
媒体内容 · 前往原文查看

Accuracy per tier, incl. sealed

Share correct per tier; Sealed = the 308 private decisions, aggregate only.
媒体内容 · 前往原文查看

Hard tier by family (v1.2 topics)

Share correct within each family of the 220 v1.2 hard-tier decisions (public and held-out).
媒体内容 · 前往原文查看

Sealed set by family

Share correct within each sealed family — system-level aggregates; the items stay private.

All values as a table

SpokeA: Jev 1.13.0B: JevK5 v0.2.0
The four score axes
Intelligence53.148.9
Calibration76.374.5
Speed83.391.1
Cost52.059.5
Accuracy per tier, incl. sealed
Easy100%100%
Standard99%96%
Judge95%95%
Hard74%70%
Sealed37%33%
Hard tier by family (v1.2 topics)
Adversarial100%92%
Ambiguous79%79%
Judge79%82%
Long policy61%50%
Multi-hop86%77%
Probability80%60%
Routing100%100%
Temporal / numeric27%37%
Trade-off92%83%
Trap100%100%
Sealed set by family
Ambiguous / abstain30%43%
Judge34%37%
Long policy28%23%
Multi-hop45%39%
Paraphrase64%14%
Probability50%39%
Safety judge38%38%
Temporal / numeric29%27%
Trade-off38%23%
Trap / adversarial42%58%

What changed in v1.4

  • Fresh sealed decisions keep the benchmark moving as public items saturate. Sealed items contribute 20% of Intelligence: I = 0.8 × I_v1.3 + 0.2 × I_sealed, where I_sealed = 100 × max(0, (acc_sealed − 0.293) / (1 − 0.293)). Public and sealed scores are published only as aggregates.
  • Calibration blends toward the sealed-inclusive result at the approved weight: C = C_v1.3 + (C_v1.4 − C_v1.3) × min(1, 0.2 / 0.35).
  • The k = 1 generalization penalty reduces Intelligence when public accuracy exceeds sealed accuracy by more than 25 percentage points: I × (1 − max(0, gap − 25) / 100). It rewards systems that generalize beyond the public half.
  • The four axes use an equal-weight harmonic mean (p = −1). Intelligence below 50 keeps its quadratic penalty; Speed and Cost each have a separate Jev-class gate below 50. Speed and Cost axis calculations are unchanged from v1.3.0.
  • The visible API flag discloses when an operator endpoint received held-out item text, without answers. Existing system notes preserve disclosures such as Hopper's public-half development and JevK5's public-set selection.

What the run says

  • Jev 1.13.0 leads v1.4.1 with 63.3: Intelligence 53.1, Calibration 76.3, Speed 83.3, Cost 52.0 ($0.040 per 1,000 decisions).
  • The best open or open-planned rebuild, JevK5 v0.2.0, is #2 at 62.0 — 1.3 points behind.
  • GPT-6 Luna has the highest Intelligence (97.4) but places #30: Speed 72.6, Cost 36.0 — the harmonic mean does not let accuracy buy back a weak axis.
  • The sealed set is hard for everyone: the best sealed accuracy among ranked systems is 95.5% (GPT-6 Luna, #30); chance is 29.3%. Large public-minus-sealed gaps above 25 points reduce Intelligence.
  • classifier.dev scores 70.8 but is not ranked: it is a service running another entrant's model.
  • swanOne, Qwen3.8 27B, Needle 3, options as tools, Needle 3 did not complete every tier; they are listed without a rank.

Jev alternatives, open source and self-hosting

The chart and table above compare the tested systems, not marketing claims. These are the practical answers readers most often need before choosing a Jev-class decision model.

What are open-source alternatives to Jev?

The highest-ranked open entrants in this run are Hopper (#3, 59.4), Winnow-12B Q8 (#4, 55.6), reflex 4B (#5, 54.0), djev (#6, 52.2). “Open” here means the tested row publishes code or weights; check the licence and exact configuration in the board before adopting one.

Which Jev-class models can I self-host in the EU or use for GDPR-sensitive work?

Open entrants with released code or weights can run on infrastructure you choose, including EU infrastructure. That can support data residency, but neither open source nor an EU server makes a deployment GDPR-compliant by itself. Assess your data, contracts, retention, subprocessors and security for the complete setup. See Benchmark Heaven's broader EU-hosting comparison.

jev-router.com offers self-hosted open decision models. Neutrality disclosure: it is run by the authors of this benchmark; it receives no scoring advantage and is not a ranked entrant.

How is JevBench scored?

The official score is the equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost. Version v1.4.1 measures 534 public and 308 sealed decisions: 20% of Intelligence comes from the sealed set, a public-minus-sealed gap above 25 points costs Intelligence, and Intelligence, Speed or Cost below 50 each pull the score down quadratically. What changed in v1.4 · method and tiers.

How do I submit my model?

Open an issue in the JevBench repository with a reproducible endpoint or runnable code, the exact model and licence, and whether public JevBench items were used during development. New entrants use the same frozen harness and appear in a new version. For private data, see the custom evaluation options.

What a decision costs

Every price here is US dollars per 1,000 decisions — not per 1,000 tokens. One decision is a whole question — state, rubric and options — about 950 input tokens for Jev 1.13.0, so at its $0.042 per million input tokens 1,000 decisions cost $0.0399.

How costs are estimated

One decision is a whole question, not a token. Jev 1.13.0 reads 950 input tokens per decision on average over the 534 v1.2 decisions. At its public tariff of $0.042 per MILLION input tokens (output tokens are free, https://docs.typesafe.ai/models), 1,000 decisions therefore cost 950 x 1,000 x $0.042 / 1,000,000 = $0.0399. That is what the Cost column shows: $0.0399 per 1,000 decisions, not per 1,000 tokens.

Systems with a public tariff (per token or per request) are priced at that tariff times the tokens we measured. Systems without one — open weights, author demos, models we ran locally — are priced as if a large inference provider hosted them: the OpenRouter list price of the same weights; if OpenRouter does not list them, the nearest larger sibling; if no model of that size class is on OpenRouter, the DeepInfra list price of the same weights or of the nearest larger model of the same class. We do not use per-minute GPU rental or our own CPU time — providers buy capacity in bulk or own the hardware, and price accordingly. Price × tokens per decision = $ per 1,000 decisions, marked “est.”.

  • JevK5 v0.2.0 — ~$0.022 est. per 1,000 decisions: Reconstructed from the frozen v1.3 Cost axis; same speed/cost measurement, not a new price observation
  • Hopper — ~$0.024 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • Winnow-12B Q8 — ~$0.037 est. per 1,000 decisions: OpenRouter google/gemma-3-12b-it hosted reference list price $0.05/M in, $0.0/M out (the nearest publicly hosted 12B Gemma sibling; Winnow reads answer logits in one forward pass and generates no answer tokens) x 393 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • reflex 4B — ~$0.022 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.0/M out (the exact base weights; one pass, no generated output) x 377 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Jev-Omni — ~$0.037 est. per 1,000 decisions: OpenRouter Gemma 3 12B input rate list price $0.05/M in, $0.0/M out (a 12B one-pass model with no generated output; the same reference the author uses in his own model card for this model) x 384 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • metask-jev-4b — ~$0.033 est. per 1,000 decisions: hosted 4B reference rate USD 0.04/M input, USD 0/M output over the exact measured prompt-token counts of all 534 attempts; no generated answer tokens
  • SemIf — ~$0.022 est. per 1,000 decisions: deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (same weights (not on OpenRouter), as open-alternative-jev in v1.1.2) x 396 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1244 in / 0 out tokens per hard decision
  • Jobe Qwen3.5-4B — ~$0.022 est. per 1,000 decisions: DeepInfra Qwen3.5-4B hosted reference list price $0.03/M in, $0.0/M out (same underlying weights; one forward pass, no generated tokens) x 396 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • local-jev Qwen3.5-4B — ~$0.030 est. per 1,000 decisions: hosted 4B reference rate list price $0.04/M in, $0.0/M out (one forward pass over measured input tokens and no generated answer tokens) x 397 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • system-one-open — ~$0.015 est. per 1,000 decisions: deepinfra google/gemma-4-E4B-it list price $0.02/M in, $0.1/M out (Gemma 4 E2B is not listed; the nearest larger sibling, Gemma 4 E4B, is listed only on DeepInfra) x 383 input and 2 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra google/gemma-4-E4B-it $0.02/M in, $0.1/M out x 1235 in / 2 out tokens per hard decision
  • spark-s1-4b-v6 — ~$0.025 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B size-class reference list price $0.03/M in, $0.0/M out (a 2-4B one-pass model with no generated output; the board's 4B open-weights reference, as for reflex 4B and decider-2b) x 505 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • jqv — ~$0.056 est. per 1,000 decisions: OpenRouter qwen/qwen3-32b list price $0.08/M in, $0.0/M out (the exact base model this system reads logits from; nothing is generated) x 359 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • decider-35b-a3b — ~$0.067 est. per 1,000 decisions: OpenRouter Qwen3.6-35B-A3B list price list price $0.1/M in, $0.0/M out (the closest public hosted 35B-A3B direct-logit model; no output is generated) x 312 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Raw Qwen3 4B Instruct 2507 direct logits — ~$0.022 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • OpenSourceJev — ~$0.016 est. per 1,000 decisions: Reconstructed from the frozen v1.3 Cost axis; same speed/cost measurement, not a new price observation
  • Raw Phi-4 mini direct logits — ~$0.048 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • JEV Qwen3.5-9B Base NVFP4 — ~$0.077 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • OpenJev — ~$0.066 est. per 1,000 decisions: openrouter google/gemma-4-26b-a4b-it list price $0.09/M in, $0.3/M out (DiffusionGemma 26B-A4B is not listed; the same-size Gemma 4 26B-A4B MoE sibling is (size class moe_26B-A4B)) x 380 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter google/gemma-4-26b-a4b-it $0.09/M in, $0.3/M out x 1222 in / 0 out tokens per hard decision
  • kev 4B — ~$0.019 est. per 1,000 decisions: DeepInfra Qwen3.5-4B size-class reference list price $0.03/M in, $0.0/M out (a 4B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Decision 2B — ~$0.018 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B size-class reference list price $0.03/M in, $0.0/M out (a 2-4B one-pass model with no generated output; the board's 4B open-weights reference, as for reflex 4B and decider-2b) x 269 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Qwen3.5-9B Jev-like data-mix v2 — ~$0.083 est. per 1,000 decisions: Reconstructed from the frozen v1.3 Cost axis; same speed/cost measurement, not a new price observation
  • SimpleJev Qwen3.8-27B — ~$0.104 est. per 1,000 decisions: OpenRouter Gemma 4 26B-A4B size-class reference list price $0.09/M in, $0.0/M out (a public 27B dense model served as a direct-logit classifier; no output is generated) x 809 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • NInfer Qwen3.8-Flash-Next mixed — ~$0.109 est. per 1,000 decisions: OpenRouter qwen/qwen3.8-flash hosted list reference list price $0.15/M in, $0.0/M out (same underlying Flash-Next weights; native one-pass option-logit readout) x 366 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • open-alternative-jev — ~$0.022 est. per 1,000 decisions: deepinfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.15/M out (as open-alternative-jev) x 383 input and 1 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) | ESTIMATE: deepinfra Qwen/Qwen3.5-4B $0.03/M in, $0.15/M out x 1235 in / 1 out tokens per hard decision
  • jev-local — ~$0.077 est. per 1,000 decisions: OpenRouter qwen/qwen3.5-9b list price $0.1/M in, $0.0/M out (the exact base weights; scored by log-probabilities, nothing is generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • Decision Fast — ~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output, the same reference the kev 0.5B/0.6B rows use) x 280 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • decider-2b — ~$0.020 est. per 1,000 decisions: DeepInfra Qwen/Qwen3.5-4B list price $0.03/M in, $0.0/M out (no hosted ~2B Qwen3.5 is listed, so the 4B price is used and errs high; one pass, no output) x 312 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • jeff — ~$0.0060 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 272 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Laya — ~$0.0029 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 205 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • lev-350m — ~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output, the same reference the kev 0.5B/0.6B rows use) x 280 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • openjev-sglang — ~$0.131 est. per 1,000 decisions: openrouter qwen/qwen3.6-35b-a3b list price $0.1/M in, $0.9/M out (same base weights) x 610 input and 2 output tokens per decision [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3.6-35b-a3b $0.1/M in, $0.9/M out x 2272 in / 2 out tokens per hard decision
  • Von — ~$0.0055 est. per 1,000 decisions: Reconstructed from the frozen v1.3 Cost axis; same speed/cost measurement, not a new price observation
  • NInfer Qwen3.8-27B NVFP4 — ~$0.145 est. per 1,000 decisions: OpenRouter Qwen3.8-27B hosted list reference list price $0.2/M in, $0.0/M out (same underlying weights; native one-pass option-logit readout) x 366 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • NInfer Qwen3.8-27B NVFP4 — ~$0.145 est. per 1,000 decisions: OpenRouter Qwen3.8-27B hosted list reference list price $0.2/M in, $0.0/M out (same underlying weights; native one-pass option-logit readout) x 366 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • kev 8B — ~$0.073 est. per 1,000 decisions: OpenRouter qwen/qwen3-8b list price list price $0.117/M in, $0.0/M out (the same-size Qwen3-8B weights; kev generates no output tokens) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • JevOne — ~$0.137 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • SimpleJev Qwen3.6-35B-A3B — ~$0.116 est. per 1,000 decisions: OpenRouter Qwen3.6-35B-A3B list price list price $0.1/M in, $0.0/M out (the same base weights served as a direct-logit classifier; no output is generated) x 809 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • kev 0.6B — ~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Raw Qwen3 8B direct logits — ~$0.087 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • system-one — ~$0.089 est. per 1,000 decisions: openrouter qwen/qwen3-8b list price $0.117/M in, $0.455/M out (same weights, listed on OpenRouter) x 412 input and 1 output tokens per decision (input tokens measured) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3-8b $0.117/M in, $0.455/M out x 1258 in / 1 out tokens per hard decision
  • OpenDecision — ~$0.0066 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 329 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • LitJev — ~$0.163 est. per 1,000 decisions: OpenRouter Qwen3.8-27B (as the reflex-27b row) list price $0.214/M in, $0.0/M out (the exact base weights; nothing is generated) x 418 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • openJev Verdict 1.4 — ~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • kev 0.5B — ~$0.0063 est. per 1,000 decisions: DeepInfra Qwen3-Embedding-0.6B size-class reference list price $0.01/M in, $0.0/M out (a <=0.6B one-pass model with no generated output) x 279 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Bespoke Nimble 9B — ~$0.166 est. per 1,000 decisions: openrouter qwen/qwen3.5-9b list price $0.1/M in, $0.15/M out (a LoRA merge of Qwen3.5-9B; the base weights are listed on OpenRouter (size class dense_9B), as in the v1.1.3 row) x 970 input and 1 output tokens per decision (input tokens measured (the system's own count))
  • openJev Verdict — ~$0.0037 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]
  • Raw Qwen3 1.7B direct logits — ~$0.015 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • reflex-27b — ~$0.181 est. per 1,000 decisions: OpenRouter Qwen3.8-27B list price list price $0.214/M in, $0.0/M out (the exact public base weights used as a direct-logit classifier; no output is generated) x 481 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • djev — ~$0.274 est. per 1,000 decisions: same-size hosted reference x 749 measured input and 690 measured output tokens per attempted decision across all 534, failures included
  • GLiNER2 large — ~$0.0077 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • OpenJev — ~$0.255 est. per 1,000 decisions: same hosted reference x 1778 billed input and 315 thought output tokens per decision
  • Qwen3.5-0.8B Decision Model — ~$0.0065 est. per 1,000 decisions: same-size DeepInfra Qwen3.5-0.8B reference tariff; measured JevLite tokenizer usage; estimated, not charged
  • open-jev-deberta-v3-large — ~$0.0073 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) $0.01/M in, $0.0/M out x 1235 in / 0 out tokens per hard decision
  • smalljev semantic-v9 — ~$0.025 est. per 1,000 decisions: submitted Qwen/Qwen2.5-3B-Instruct hosted reference list price $0.04/M in, $0.0/M out (the author's documented reference for the same approximate size class; one forward pass, nothing generated) x 329 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • GLiNER2 — ~$0.0037 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 383 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]
  • Open-Jev 9B — ~$0.249 est. per 1,000 decisions: OpenRouter Qwen3.5-9B list price read 2026-09-21 list price $0.1/M in, $0.0/M out (the exact 9B base and a conservative same-family proxy for the unlisted 2B; the decision head generates no output tokens) x 1439 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Open-Jev 2B — ~$0.249 est. per 1,000 decisions: OpenRouter Qwen3.5-9B list price read 2026-09-21 list price $0.1/M in, $0.0/M out (the exact 9B base and a conservative same-family proxy for the unlisted 2B; the decision head generates no output tokens) x 1439 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • GLiNER2.5 multi — ~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • SimpleJev — ~$0.011 est. per 1,000 decisions: Same nonzero hosted size-class reference and exact prompt-token accounting; see RESULT.md
  • GLiNER2.5 small — ~$0.0039 est. per 1,000 decisions: deepinfra base-size encoders (bge-base, e5-base, gte-base, all-mpnet-base) list price $0.005/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 452 input and 0 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts)
  • Raw Qwen3 0.6B direct logits — ~$0.0074 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • Mirror — ~$0.0077 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • Certo v1 — ~$0.0010 est. per 1,000 decisions: deepinfra encoders of the same size (bge-large, e5-large, Qwen3-Embedding-0.6B) list price $0.01/M in, $0.0/M out (an encoder of the same size class; one forward pass, nothing generated) x 86 input and 0 output tokens per decision (input tokens measured (the system's own count))
  • Open Jev JSON Canvas — ~$0.065 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • classifier.dev — ~$0.0033 est. per 1,000 decisions: ESTIMATE from the published paid plan (the free tier was used): classifier.dev Pro $20/month for 200,000 fast classifications a day (https://classifier.dev/pricing, read 2026-09-19) = $0.0033 per 1,000 decisions at full use; one decision = one classification. Lower use costs more per decision: at a tenth of that allowance it is $0.033 per 1,000, and the free tier (20,000 fast classifications a day, which is what this run used) costs nothing.
  • swanOne — ~$0.111 est. per 1,000 decisions: Nonzero hosted-comparable estimate; see RESULT.md
  • Qwen3.8 27B — ~$2.669 est. per 1,000 decisions: openrouter qwen/qwen3.8-27b list price $0.214/M in, $2.55/M out (same weights; our run used a flat-rate Chutes subscription) x 416 input and 393 output tokens per decision [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter qwen/qwen3.8-27b $0.214/M in, $2.55/M out x 1592 in / 1833 out tokens per hard decision
  • Needle 3, options as tools — ~$0.014 est. per 1,000 decisions: same per-token price as Needle 3 (openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out) x 383 input and 20 output tokens per decision, over the 314 easy/standard/judge decisions it ran (no hard-tier run). The v1.2 score lab had no price for this row and scored it 100; fixed. [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json]
  • Needle 3 — ~$0.024 est. per 1,000 decisions: openrouter meta-llama/llama-3.2-1b-instruct list price $0.027/M in, $0.201/M out (no generative model under 1B is listed; the smallest listed one (1B) errs high; about 20 generated tokens for one tool call) x 383 input and 20 output tokens per decision (input tokens counted from the gemini-3.1-flash-lite run, same prompts) [corrected in v1.2.3: the price now averages each of the 314 v1.1 decisions once; see results/v1.2/cost-correction-v1.2.3.json] | ESTIMATE: openrouter meta-llama/llama-3.2-1b-instruct $0.027/M in, $0.201/M out x 1235 in / 20 out tokens per hard decision

Reference prices by size class ($ per million input / output tokens)

  • dense 2-4B: deepinfra Qwen/Qwen3.5-4B $0.03 / $0.15; deepinfra google/gemma-4-E4B-it $0.02 / $0.1; openrouter google/gemma-3-4b-it $0.05 / $0.1; openrouter meta-llama/llama-3.2-3b-instruct $0.05 / $0.33
  • dense 27B: openrouter qwen/qwen3.5-27b $0.195 / $1.56; openrouter qwen/qwen3.6-27b $0.3 / $2; openrouter qwen/qwen3.8-27b $0.214 / $2.55
  • dense 9B: openrouter qwen/qwen3.5-9b $0.1 / $0.15
  • encoder classifier <=0.6B: BAAI/bge-large-en-v1.5 (335M) $0.01; Qwen/Qwen3-Embedding-0.6B $0.01; intfloat/e5-large-v2 (335M) $0.01; intfloat/multilingual-e5-large (560M) $0.01; thenlper/gte-base (110M) $0.005
  • generative <=1B: deepinfra meta-llama/Llama-3.2-1B-Instruct $0.005 / $0.01; openrouter meta-llama/llama-3.2-1b-instruct $0.027 / $0.201
  • moe 26B-A4B: deepinfra google/gemma-4-26B-A4B-it $0.07 / $0.34; openrouter google/gemma-4-26b-a4b-it $0.09 / $0.3
  • moe 35B-A3B: deepinfra Qwen/Qwen3.6-35B-A3B $0.1 / $0.95; openrouter qwen/qwen3.5-35b-a3b $0.1625 / $1.3; openrouter qwen/qwen3.6-35b-a3b $0.1 / $0.9

Sources: OpenRouter https://openrouter.ai/api/v1/models and DeepInfra https://api.deepinfra.com/models/list (both read 2026-09-19).

Correction, v1.2.3 (20 September 2026): every price recomputed, each decision counted once

usd_per_1000_v11_tiers = 1000 x (mean input tokens per decision x $/M in + output tokens charged x $/M out) / 1e6, over all 314 v1.1 decisions (72 easy + 242 standard+judge), each decision counted exactly once and priced exactly once. A metered row uses the provider's own tariff and its own measured token counts, including the requests whose answer could not be parsed; an estimated row uses the reference tariff for its weights or size class and, when the run reports no usage, the input tokens of the gemini-3.1-flash-lite run on the same prompts over the same 314 decisions. usd_per_1000 = (v11 x 314 + hard x 220) / 534.

  • The v1.1 and v1.1.3 aggregations built their cost average from a row list that contained the 242-decision standard+judge run twice (once as the standard tier, once as the judge tier) and the 72 easy decisions once: 556 rows instead of 314. The standard and judge tiers were therefore over-weighted in the price, which made the affected rows look 1.5-3.3 % more expensive than they are.
  • Rows without their own token counts were priced at the input tokens of the gemini-3.1-flash-lite run measured on the 242 standard+judge decisions only (452 per decision) and that figure was applied to all 314 v1.1 decisions, which excludes the shorter easy tier. Over all 314 decisions the same run averages 383.41 input tokens, which is the figure used from v1.2.3 on. This made the affected rows look 4-11 % more expensive.
  • A metered row's price left out the requests whose answer came back unparseable. Those requests returned HTTP 200 with generated tokens and were billed, and JevBench already counts them as wrong answers, so from v1.2.3 they are priced too. Only DeepSeek V4.1 Flash had any (9 of its 314 v1.1 decisions); its price rises by 2.6 %.
  • No tariff was wrong. The hard-tier costs, and classifier.dev's flat plan price, were already correct.

No tariff, measurement, item, answer or rank changed. The prices before and after:

  • Jev 1.13.0 — $0.0406 → $0.0399 (-1.72 %)
  • SemIf — $0.0230 → $0.0224 (-2.31 %)
  • system-one-open — $0.0157 → $0.0149 (-5.13 %)
  • OpenJev — $0.0672 → $0.0656 (-2.36 %)
  • GPT-5.6 Luna — $0.2473 → $0.2419 (-2.17 %)
  • openjev-sglang — $0.1346 → $0.1313 (-2.48 %)
  • Bespoke Nimble 9B — $0.1085 → $0.1049 (-3.28 %)
  • Gemini 3.1 Flash-Lite — $0.2682 → $0.2638 (-1.65 %)
  • DeepSeek V4.1 Flash — $0.5788 → $0.5937 (+2.57 %)
  • system-one — $0.0915 → $0.0894 (-2.29 %)
  • openJev Verdict — $0.0039 → $0.0037 (-5.21 %)
  • GLiNER2 — $0.0039 → $0.0037 (-5.21 %)
  • open-jev-deberta-v3-large — $0.0077 → $0.0073 (-5.20 %)
  • Qwen3.8 27B — $2.7110 → $2.6691 (-1.55 %)
  • Needle 3, options as tools — $0.0162 → $0.0144 (-11.39 %)
  • Needle 3 — $0.0249 → $0.0238 (-4.36 %)

Who could not be measured, and why

An exclusion is an availability fact about our run — hardware, access, terms — never a quality verdict. Partial runs are in the table above, greyed and without a rank; so are the honorable mentions, which are complete runs that simply are not ranked.

  • open-jev (Dasein Labs) — MLX on Apple Silicon only. Its own README says Linux containers cannot reach the Apple GPU, so a RunPod NVIDIA GPU cannot run it.
  • open-jev (JoshuaSP) — A DiffusionGemma 26B-A4B serving wrapper rather than new trained weights. It was demonstrated on an H100; no public endpoint exists and no suitable 80 GB RunPod host was available in this round.
  • mini-jev (Mikhail Rakutko (r-ms)) — Public weights exist and fit a normal GPU, but the implementation covers Choice/Noul and explicitly does not measure Score. A faithful full-suite adapter would require new interface work rather than a mechanical endpoint adapter.
  • system-one-gemma (Akash Kamat) — The adapter is public, but its Gemma base is gated behind Google’s licence terms. We do not accept binding terms on Florian’s behalf.
  • jevlike (Vincent Wang-Maścianica) — Only Doom and chess vision checkpoints are released; there is no general text-decision checkpoint for this suite.
  • AlexWortega/openjev (Alex Wortega) — Its released NLI and task-specific heads do not define a distribution over an arbitrary supplied label set. Inventing that mapping would measure our assumption.
  • Needle 3 (Cactus Compute) — Its native response is a chosen label plus one accept/refuse confidence, not a categorical distribution over the supplied labels. The options-as-tools adaptation remains published as a partial run.
  • Succinct Router 14M (Pedro Marques) — A router over three fixed GPT settings, not a general typed-decision model.
  • jev-model-router, Director, Loki (various) — Applications built on decision models, not decision models themselves.
  • ProgramAsWeights (ProgramAsWeights) — The compiler still requires GitHub authentication and the available path would expose held-out rubrics to a third party. No public weights or anonymous endpoint are available.
  • EigenJev (EigenJev) — The endpoint requires authentication and no public weights or runnable implementation are published.
  • NanoJev (NanoJev) — Public weights exist, but the server exposes a different schema (including boolean rather than Noul) and lacks the full structured/null contract. It needs substantive compatibility work before a fair full-suite run.
  • Werr (pCwOrM) — Its documented server imports a module (scratch.jevbench_eval.optimize_werr_jevbench) that is not in the public repository, so the submitted configuration cannot be started; its engine also sends telemetry about each request to an outside server by default.
  • DIY Jev (VakeDomen) — The repository named in the request (github.com/VakeDomen/DIY-Jev) answers 404, so there is nothing to run.
  • SimpleJev RWKV variants (SimpleJev) — The public demo exposes RWKV IDs, but it does not identify their exact checkpoints or licences. Without reproducible model provenance, we do not publish benchmark rows for them.

Method and tiers

Built and run by us, not collected from someone else's leaderboard; the results describe the tested configurations, not every application.

JevBench Score. Equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost (power mean p=-1). If Intelligence <50 multiply by (I/50)^2. For Speed and Cost separately, if below 50 multiply by (axis/50)^2.

Intelligence. 0.8 × v1.3 chance-corrected Intelligence on the frozen v1.2 items + 0.2 × 100 × max(0, (sealed accuracy − 0.293)/(1 − 0.293)); then multiply by 1 − max(0, public-minus-sealed accuracy gap in percentage points − 25)/100. Original tier weights: easy .14, standard .28, judge .28, hard .30.

Revision v1.4.1. v1.4.1 adds 6 systems omitted from the v1.4.0 freeze. The v1.4 scoring formulas and prior system measurements are unchanged.

  • easy: 72 clear-cut decisions (intent, explicit yes/no fact, enum extraction, one-obvious-tool selection); new in v1.1
  • standard: 96 authored decisions from v1.0 (policy, intent, extraction, ordinal, adequacy, routing), unchanged
  • judge: 146 imported decisions from v1.0 (routing real task prompts into 9 categories; judging whether a saved math answer is correct), unchanged
  • hard: 220 new decisions (111 public, 109 held out): long multi-condition policy documents (2-6k tokens), priority trade-offs, deliberately ambiguous cases with a 'no clear answer' label, traps, multi-hop lookups, date/number reasoning, adversarial distractors, subtle answer-judging, overlapping routing, and probability items with an exact gold distribution. Half written by Claude Opus 5, half by GPT-5.6 Sol; each item reviewed blind and then against its gold by the other model; one discussion round; frozen and hashed before any benchmarked system saw an item. No item was selected on any system's answers.
  • sealed: 308 fresh private decisions across ten families, run once per system. Only system-level aggregates — overall and per-family accuracy, calibration — are published; the item text, answers and per-item results stay private and rotate between versions.

Every system sees the same state, instructions, rubric and exact label set; only the transport differs. Requests go out one at a time with no retries, so latency includes the network. Estimated costs are hosted-provider prices for the same weights or size class and are marked “est.” — hover one for its basis, or see how costs are estimated. Every system has a price; none gets a free 100.

A service running another entrant's model is listed, but not ranked against the models. Ranking it would rank the same model twice, once at the model's own price and once at the service's. The row keeps every number, axis, cost basis and per-task outcome; it carries no rank number. Which rows this applies to, and why: Honorable mentions — services built on another entrant's model.

v1.4 scores are not comparable with v1.3 or earlier (sealed blend, gap penalty and harmonic mean). The v1.3.0 board stays below as history; the v1.0 page keeps its own numbers, calibration plots and per-family tables.

Limits

  • 842 decisions (534 public, 308 sealed) is a pilot, not a census, and it is English-only. The v1.4 sealed set is very hard: most systems score close to chance on it, so sealed accuracy separates the field less than public accuracy does.
  • The weights are a choice. The JevBench Score weights the four axes equally and uses a harmonic mean, so the weakest axis dominates; if a wrong decision costs you more than a slow or expensive one, read the Intelligence column and the accuracy radars rather than the score alone.
  • The latency adjustment (×2, +0.15 s on our own servers) is an assumption, not a measurement. We ran the self-hosted and demo endpoints one request at a time (parallelism 1, no other load), so their latency is likely better than the same model on a busy production server. The official Jev API is presumably under high load, given the public interest. Serving under load trades per-user speed for throughput: in the NVIDIA chart shown by SemiAnalysis, moving to the throughput-maximising setting cuts per-user tokens per second by far more than 2×. That chart is a 1.8T mixture-of-experts model on GPU clusters, not a 4B model on one GPU, so it supports the direction and size of the effect, not our exact factor. The +0.15 s stands for infrastructure our self-hosted tests lacked: authentication, load balancing, logging, billing and an API gateway. Both numbers are assumptions; raw p50/p95 latencies are in the table and the repo, and a measurement under load is planned.
  • Held-out decisions are sent to the evaluated services to get predictions. Not public is not the same as not seen.
  • Latency is one origin at one time of day; hosted endpoints, public demos and a local CPU are different kinds of latency. Public demo endpoints are shared with everyone else using them.
  • Estimated costs describe what a large inference provider would charge for a model of that size, not what the author pays; a system on a tariff pays its tariff.

Credit

Harness, public tasks and every scoring rule: github.com/fstandhartinger/jevbench (MIT). Each project links its author's repository or vendor page.

Authors: if we tested the wrong configuration, tell us and we will rerun it. New entrants become a new version rather than silently changing this one.

JevBench v1.4.1 · additional views

Capability, cost and speed

Capability is the arithmetic mean of Intelligence and Calibration: (Intelligence + Calibration) / 2, on a 0–100 scale. Cost is USD per 1,000 decisions; its axis is logarithmic, and lower is better. Speed uses the JevBench Speed axis, where higher is faster. Estimated costs are marked.

3 placeholder rows have no published Intelligence or Calibration values and are omitted.

媒体内容 · 前往原文查看

Top 20 by Capability

Capability with cost alongside

Each system has a wide Capability bar and a narrower cost bar. The cost scale is logarithmic: longer bars mean higher cost, so shorter is cheaper.

020406080100I · C inputs
  1. 1GPT-6 Luna (medium)API95.4Cost $0.14 · I 97.4 · C 93.5
  2. 2DeepSeek V4.1 FlashAPI94.7Cost $0.59 · I 94.0 · C 95.5
  3. 3GPT-6 Luna (low)API93.9Cost $0.13 · I 95.8 · C 92.0
  4. 4GPT-5.6 LunaAPI90.3Cost $0.24 · I 93.1 · C 87.4
  5. 5djev79.7Cost $0.27 est. · I 71.6 · C 87.8
  6. 6Qwen3.8 27BAPI67.0Cost $2.67 est. · I 40.4 · C 93.6
  7. 7Jev 1.13.0API64.7Cost $0.040 · I 53.1 · C 76.3
  8. 8NInfer Qwen3.8-Flash-Next mixed64.1Cost $0.11 est. · I 49.5 · C 78.6
  9. 9NInfer Qwen3.8-27B NVFP463.7Cost $0.14 est. · I 51.5 · C 76.0
  10. 10Hopper63.5Cost $0.024 est. · I 48.0 · C 79.1
  11. 11SimpleJev Qwen3.8-27BAPI63.0Cost $0.10 est. · I 51.6 · C 74.5
  12. 12JevOne62.6Cost $0.14 est. · I 47.6 · C 77.6
  13. 13classifier.devAPI62.0Cost $0.0033 est. · I 51.6 · C 72.4
  14. 14reflex-27b61.8Cost $0.18 est. · I 46.4 · C 77.2
  15. 15JevK5 v0.2.061.7Cost $0.022 est. · I 48.9 · C 74.5
  16. 16LitJev61.4Cost $0.16 est. · I 46.3 · C 76.6
  17. 17openjev-sglangAPI59.3Cost $0.13 est. · I 49.4 · C 69.2
  18. 18NInfer Qwen3.8-27B NVFP459.3Cost $0.14 est. · I 51.5 · C 67.2
  19. 19jqv59.0Cost $0.056 est. · I 46.4 · C 71.6
  20. 20reflex 4B58.9Cost $0.022 est. · I 47.5 · C 70.4
Show all 79 systems (59 more)
  1. 21ZeroEntropy zerank-258.9Cost $0.047 · I 42.1 · C 75.8
  2. 22local-jev Qwen3.5-4B58.9Cost $0.030 est. · I 44.4 · C 73.3
  3. 23OpenJev58.1Cost $0.25 est. · I 58.1 · C 58.1
  4. 24JEV Qwen3.5-9B Base NVFP457.3Cost $0.077 est. · I 46.8 · C 67.7
  5. 25Gemini 3.1 Flash-LiteAPI56.9Cost $0.26 · I 54.5 · C 59.3
  6. 26Winnow-12B Q856.6Cost $0.037 est. · I 48.3 · C 64.8
  7. 27Decision 2B56.4Cost $0.018 est. · I 38.8 · C 74.1
  8. 28decider-35b-a3b56.2Cost $0.067 est. · I 47.2 · C 65.3
  9. 29metask-jev-4b55.8Cost $0.033 est. · I 44.7 · C 66.9
  10. 30SemIf55.6Cost $0.022 est. · I 44.4 · C 66.8
  11. 31Jev-Omni55.4Cost $0.037 est. · I 46.8 · C 64.1
  12. 32Von55.1Cost $0.0055 est. · I 34.5 · C 75.7
  13. 33Jobe Qwen3.5-4B55.1Cost $0.022 est. · I 44.1 · C 66.1
  14. 34Qwen3-Reranker-4B54.9Cost $0.050 · I 44.6 · C 65.2
  15. 35decision-machine-1API54.8Cost $0.035 · I 41.3 · C 68.3
  16. 36jev-local54.7Cost $0.077 est. · I 45.2 · C 64.2
  17. 37Qwen3.5-9B Jev-like data-mix v254.4Cost $0.083 est. · I 47.4 · C 61.3
  18. 38Open-Jev 9B53.0Cost $0.25 est. · I 44.2 · C 61.8
  19. 39SimpleJev Qwen3.6-35B-A3BAPI52.8Cost $0.12 est. · I 45.7 · C 59.8
  20. 40lev-350m52.7Cost $0.0063 est. · I 34.8 · C 70.6
  21. 41jeff52.3Cost $0.0060 est. · I 36.8 · C 67.9
  22. 42Bespoke Nimble 9B51.4Cost $0.17 est. · I 46.3 · C 56.4
  23. 43Decision Fast51.2Cost $0.0063 est. · I 37.1 · C 65.3
  24. 44djev51.2Cost $0.026 · I 47.0 · C 55.4
  25. 45OpenSourceJev51.1Cost $0.016 est. · I 41.8 · C 60.3
  26. 46openJev Verdict 1.450.7Cost $0.0039 est. · I 29.4 · C 72.0
  27. 47Raw Phi-4 mini direct logits50.3Cost $0.048 est. · I 41.8 · C 58.8
  28. 48OpenJev50.2Cost $0.066 est. · I 45.4 · C 55.0
  29. 49Laya49.9Cost $0.0029 est. · I 36.1 · C 63.7
  30. 50system-one-openAPI49.5Cost $0.015 est. · I 44.2 · C 54.9
  31. 51Open-Jev 2B48.8Cost $0.25 est. · I 42.3 · C 55.3
  32. 52open-alternative-jev48.6Cost $0.022 est. · I 38.6 · C 58.7
  33. 53Qwen3.5-0.8B Decision Model48.2Cost $0.0065 est. · I 28.1 · C 68.2
  34. 54spark-s1-4b-v646.3Cost $0.025 est. · I 45.1 · C 47.6
  35. 55open-jev-deberta-v3-large46.1Cost $0.0073 est. · I 25.6 · C 66.6
  36. 56Mixedbread mxbai-rerank-base-v245.5Cost $0.012 · I 6.8 · C 84.1
  37. 57BAAI bge-reranker-v2-m344.6Cost $0.0077 · I 5.0 · C 84.2
  38. 58OpenDecision44.4Cost $0.0066 est. · I 31.8 · C 57.1
  39. 59smalljev semantic-v942.4Cost $0.025 est. · I 25.7 · C 59.2
  40. 60kev 0.6B42.1Cost $0.0063 est. · I 34.2 · C 50.0
  41. 61Alibaba GTE Reranker ModernBERT-base41.9Cost $0.010 · I 4.8 · C 78.9
  42. 62Certo v141.5Cost $0.00097 est. · I 0.1 · C 83.0
  43. 63decider-2b41.0Cost $0.020 est. · I 38.5 · C 43.5
  44. 64kev 8B41.0Cost $0.073 est. · I 41.8 · C 40.2
  45. 65kev 4B40.9Cost $0.019 est. · I 42.1 · C 39.6
  46. 66GLiNER2.5 multi40.1Cost $0.0039 est. · I 23.1 · C 57.2
  47. 67kev 0.5B40.1Cost $0.0063 est. · I 30.5 · C 49.7
  48. 68openJev Verdict38.5Cost $0.0037 est. · I 30.0 · C 47.0
  49. 69system-one38.2Cost $0.089 est. · I 43.6 · C 32.8
  50. 70Raw Qwen3 4B Instruct 2507 direct logits37.7Cost $0.022 est. · I 46.4 · C 29.1
  51. 71GLiNER2.5 small35.6Cost $0.0039 est. · I 20.5 · C 50.7
  52. 72SimpleJev35.3Cost $0.011 est. · I 21.5 · C 49.1
  53. 73Raw Qwen3 8B direct logits34.9Cost $0.087 est. · I 45.7 · C 24.1
  54. 74Raw Qwen3 1.7B direct logits28.8Cost $0.015 est. · I 33.2 · C 24.3
  55. 75GLiNER2 large28.0Cost $0.0077 est. · I 31.1 · C 24.8
  56. 76GLiNER226.3Cost $0.0037 est. · I 27.4 · C 25.2
  57. 77Open Jev JSON Canvas24.1Cost $0.065 est. · I 48.2 · C 0.0
  58. 78Raw Qwen3 0.6B direct logits21.8Cost $0.0074 est. · I 22.8 · C 20.7
  59. 79Mirror19.8Cost $0.0077 est. · I 13.6 · C 26.0
$0.0010$0.010$0.10$1.00
Cost per 1,000 decisions · logarithmic · lower is better; free is at the left edge.
Capability 0–100100Cost: $0.00097–$2.67 / 1k
  • Jev (TypeSafe, closed)
  • Jev rebuild
  • Instruction model, JSON schema
  • Service built on Jev
  • Zero-shot classifier
  • Closed decision API
  • Reranker (neutral adapter)
  • Raw-logit control (base model)
  • Native-logit decision engine
  • Cost per 1,000 decisions · log scale
Cost bars use the right-hand scale, from $0.00097 to $2.67 per 1,000 decisions. Free cost is placed at the cheapest edge; missing cost is shown as —.
媒体内容 · 前往原文查看

Capability vs cost

The upper-left is the more attractive area: higher Capability and lower cost.

  1. Capability #1 · GPT-6 Luna (medium)
    Capability 95.4 · Cost $0.14 / 1,000 · Speed 72.6 · JevBench #30
  2. Capability #2 · DeepSeek V4.1 Flash
    Capability 94.7 · Cost $0.59 / 1,000 · Speed 71.6 · JevBench #71
  3. Capability #3 · GPT-6 Luna (low)
    Capability 93.9 · Cost $0.13 / 1,000 · Speed 73.7 · JevBench #27
  4. Capability #4 · GPT-5.6 Luna
    Capability 90.3 · Cost $0.24 / 1,000 · Speed 77.5 · JevBench #53
  5. Capability #5 · djev
    Capability 79.7 · Cost $0.27 estimated / 1,000 · Speed 75.2 · JevBench #57
79 systems plotted. Hover or focus a point to read its values.
媒体内容 · 前往原文查看

Capability vs speed

The upper-right is the more attractive area: higher Capability and higher Speed.

  1. Capability #1 · GPT-6 Luna (medium)
    Capability 95.4 · Cost $0.14 / 1,000 · Speed 72.6 · JevBench #30
  2. Capability #2 · DeepSeek V4.1 Flash
    Capability 94.7 · Cost $0.59 / 1,000 · Speed 71.6 · JevBench #71
  3. Capability #3 · GPT-6 Luna (low)
    Capability 93.9 · Cost $0.13 / 1,000 · Speed 73.7 · JevBench #27
  4. Capability #4 · GPT-5.6 Luna
    Capability 90.3 · Cost $0.24 / 1,000 · Speed 77.5 · JevBench #53
  5. Capability #5 · djev
    Capability 79.7 · Cost $0.27 estimated / 1,000 · Speed 75.2 · JevBench #57
79 systems plotted. Hover or focus a point to read its values.

All three at once

The 3D view plots Capability vertically, lower cost to the right, and higher Speed toward you. Sphere size follows the JevBench score. Drag to rotate; pinch or scroll to zoom. The view loads when it scrolls into view.

Scroll here to load the interactive 3D view.

  1. Capability #1 · GPT-6 Luna (medium)
    Capability 95.4 · Cost $0.14 / 1,000 decisions · Speed 72.6 · JevBench #30
  2. Capability #2 · DeepSeek V4.1 Flash
    Capability 94.7 · Cost $0.59 / 1,000 decisions · Speed 71.6 · JevBench #71
  3. Capability #3 · GPT-6 Luna (low)
    Capability 93.9 · Cost $0.13 / 1,000 decisions · Speed 73.7 · JevBench #27
  4. Capability #4 · GPT-5.6 Luna
    Capability 90.3 · Cost $0.24 / 1,000 decisions · Speed 77.5 · JevBench #53
  5. Capability #5 · djev
    Capability 79.7 · Cost $0.27 / 1,000 decisions · Speed 75.2 · JevBench #57

Vertical: Capability · Right: cheaper · Toward you: faster

The interactive 3D view loads when this panel scrolls into view.

79 systems plotted; systems missing cost or Speed are omitted. three.js r128 is included under its MIT license.

JevBench v1.4.1 public split · input capacity and long inputs

Context length

Context length is the amount of input a model or service can accept in one request. It matters when an app sends a long conversation state, policy set, or document: a smaller window can force truncation or chunking. A larger window is a capacity ceiling, not a promise that the system will use every token well.

Across 82 v1.4.1 rows (77 ranked systems and 5 unranked additions), supported published values range from 512 tokens to 1,050,000 tokens; 8 have no published maximum we could verify. For Jev-class rebuilds, the chart and table separate a model's trained length from any published serving cap and show training truncation limits where available.

媒体内容 · 前往原文查看

Public accuracy by actual input length

Each line is one of the 13 top-15 systems with reconciled public outcomes and input-token counts. A point's tooltip shows the system, accuracy, correct answers and bucket size.

  • Top five · #1 Jev 1.13.0 (TypeSafe AI)
  • Top five · #3 Hopper
  • Top five · #4 Winnow-12B Q8
  • Top five · #5 reflex 4B (kshetrajna12)
  • Top five · #6 djev (Maisa, diffusion-gemma)
  • #7 Jev-Omni (akhilaaa3, Gemma-4-12B merged)
  • #8 metask-jev-4b
  • #9 SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)
  • #10 Jobe Qwen3.5-4B (frozen)
  • #11 local-jev Qwen3.5-4B
  • #13 spark-s1-4b-v6 (Open Spark Jev, abhishek085)
  • #14 jqv (Qwen3-32B zero-shot)
  • #15 Qwen3-Reranker-4B
Bucket counts vary because only stored usage telemetry is available. The chart describes these benchmark items; it does not show that context length alone caused a score change.

Exact correct counts and denominators by bucket

SystemMean input tokensLength n<2k2–8k8–16k16–64k64–256k256k–1M≥1M
#1 Jev 1.13.0 (TypeSafe AI)1,057.8183/231125/146 · 85.6%27/37 · 73.0%No dataNo dataNo dataNo dataNo data
#3 Hopper739.2231/231167/195 · 85.6%23/36 · 63.9%No dataNo dataNo dataNo dataNo data
#4 Winnow-12B Q8692.3231/231170/194 · 87.6%28/37 · 75.7%No dataNo dataNo dataNo dataNo data
#5 reflex 4B (kshetrajna12)696.8231/231162/195 · 83.1%21/36 · 58.3%No dataNo dataNo dataNo dataNo data
#6 djev (Maisa, diffusion-gemma)692.3231/231169/194 · 87.1%25/37 · 67.6%No dataNo dataNo dataNo dataNo data
#7 Jev-Omni (akhilaaa3, Gemma-4-12B merged)686.5231/231175/194 · 90.2%30/37 · 81.1%No dataNo dataNo dataNo dataNo data
#8 metask-jev-4b761.2230/231165/194 · 85.1%19/36 · 52.8%No dataNo dataNo dataNo dataNo data
#9 SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)698.6231/231166/195 · 85.1%21/36 · 58.3%No dataNo dataNo dataNo dataNo data
#10 Jobe Qwen3.5-4B (frozen)698.6231/231166/195 · 85.1%21/36 · 58.3%No dataNo dataNo dataNo dataNo data
#11 local-jev Qwen3.5-4B696.7231/231163/195 · 83.6%23/36 · 63.9%No dataNo dataNo dataNo dataNo data
#13 spark-s1-4b-v6 (Open Spark Jev, abhishek085)794.6231/231161/194 · 83.0%22/37 · 59.5%No dataNo dataNo dataNo dataNo data
#14 jqv (Qwen3-32B zero-shot)662.7231/231167/195 · 85.6%18/36 · 50.0%No dataNo dataNo dataNo dataNo data
#15 Qwen3-Reranker-4B2,401.9231/231134/177 · 75.7%7/24 · 29.2%13/27 · 48.1%3/3 · 100.0%No dataNo dataNo data

Coverage: 13 of the top 15 systems are shown. Exclusions: JevK5 v0.2.0 (excluded: per-item accuracy 199/231 does not match published v1.4.1); system-one-open (Gemma 4 E2B LoRA on an L4) (excluded: no per-item token counts). All 13 included systems exactly reproduce their published public accuracy; stored lengths cover 183–231 decisions per system.

Long-policy tasks show a separate stress point

Across 19 public items in the long_policy family, several systems scored well below their full public-set accuracy. The comparison uses the family label, not only the token buckets.

  • metask-jev-4b
    Long policy 6/19 (31.6%) vs 184/231 overall (79.7%): −48.1 pp.
  • spark-s1-4b-v6 (Open Spark Jev, abhishek085)
    Long policy 7/19 (36.8%) vs 183/231 overall (79.2%): −42.4 pp.
  • system-one-open (Gemma 4 E2B LoRA on an L4)
    Long policy 6/19 (31.6%) vs 169/231 overall (73.2%): −41.6 pp.
  • Winnow-12B Q8
    Long policy 15/19 (78.9%) vs 198/231 overall (85.7%): −6.8 pp.
  • Jev-Omni (akhilaaa3, Gemma-4-12B merged)
    Long policy 15/19 (78.9%) vs 205/231 overall (88.7%): −9.8 pp.

Show long_policy results for all 14 matched systems

SystemOverallLong policy (19 items)Change
#8 metask-jev-4b184/231 · 79.7%6/19 · 31.6%−48.1 pp
#13 spark-s1-4b-v6 (Open Spark Jev, abhishek085)183/231 · 79.2%7/19 · 36.8%−42.4 pp
#12 system-one-open (Gemma 4 E2B LoRA on an L4)169/231 · 73.2%6/19 · 31.6%−41.6 pp
#14 jqv (Qwen3-32B zero-shot)185/231 · 80.1%9/19 · 47.4%−32.7 pp
#6 djev (Maisa, diffusion-gemma)194/231 · 84.0%10/19 · 52.6%−31.4 pp
#9 SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)187/231 · 81.0%10/19 · 52.6%−28.3 pp
#10 Jobe Qwen3.5-4B (frozen)187/231 · 81.0%10/19 · 52.6%−28.3 pp
#5 reflex 4B (kshetrajna12)183/231 · 79.2%10/19 · 52.6%−26.6 pp
#3 Hopper190/231 · 82.3%11/19 · 57.9%−24.4 pp
#1 Jev 1.13.0 (TypeSafe AI)200/231 · 86.6%12/19 · 63.2%−23.4 pp
#11 local-jev Qwen3.5-4B186/231 · 80.5%12/19 · 63.2%−17.4 pp
#15 Qwen3-Reranker-4B157/231 · 68.0%10/19 · 52.6%−15.3 pp
#7 Jev-Omni (akhilaaa3, Gemma-4-12B merged)205/231 · 88.7%15/19 · 78.9%−9.8 pp
#4 Winnow-12B Q8198/231 · 85.7%15/19 · 78.9%−6.8 pp

For example, metask-jev-4b scored 31.6% on long_policy versus 79.7% overall (change −48.1 pp), while Winnow-12B Q8 scored 78.9% versus 85.7% (change −6.8 pp). These are descriptive public-set comparisons. Prompt wrappers and tokenizers differ by system, and the 19-item family is small, so the results do not isolate context length as the cause. Only public item results were used; sealed-set item rows were not used.

媒体内容 · 前往原文查看

Published context limits · logarithmic scale

Each row shows the system's exact published maximum input context. API/serving caps, hard limits and trained lengths use different bar colors; training configuration limits appear as separate markers and values where published.

5122k8k32k128k512k1.05M
  1. #1 Jev 1.13.0 (TypeSafe AI)64,000 · API cap
  2. #2 JevK5 v0.2.0262,144 · Trained lengthTraining configuration: max_seq_len 2,048 tokens.
  3. #3 Hopper262,144 · Trained length
  4. #4 Winnow-12B Q865,536 · API cap
  5. #5 reflex 4B (kshetrajna12)262,144 · Trained length
  6. #6 djev (Maisa, diffusion-gemma)32,768 · API cap
  7. #7 Jev-Omni (akhilaaa3, Gemma-4-12B merged)262,144 · Trained length
  8. #8 metask-jev-4b262,144 · Trained length
  9. #9 SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)262,144 · Trained length
  10. #10 Jobe Qwen3.5-4B (frozen)262,144 · Trained lengthTraining configuration: max_seq_len 4,096 tokens.
  11. #11 local-jev Qwen3.5-4B262,144 · Trained length
  12. #12 system-one-open (Gemma 4 E2B LoRA on an L4)131,072 · Trained lengthTraining configuration: state limit 2,048 tokens.
  13. #13 spark-s1-4b-v6 (Open Spark Jev, abhishek085)262,144 · Trained lengthTraining configuration: max_seq_len 2,048 tokens.
  14. #14 jqv (Qwen3-32B zero-shot)32,768 · Trained length
  15. #15 Qwen3-Reranker-4B32,768 · Trained length
  16. #16 decider-35b-a3b (Mapika)262,144 · Hard limit
  17. #17 Raw Qwen3 4B Instruct 2507 direct logits262,144 · Trained length
  18. #18 OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp)32,768 · Trained length
  19. #19 ZeroEntropy zerank-232,768 · Trained length
  20. #20 decision-machine-1 (milliseconds.ai)Unknown
  21. #21 Raw Phi-4 mini direct logits131,072 · Trained length
  22. #22 JEV Qwen3.5-9B Base NVFP4262,144 · Trained length
  23. #23 OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)65,536 · API cap
  24. #24 kev 4B (research preview)32,768 · Trained length
  25. #25 Decision 2B (FlyMy.AI, v59)131,072 · Trained length
  26. #26 Qwen3.5-9B Jev-like data-mix v2262,144 · Trained length
  27. #27 GPT-6 Luna (low reasoning effort)1,050,000 · API cap
  28. #28 SimpleJev Qwen3.8-27B2,000 · API cap
  29. #29 NInfer Qwen3.8-Flash-Next mixed262,144 · Trained length
  30. #30 GPT-6 Luna (default medium reasoning effort)1,050,000 · API cap
  31. #31 open-alternative-jev (Qwen3.5-4B, IkerMoel)262,144 · Trained length
  32. #32 jev-local (Qwen3.5-9B)262,144 · Trained length
  33. #33 Decision Fast (FlyMy.AI, v53a)32,768 · Trained length
  34. #34 decider-2b (Mapika)262,144 · Hard limit
  35. #35 jeff (Logan Markewich, GLiFormer 400M)8,192 · Hard limit
  36. #36 Laya (Convai Innovations, ModernBERT-large 421M)512 · Hard limit
  37. #37 lev-350m (Franck Verrot, LFM2.5-350M)32,768 · Trained length
  38. #38 openjev-sglang (Qwen3.6-35B-A3B on SGLang)32,768 · API cap
  39. #39 Von (wfzyx, Option-Marker 395M)8,192 · Trained length
  40. #40 NInfer Qwen3.8-27B NVFP4 (T=1.5)262,144 · Trained length
  41. #41 NInfer Qwen3.8-27B NVFP4262,144 · Trained length
  42. #42 kev 8B (research preview)32,768 · Trained length
  43. #43 JevOne262,144 · Trained length
  44. #44 SimpleJev Qwen3.6-35B-A3B262,144 · Trained length
  45. #45 kev 0.6B (research preview)32,768 · Trained length
  46. #46 Raw Qwen3 8B direct logits32,768 · Trained length
  47. #47 system-one (Qwen3-8B, Sean Goedecke)32,768 · Trained length
  48. #48 OpenDecision (ModernBERT-large zero-shot)8,192 · Hard limit
  49. #49 LitJev (Qwen3.8-27B)262,144 · Trained length
  50. #50 openJev Verdict 1.4512 · API cap
  51. #51 kev 0.5B8,192 · API cap
  52. #52 Bespoke Nimble 9B (Bespoke Labs)8,192 · API capTraining configuration: max_seq_len 2,048 tokens.
  53. #53 GPT-5.6 Luna (low reasoning effort)1,050,000 · API cap
  54. #54 openJev Verdict (heman10x, ModernBERT-base 151M)8,192 · Hard limit
  55. #55 Raw Qwen3 1.7B direct logits32,768 · Trained length
  56. #56 reflex-27b (Qwen3.8-27B)262,144 · Trained length
  57. #57 djev (thinking)262,144 · Trained length
  58. #58 GLiNER2 large (Fastino)Unknown
  59. #59 OpenJev (thinking, BF16)65,536 · API cap
  60. #60 Qwen3.5-0.8B Decision Model (Mourad Ghafiri)262,144 · Hard limit
  61. #61 Gemini 3.1 Flash-Lite1,048,576 · API cap
  62. #62 open-jev-deberta-v3-large (local CPU)512 · Hard limit
  63. #63 smalljev semantic-v9131,072 · Trained length
  64. #64 GLiNER2 (Fastino, gliner2.5-base)Unknown
  65. #65 Open-Jev 9B (Zefan Cai)262,144 · Trained length
  66. #66 Open-Jev 2B (Zefan Cai)262,144 · Trained length
  67. #67 GLiNER2.5 multi (Fastino, 287M)Unknown
  68. #68 SimpleJev (Qwen3.5-0.8B, CPU)262,144 · Trained length
  69. #69 GLiNER2.5 small (Fastino, 74M)Unknown
  70. #70 Raw Qwen3 0.6B direct logits32,768 · Trained length
  71. #71 DeepSeek V4.1 Flash (thinking default)1,000,000 · API cap
  72. #72 Mirror512 · Hard limit
  73. #73 Mixedbread mxbai-rerank-base-v232,768 · Hard limit
  74. #74 BAAI bge-reranker-v2-m38,192 · Hard limit
  75. #75 Alibaba GTE Reranker ModernBERT-base8,192 · Hard limit
  76. #76 Certo v1 (AltSlate Labs)8,192 · Hard limit
  77. #77 Open Jev JSON Canvas (JoshuaSP)262,144 · Trained length
  78. Unranked classifier.dev (fast tier)Unknown
  79. Unranked Needle 3 (Cactus, 2-bit, local CPU)Unknown
  80. Unranked Needle 3, options as tools (post-hoc adapter mode)Unknown
  81. Unranked Qwen3.8 27B (Chutes TEE)262,144 · API cap
  82. Unranked swanOne262,144 · Trained length
  • API / serving cap
  • Hard limit
  • Trained length
  • Training max_seq_len
  • Training state limit
The scale runs from 512 tokens to 1,050,000 tokens. Source links, dates, evidence notes and exact training details remain in the table below.

Context limits by system

82 systems · sources checked 24 Sept 2026

Sort by selecting a column heading. Source links open the primary model card, vendor documentation, or API documentation.

#1Jev 1.13.0 (TypeSafe AI)Basis, training and serving notes

Evidence: TypeSafe's official Jev 1.13 docs specify a 64K request budget and a 32K state-plus-longest-question budget.

System repository

64,000 total request; 32,000 state + longest questionOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026API cap
#2JevK5 v0.2.0Basis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

Training max_seq_len: 2,048 tokens.

Training utility defaults --max-len to 2,048 and skips longer rows. Runtime has no smaller total context cap documented; native Qwen3.5-4B window is 262,144.

System repository · 23 Sept 2026

Training configuration · 23 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#3HopperBasis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

JevBench registry identifies this as a LoRA on Qwen3.5-4B. Its public model card does not specify a shorter max_seq_len or serving truncation, so the base model's native 262,144 window is listed; adapter training length is unknown.

System repository · 22 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#4Winnow-12B Q8Basis, training and serving notes

Base model: google/gemma-4-12B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: The Winnow model card documents a 65,536-position configured Q8 test profile; Gemma 4 12B base supports 262,144.

System repository · 21 Sept 2026

Base model source · 20 Jul 2026

65,536 configured/tested; base 262,144Open primary sourceSource date: 21 Sept 2026 · checked: 23 Sept 2026API cap
#5reflex 4B (kshetrajna12)Basis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

Server refuses an input above the base model's context window rather than truncating it. The repo's 8,192 max-pack-tokens is a batching budget, not the per-request context cap.

System repository · 23 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#6djev (Maisa, diffusion-gemma)Basis, training and serving notes

Base model: google/diffusiongemma-26B-A4B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Runtime setting is bounded to 1,024–32,768; DiffusionGemma base is 262,144.

System repository · 19 Sept 2026

Base model source · 15 Jul 2026

32,768 (prompt + reserved canvas)Open primary sourceSource date: 19 Sept 2026 · checked: 23 Sept 2026API cap
#7Jev-Omni (akhilaaa3, Gemma-4-12B merged)Basis, training and serving notes

Base model: google/gemma-4-12B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Gemma 4 12B card states a 256K context window.

System repository · 22 Sept 2026

262,144Open primary sourceSource date: 20 Jul 2026 · checked: 23 Sept 2026Trained length
#8metask-jev-4bBasis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

Reported validation point: 4,096 tokens. This is not automatically the maximum accepted input.

Model card reports 4,096 as the validated evaluation point, not an architectural limit; the same card/config says native 262,144. Published validated evaluation point: 4,096 tokens.

System repository · 22 Sept 2026

Training configuration · 22 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#9SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ)Basis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

System repository · 23 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#10Jobe Qwen3.5-4B (frozen)Basis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

Training max_seq_len: 4,096 tokens.

The optional adapter-training helper uses max_tokens=4,096 and rejects longer training examples. The ranked v1.4.1 entry is the frozen Qwen3.5-4B backbone, so this optional training helper does not set its inference window.

System repository · 23 Sept 2026

Training configuration · 23 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#11local-jev Qwen3.5-4BBasis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

The README documents long-state shortening. Its 32,768-token context_tokens value is only an illustrative custom-model card; the effective Qwen3.5 deployment cap is not published. The listed 262,144 is the base model window.

System repository · 21 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#12system-one-open (Gemma 4 E2B LoRA on an L4)Basis, training and serving notes

Base model: google/gemma-4-E2B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Gemma 4 E2B card/config state 131,072 (128K) positions.

Training state limit: 2,048 tokens.

Training batch token budget: 24,576 tokens.

Full training profile caps state at 2,048 tokens and total batch tokens at 24,576; this is a training profile, not an inference limit.

System repository · 17 Sept 2026

Training configuration · 17 Sept 2026

131,072Open primary sourceSource date: 20 Jul 2026 · checked: 23 Sept 2026Trained length
#13spark-s1-4b-v6 (Open Spark Jev, abhishek085)Basis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

Training max_seq_len: 2,048 tokens.

Published RLCD configs use max_len=2,048 for training; Qwen3.5-4B native inference window is 262,144.

System repository · 22 Sept 2026

Training configuration · 22 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#14jqv (Qwen3-32B zero-shot)Basis, training and serving notes

Base model: Qwen/Qwen3-32B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3 card specifies 32,768 native. The config's 40,960 positions reserve output space; 131,072 requires YaRN.

System repository · 23 Sept 2026

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#15Qwen3-Reranker-4BBasis, training and serving notes

Base model: Qwen/Qwen3-Reranker-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official reranker card states 32K context; config has extra positions reserved for prompt/output.

Official card's stated context is 32K; do not substitute the larger config allocation because the card is explicit.

System repository · 16 Apr 2026

32,768Open primary sourceSource date: 16 Apr 2026 · checked: 23 Sept 2026Trained length
#16decider-35b-a3b (Mapika)Basis, training and serving notes

Base model: Qwen/Qwen3.5-35B-A3B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: The submitted model's config.json sets 262,144 positions; it is a decision readout on Qwen3.5-35B-A3B-Base, whose native context is also 262,144.

System repository · 23 Sept 2026

Base model source · 23 Apr 2026

262,144Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Hard limit
#17Raw Qwen3 4B Instruct 2507 direct logitsBasis, training and serving notes

Base model: Qwen/Qwen3-4B-Instruct-2507. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official model card states 262,144 natively.

System repository · 17 Sept 2025

262,144Open primary sourceSource date: 17 Sept 2025 · checked: 23 Sept 2026Trained length
#18OpenSourceJev (Qwen3.5-4B Q4_K_M, native llama.cpp)Basis, training and serving notes

Base model: Qwen/Qwen3-1.7B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3-1.7B card specifies 32,768 context.

System repository · 23 Sept 2026

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#19ZeroEntropy zerank-2Basis, training and serving notes

Base model: zeroentropy/zerank-2-reranker. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official model card states 32,768 context.

System repository · 24 Jul 2026

32,768Open primary sourceSource date: 24 Jul 2026 · checked: 23 Sept 2026Trained length
#20decision-machine-1 (milliseconds.ai)Basis, training and serving notes

Evidence: No primary public model-card or API context limit was found.

System repository

UnknownOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Unknown
#21Raw Phi-4 mini direct logitsBasis, training and serving notes

Base model: microsoft/Phi-4-mini-instruct. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Microsoft card states 128K context; config uses long-RoPE scaling.

System repository · 10 Dec 2025

131,072Open primary sourceSource date: 10 Dec 2025 · checked: 23 Sept 2026Trained length
#22JEV Qwen3.5-9B Base NVFP4Basis, training and serving notes

Base model: Qwen/Qwen3.5-9B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-9B: native context length 262,144; optional YaRN extension is not enabled by default.

System repository · 22 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#23OpenJev (DiffusionGemma 26B-A4B NVFP4, razorback16)Basis, training and serving notes

Base model: google/diffusiongemma-26B-A4B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: OPENJEV_MAX_MODEL_LEN defaults to 65,536 in the submitted runner; DiffusionGemma base supports 262,144.

System repository · 23 Sept 2026

Base model source · 15 Jul 2026

65,536; base 262,144Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026API cap
#24kev 4B (research preview)Basis, training and serving notes

Base model: Qwen/Qwen3-4B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3-4B-Base is in the Qwen3 family with a 32,768-token native window.

System repository · 23 Sept 2026

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#25Decision 2B (FlyMy.AI, v59)Basis, training and serving notes

Base model: openbmb/MiniCPM5-2B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: MiniCPM5-2B card/config states 131,072 context.

System repository · 23 Sept 2026

131,072Open primary sourceSource date: 12 Sept 2026 · checked: 23 Sept 2026Trained length
#26Qwen3.5-9B Jev-like data-mix v2Basis, training and serving notes

Base model: Qwen/Qwen3.5-9B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-9B: native context length 262,144; optional YaRN extension is not enabled by default.

System repository · 22 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#27GPT-6 Luna (low reasoning effort)Basis, training and serving notes

Evidence: OpenAI model docs: 1,050,000 context window and 128,000 maximum output.

1,050,000 context window; 128,000 max outputOpen primary sourceSource date: 22 Sept 2026 · checked: 23 Sept 2026API cap
#28SimpleJev Qwen3.8-27BBasis, training and serving notes

Base model: Qwen/Qwen3.8-27B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: The tested public Simple Jev demo API documents a 2,000-token context limit; the Qwen3.8-27B base window is 262,144.

System repository · 23 Sept 2026

Base model source · 14 Aug 2026

2,000 demo API context; base 262,144Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026API cap
#29NInfer Qwen3.8-Flash-Next mixedBasis, training and serving notes

Base model: Qwen/Qwen3.8-Flash-Next. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.8-Flash-Next card: 262,144 native; the 1M extension requires YaRN.

System repository · 22 Sept 2026

262,144Open primary sourceSource date: 27 Aug 2026 · checked: 23 Sept 2026Trained length
#30GPT-6 Luna (default medium reasoning effort)Basis, training and serving notes

Evidence: OpenAI model docs: 1,050,000 context window and 128,000 maximum output.

1,050,000 context window; 128,000 max outputOpen primary sourceSource date: 22 Sept 2026 · checked: 23 Sept 2026API cap
#31open-alternative-jev (Qwen3.5-4B, IkerMoel)Basis, training and serving notes

Base model: Qwen/Qwen3.5-4B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-4B: native context length 262,144; optional YaRN extension is not enabled by default.

System repository · 22 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#32jev-local (Qwen3.5-9B)Basis, training and serving notes

Base model: Qwen/Qwen3.5-9B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-9B: native context length 262,144; optional YaRN extension is not enabled by default.

System repository · 18 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#33Decision Fast (FlyMy.AI, v53a)Basis, training and serving notes

Base model: Qwen/Qwen3-0.6B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3-0.6B-Base config specifies 32,768 positions.

System repository · 23 Sept 2026

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#34decider-2b (Mapika)Basis, training and serving notes

Base model: Qwen/Qwen3.5-2B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: The submitted model's config.json sets 262,144 positions; it is a decision readout on Qwen3.5-2B-Base, whose native context is also 262,144.

System repository · 23 Sept 2026

Base model source · 23 Apr 2026

262,144Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Hard limit
#35jeff (Logan Markewich, GLiFormer 400M)Basis, training and serving notes

Base model: knowledgator/gliformer-large-v1. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: GLiFormer card states configured max_len=8,192.

System repository · 20 Sept 2026

8,192Open primary sourceSource date: 18 Sept 2026 · checked: 23 Sept 2026Hard limit
#36Laya (Convai Innovations, ModernBERT-large 421M)Basis, training and serving notes

Base model: convaiinnovations/laya. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Laya model card documents a 512-token base context; head_max_len=192 is its answer-candidate budget.

System repository · 23 Sept 2026

512Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Hard limit
#37lev-350m (Franck Verrot, LFM2.5-350M)Basis, training and serving notes

Base model: LiquidAI/LFM2.5-350M. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: LiquidAI model card states 32,768 context; config has a larger positional allocation.

LiquidAI card states a 32,768 context length. Config has a larger positional allocation; no larger trained/evaluated sequence is claimed.

System repository · 21 Sept 2026

32,768Open primary sourceSource date: 5 Aug 2026 · checked: 23 Sept 2026Trained length
#38openjev-sglang (Qwen3.6-35B-A3B on SGLang)Basis, training and serving notes

Base model: Qwen/Qwen3.6-35B-A3B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Runtime defaults max_input_tokens=32,768 and max_total_input_tokens=262,144.

System repository · 21 Sept 2026

Base model source · 24 Apr 2026

32,768 per question; 262,144 total across questionsOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026API cap
#39Von (wfzyx, Option-Marker 395M)Basis, training and serving notes

Evidence: The Von model card states an 8,192-token context for its ModernBERT-large scoring model and describes accurate premise reading to about 2,048 tokens.

System repository · 23 Sept 2026

8,192 model context; reads well to about 2,048Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Trained length
#40NInfer Qwen3.8-27B NVFP4 (T=1.5)Basis, training and serving notes

Base model: Qwen/Qwen3.8-27B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.8-27B card: 262,144 native; one-million-token extension requires YaRN.

System repository · 22 Sept 2026

262,144Open primary sourceSource date: 14 Aug 2026 · checked: 23 Sept 2026Trained length
#41NInfer Qwen3.8-27B NVFP4Basis, training and serving notes

Base model: Qwen/Qwen3.8-27B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.8-27B card: 262,144 native; one-million-token extension requires YaRN.

System repository · 22 Sept 2026

262,144Open primary sourceSource date: 14 Aug 2026 · checked: 23 Sept 2026Trained length
#42kev 8B (research preview)Basis, training and serving notes

Base model: Qwen/Qwen3-8B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3 card specifies 32,768 native; 131,072 requires YaRN.

System repository · 23 Sept 2026

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#43JevOneBasis, training and serving notes

Base model: Qwen/Qwen3.6-35B-A3B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.6-35B-A3B card: 262,144 native; optional YaRN extension is not assumed.

JevOne is published as Qwen3.6-35B-A3B BF16 with a bidirectional option-logit mapping; no smaller serving or training sequence cap is documented.

System repository · 23 Sept 2026

Base model source · 24 Apr 2026

262,144Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Trained length
#44SimpleJev Qwen3.6-35B-A3BBasis, training and serving notes

Base model: Qwen/Qwen3.6-35B-A3B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.6-35B-A3B card: 262,144 native; optional YaRN extension is not assumed.

System repository · 23 Sept 2026

262,144Open primary sourceSource date: 24 Apr 2026 · checked: 23 Sept 2026Trained length
#45kev 0.6B (research preview)Basis, training and serving notes

Base model: Qwen/Qwen3-0.6B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3-0.6B-Base config specifies 32,768 positions.

System repository · 23 Sept 2026

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#46Raw Qwen3 8B direct logitsBasis, training and serving notes

Base model: Qwen/Qwen3-8B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3 card specifies 32,768 native; 131,072 requires YaRN.

System repository · 26 Jul 2025

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#47system-one (Qwen3-8B, Sean Goedecke)Basis, training and serving notes

Base model: Qwen/Qwen3-8B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3 card specifies 32,768 native; 131,072 requires YaRN.

System repository · 18 Sept 2026

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#48OpenDecision (ModernBERT-large zero-shot)Basis, training and serving notes

Base model: MoritzLaurer/ModernBERT-large-zeroshot-v2.0. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Config has 8,192 positions. Card says v2.0 may not fully use the 8K window; exact trained length is not stated.

ModernBERT config allows 8,192 positions. The v2.0 card says the older zero-shot checkpoint may not fully use the long window; it does not state a smaller exact trained limit.

System repository · 21 Sept 2026

8,192Open primary sourceSource date: 16 Jan 2025 · checked: 23 Sept 2026Hard limit
#49LitJev (Qwen3.8-27B)Basis, training and serving notes

Base model: Qwen/Qwen3.8-27B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.8-27B card: 262,144 native; one-million-token extension requires YaRN.

System repository · 21 Sept 2026

262,144Open primary sourceSource date: 14 Aug 2026 · checked: 23 Sept 2026Trained length
#50openJev Verdict 1.4Basis, training and serving notes

Base model: heman10x/rlcd-modernbert-151m. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: The fixed v1.4 inference path sets its context budget to 512; underlying model config is 8,192.

System repository · 20 Sept 2026

Base model source · 20 Sept 2026

512 service budget; model config 8,192Open primary sourceSource date: 20 Sept 2026 · checked: 23 Sept 2026API cap
#51kev 0.5BBasis, training and serving notes

Base model: Qwen/Qwen2.5-0.5B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Kev 0.5B model card specifies an 8,192-token serving allowance per branch and a 32K backbone window.

System repository · 23 Sept 2026

Base model source · 25 Sept 2024

8,192 per branch; backbone 32,768Open primary sourceSource date: 20 Sept 2026 · checked: 23 Sept 2026API cap
#52Bespoke Nimble 9B (Bespoke Labs)Basis, training and serving notes

Base model: Qwen/Qwen3.5-9B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Serving code defaults NIMBLE_MAX_PROMPT_TOKENS to 8,192 and launches the backend with that request cap.

Training max_seq_len: 2,048 tokens.

The published adapter-training recipe uses --max-length=2,048; the ranked serving profile has an 8,192-token prompt cap.

System repository · 23 Sept 2026

Base model source · 2 Mar 2026

Training configuration · 23 Sept 2026

8,192 prompt cap; base 262,144Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026API cap
#53GPT-5.6 Luna (low reasoning effort)Basis, training and serving notes

Evidence: OpenAI model docs: 1,050,000 context window and 128,000 maximum output.

1,050,000 context window; 128,000 max outputOpen primary sourceSource date: 9 Jul 2026 · checked: 23 Sept 2026API cap
#54openJev Verdict (heman10x, ModernBERT-base 151M)Basis, training and serving notes

Base model: knowledgator/gliclass-modern-base-v2.0. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Model tokenizer config sets an 8,192-token maximum.

Model config/tokenizer sets 8,192; its current separate v1.4 inference-engine cap is not documented in the available primary sources.

System repository · 20 Sept 2026

8,192Open primary sourceSource date: 12 Aug 2025 · checked: 23 Sept 2026Hard limit
#55Raw Qwen3 1.7B direct logitsBasis, training and serving notes

Base model: Qwen/Qwen3-1.7B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3-1.7B card specifies 32,768 context.

System repository · 26 Jul 2025

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#56reflex-27b (Qwen3.8-27B)Basis, training and serving notes

Base model: Qwen/Qwen3.8-27B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.8-27B card: 262,144 native; one-million-token extension requires YaRN.

System repository · 23 Sept 2026

262,144Open primary sourceSource date: 14 Aug 2026 · checked: 23 Sept 2026Trained length
#57djev (thinking)Basis, training and serving notes

Base model: google/diffusiongemma-26B-A4B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: DiffusionGemma card/config states 256K context.

The standard djev runtime caps requests at 32,768, but the separately measured full-generation thinking run has no matching runtime cap published. The listed 262,144 is the DiffusionGemma base window, not a verified cap for this run.

System repository · 19 Sept 2026

262,144Open primary sourceSource date: 15 Jul 2026 · checked: 23 Sept 2026Trained length
#58GLiNER2 large (Fastino)Basis, training and serving notes

Base model: microsoft/deberta-v3-large. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official docs provide chunked long-document helpers but no maximum total document size; the named DeBERTa-v3-large encoder has 512 positions.

System repository · 17 Sept 2026

Base model source · 19 Mar 2023

UnknownOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Unknown
#59OpenJev (thinking, BF16)Basis, training and serving notes

Base model: google/diffusiongemma-26B-A4B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: OpenJev runtime defaults OPENJEV_MAX_MODEL_LEN to 65,536; the 512-token thinking allowance is generated output, not input context.

System repository · 23 Sept 2026

Base model source · 15 Jul 2026

65,536; base 262,144Open primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026API cap
#60Qwen3.5-0.8B Decision Model (Mourad Ghafiri)Basis, training and serving notes

Base model: Qwen/Qwen3.5-0.8B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: The submitted decision model's config.json sets 262,144 positions; it is based on Qwen3.5-0.8B-Base, whose native context is 262,144.

System repository · 22 Sept 2026

Base model source · 23 Apr 2026

262,144Open primary sourceSource date: 22 Sept 2026 · checked: 23 Sept 2026Hard limit
#61Gemini 3.1 Flash-LiteBasis, training and serving notes

Evidence: Google model docs explicitly list a 1,048,576 input-token limit and 65,536 output-token limit.

1,048,576 input token limitOpen primary sourceSource date: 21 Jul 2026 · checked: 23 Sept 2026API cap
#62open-jev-deberta-v3-large (local CPU)Basis, training and serving notes

Base model: microsoft/deberta-v3-large. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Microsoft config max_position_embeddings=512.

System repository · 22 Sept 2026

512Open primary sourceSource date: 19 Mar 2023 · checked: 23 Sept 2026Hard limit
#63smalljev semantic-v9Basis, training and serving notes

Base model: openbmb/MiniCPM5-2B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: MiniCPM5-2B card/config states 131,072 context.

System repository · 21 Sept 2026

131,072Open primary sourceSource date: 12 Sept 2026 · checked: 23 Sept 2026Trained length
#64GLiNER2 (Fastino, gliner2.5-base)Basis, training and serving notes

Base model: microsoft/deberta-v3-large. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official docs provide chunked long-document helpers but no maximum total document size; underlying DeBERTa-v3-base encoder has 512 positions.

System repository · 23 Sept 2026

Base model source · 19 Mar 2023

UnknownOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Unknown
#65Open-Jev 9B (Zefan Cai)Basis, training and serving notes

Base model: Qwen/Qwen3.5-9B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.5-9B: native context length 262,144; optional YaRN extension is not enabled by default.

System repository · 23 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#66Open-Jev 2B (Zefan Cai)Basis, training and serving notes

Base model: Qwen/Qwen3.5-2B-Base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Mapika model config and Qwen3.5 family card give 262,144 positions.

System repository · 23 Sept 2026

262,144Open primary sourceSource date: 23 Apr 2026 · checked: 23 Sept 2026Trained length
#67GLiNER2.5 multi (Fastino, 287M)Basis, training and serving notes

Evidence: Model docs say max_len truncates and long-context helpers chunk documents; no fixed total input ceiling is published.

System repository · 20 Sept 2026

UnknownOpen primary sourceSource date: 20 Sept 2026 · checked: 23 Sept 2026Unknown
#68SimpleJev (Qwen3.5-0.8B, CPU)Basis, training and serving notes

Base model: Qwen/Qwen3.5-0.8B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official Qwen3.5-0.8B card/config reports 262,144 natively; optional YaRN extends the window.

System repository · 23 Sept 2026

262,144Open primary sourceSource date: 2 Mar 2026 · checked: 23 Sept 2026Trained length
#69GLiNER2.5 small (Fastino, 74M)Basis, training and serving notes

Evidence: Model docs say max_len truncates and long-context helpers chunk documents; no fixed total input ceiling is published.

System repository · 20 Sept 2026

UnknownOpen primary sourceSource date: 20 Sept 2026 · checked: 23 Sept 2026Unknown
#70Raw Qwen3 0.6B direct logitsBasis, training and serving notes

Base model: Qwen/Qwen3-0.6B. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3-0.6B card specifies 32,768 context.

System repository · 26 Jul 2025

32,768Open primary sourceSource date: 26 Jul 2025 · checked: 23 Sept 2026Trained length
#71DeepSeek V4.1 Flash (thinking default)Basis, training and serving notes

Evidence: DeepSeek API model/pricing docs list DeepSeek V4.1 Flash at 1M context and 384K maximum output.

1,000,000 context; 384,000 max outputOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026API cap
#72MirrorBasis, training and serving notes

Base model: microsoft/deberta-v3-large. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Microsoft config max_position_embeddings=512.

System repository

512Open primary sourceSource date: 19 Mar 2023 · checked: 23 Sept 2026Hard limit
#73Mixedbread mxbai-rerank-base-v2Basis, training and serving notes

Base model: mixedbread-ai/mxbai-rerank-base-v2. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official config max_position_embeddings=32,768.

System repository · 8 Apr 2026

32,768Open primary sourceSource date: 8 Apr 2026 · checked: 23 Sept 2026Hard limit
#74BAAI bge-reranker-v2-m3Basis, training and serving notes

Base model: BAAI/bge-reranker-v2-m3. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official tokenizer limit is 8,192; config has 8,194 positions.

System repository · 24 Jun 2024

8,192Open primary sourceSource date: 24 Jun 2024 · checked: 23 Sept 2026Hard limit
#75Alibaba GTE Reranker ModernBERT-baseBasis, training and serving notes

Base model: Alibaba-NLP/gte-reranker-modernbert-base. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Official config max_position_embeddings=8,192.

System repository · 4 Jul 2025

8,192Open primary sourceSource date: 4 Jul 2025 · checked: 23 Sept 2026Hard limit
#76Certo v1 (AltSlate Labs)Basis, training and serving notes

Base model: altslate/certo-decision-model. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Submitted model config/tokenizer sets 8,192.

System repository · 21 Sept 2026

8,192Open primary sourceSource date: 21 Sept 2026 · checked: 23 Sept 2026Hard limit
#77Open Jev JSON Canvas (JoshuaSP)Basis, training and serving notes

Base model: google/diffusiongemma-26B-A4B-it. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: DiffusionGemma card/config states 256K context.

System repository · 16 Sept 2026

262,144Open primary sourceSource date: 15 Jul 2026 · checked: 23 Sept 2026Trained length
Unrankedclassifier.dev (fast tier)Basis, training and serving notes

Evidence: The Jev model behind the service has a documented limit, but no primary source documents a narrower or matching classifier.dev endpoint cap.

System repository

UnknownOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Unknown
UnrankedNeedle 3 (Cactus, 2-bit, local CPU)Basis, training and serving notes

Evidence: Official Needle 3 docs describe text input but publish no maximum context/token limit.

System repository

UnknownOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Unknown
UnrankedNeedle 3, options as tools (post-hoc adapter mode)Basis, training and serving notes

Evidence: Official Needle 3 docs describe tool inputs but publish no maximum context/token limit.

System repository

UnknownOpen primary sourceSource date: 23 Sept 2026 · checked: 23 Sept 2026Unknown
UnrankedQwen3.8 27B (Chutes TEE)Basis, training and serving notes

Evidence: Chutes model catalog lists Qwen3.8-27B-TEE at 262K context.

262,144 contextOpen primary sourceSource date: 17 Aug 2026 · checked: 23 Sept 2026API cap
UnrankedswanOneBasis, training and serving notes

Base model: Qwen/Qwen3.8-Flash-Next. A Jev-class adapter normally inherits this window unless its training or serving setup truncates input.

Evidence: Qwen3.8-Flash-Next card: 262,144 native; the 1M extension requires YaRN.

The submitted NVFP4 runner is based on Qwen3.8-Flash-Next; 262,144 is its native window. No larger runtime setting is documented for the submitted patch.

System repository

262,144Open primary sourceSource date: 27 Aug 2026 · checked: 23 Sept 2026Trained length

“Hard limit” is an explicit model or tokenizer ceiling; “Trained length” is a published base-model or training length; “API cap” is a published service limit. These are different kinds of evidence. A base-model window does not prove that a particular hosted endpoint accepts the same length; row notes identify cases where its serving cap is unpublished. Some API docs publish a combined context window and a separate output ceiling, so the usable input can be lower when output tokens share that window. “Unknown” means no supported maximum was found.

The input-length chart uses each system's existing usage.input_tokens telemetry and public item outcomes; no new model runs were made. Source dates are listed beside each primary-source link, and every source was checked 24 Sept 2026. Training limits and inference or API caps are shown separately where published.

Historical v1.3.0 board: weightings, per-task grid, topic radars and held-out diagnostics

The following public-only tables and diagnostics preserve the earlier JevBench v1.3.0 view. The ranking above is the current v1.4.1 result.

媒体内容 · 前往原文查看

JevBench v1.3.0 · 534 decisions per system

JevBench Score (Intelligence, Calibration, Speed, Cost — 25 % each)

OfficialIntelligence above chance, Calibration, Speed, Cost — 25 % each, geometric mean; below 50 Intelligence receives a growing near-chance penalty. Change the weighting ↓

Intel.Calib.SpeedCost$/1k dec.
  1. 1Jev 1.13.074.4I 86 · C 83 · S 83 · K 52 · $0.040
  2. 2SemIf (Qwen3.5-4B)73.1I 79 · C 73 · S 84 · K 59 · ~$0.022 est.
  3. 3djev (Maisa, diffusion-gemma)†73.0I 83 · C 65 · S 91 · K 58 · $0.026 ann.
  4. 4Winnow-12B Q8†71.2I 82 · C 72 · S 82 · K 53 · ~$0.037 est.
  5. 5reflex 4B†70.3I 80 · C 75 · S 68 · K 60 · ~$0.022 est.
  6. 6jqv†68.6I 79 · C 79 · S 75 · K 47 · ~$0.056 est.
  7. 7decision-machine-1†68.3I 62 · C 70 · S 93 · K 54 · $0.035
  8. 8decider-35b-a3b†67.6I 80 · C 72 · S 81 · K 45 · ~$0.067 est.
  9. 9open-alternative-jev (Qwen3.5-4B)†67.0I 64 · C 63 · S 83 · K 60 · ~$0.022 est.
  10. 10system-one-open66.6I 70 · C 57 · S 77 · K 65 · ~$0.015 est.
  11. 11OpenJev (razorback16)66.4I 79 · C 65 · S 83 · K 45 · ~$0.066 est.
  12. 12SimpleJev Qwen3.8-27B†66.3I 85 · C 81 · S 71 · K 39 · ~$0.104 est.
  13. 13ZeroEntropy zerank-2†66.0I 63 · C 76 · S 79 · K 50 · $0.047
  14. 14GPT-5.6 Luna (low)65.9I 95 · C 90 · S 78 · K 28 · $0.242
  15. 15openjev-sglang65.3I 83 · C 77 · S 77 · K 36 · ~$0.131 est.
  16. 16Qwen3-Reranker-4B†63.8I 64 · C 67 · S 79 · K 49 · $0.050
  17. 17reflex-27b†63.3I 86 · C 86 · S 67 · K 32 · ~$0.181 est.
  18. 18LitJev†62.7I 82 · C 84 · S 67 · K 34 · ~$0.163 est.
  19. 19kev 0.6B†62.5I 52 · C 51 · S 76 · K 76 · ~$0.0063 est.
  20. 20SimpleJev Qwen3.6-35B-A3B†62.5I 80 · C 67 · S 75 · K 38 · ~$0.116 est.
  21. 21djev†62.4I 81 · C 93 · S 75 · K 27 · ~$0.274 est.
  22. 22jev-local†61.8I 71 · C 69 · S 69 · K 43 · ~$0.077 est.
  23. 23decider-2b†61.7I 61 · C 47 · S 83 · K 61 · ~$0.020 est.
  24. 24Bespoke Nimble 9B†60.5I 78 · C 65 · S 79 · K 33 · ~$0.166 est.
  25. 25Gemini 3.1 Flash-Lite60.1I 86 · C 68 · S 82 · K 27 · $0.264
  26. 26OpenJev†60.0I 88 · C 70 · S 76 · K 28 · ~$0.255 est.
  27. 27kev 4B†59.7I 65 · C 42 · S 76 · K 62 · ~$0.019 est.
  28. 28DeepSeek V4.1 Flash57.5I 94 · C 97 · S 72 · K 17 · $0.594
  29. 29kev 8B†56.4I 69 · C 44 · S 75 · K 44 · ~$0.073 est.
  30. 30Open-Jev 9B†55.0I 71 · C 63 · S 72 · K 28 · ~$0.249 est.
  31. 31system-one54.8I 70 · C 37 · S 84 · K 41 · ~$0.089 est.
  32. 32jeff†54.4I 47 · C 65 · S 63 · K 77 · ~$0.0060 est.
  33. 33Laya†54.4I 46 · C 62 · S 71 · K 86 · ~$0.0029 est.
  34. 34Open-Jev 2B†51.3I 61 · C 55 · S 73 · K 28 · ~$0.249 est.
  35. 35OpenDecision†40.6I 41 · C 56 · S 80 · K 75 · ~$0.0066 est.
  36. 36openJev Verdict 1.4†38.9I 39 · C 74 · S 78 · K 82 · ~$0.0039 est.
  37. 37openJev Verdict†38.1I 40 · C 51 · S 77 · K 83 · ~$0.0037 est.
  38. 38kev 0.5B†33.2I 38 · C 47 · S 77 · K 76 · ~$0.0063 est.
  39. 39GLiNER2 large†29.6I 40 · C 24 · S 62 · K 73 · ~$0.0077 est.
  40. 40smalljev semantic-v9†27.4I 35 · C 59 · S 80 · K 58 · ~$0.025 est.
  41. 41GLiNER2†24.0I 36 · C 24 · S 72 · K 83 · ~$0.0037 est.
  42. 42open-jev-deberta-v3-large23.1I 32 · C 66 · S 66 · K 74 · ~$0.0073 est.
  43. 43GLiNER2.5 multi†16.6I 28 · C 56 · S 68 · K 82 · ~$0.0039 est.
  44. 44GLiNER2.5 small†13.8I 26 · C 47 · S 78 · K 82 · ~$0.0039 est.
  45. 45Mixedbread mxbai-rerank-base-v2†0.8I 7 · C 83 · S 88 · K 68 · $0.012
  46. 46BAAI bge-reranker-v2-m3†0.7I 6 · C 84 · S 90 · K 73 · $0.0077
  47. 47Alibaba GTE Reranker ModernBERT-base†0.3I 5 · C 77 · S 91 · K 70 · $0.010
  48. 48Certo v1†0.0I 0 · C 82 · S 94 · K 100 · ~$0.0010 est.
  49. classifier.dev (fast tier)† (honorable mention)83.6I 85 · C 78 · S 88 · K 84 · ~$0.0033 est.
  50. Qwen3.8 27B (partial run)24.8I 67 · C 92 · S 61 · K 0 · ~$2.669 est.
  51. Needle 3, options as tools (partial run)1.1I 14 · C – · S 53 · K 65 · ~$0.014 est.
  52. Needle 3 (partial run)0.1I 5 · C – · S 60 · K 59 · ~$0.024 est.
020406080100

Score = Intelligence0.25 × Calibration0.25 × Speed0.25 × Cost0.25 (each 0–100; geometric mean; below 50 Intelligence, × (I / 50)²)

  • Jev (TypeSafe, closed)
  • Jev rebuild (open, or open source planned)
  • Instruction model, JSON schema
  • Small tool-calling model
  • Service built on Jev
  • Zero-shot classifier (not a Jev rebuild)
  • Closed decision model (API only, not Jev)
  • Shown, not ranked — honorable mention (runs another entrant's model) · partial run
⏱ Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.I, C, S, K = Intelligence, Calibration, Speed, Cost; est./ann. = estimated/announced cost; † = see note.Legend and notes
  • ~ est. = no measured bill; priced like a large inference provider (how costs are estimated).
  • ann. = the provider’s announced price, not yet charged.
  • Names link to each project.
  • A label-only system has no calibration (–, counted as 0).
  • djev (Maisa, diffusion-gemma): The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated).
  • Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
  • reflex 4B: The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
  • jqv: A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
  • decision-machine-1: A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
  • decider-35b-a3b: The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
  • open-alternative-jev (Qwen3.5-4B): With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
  • SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • reflex-27b: The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
  • LitJev: The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
  • kev 0.6B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • djev: Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill.
  • jev-local: The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
  • decider-2b: The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
  • Bespoke Nimble 9B: Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
  • OpenJev: OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
  • kev 4B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • kev 8B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • Open-Jev 9B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • jeff: Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
  • Laya: The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
  • Open-Jev 2B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • OpenDecision: A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
  • openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
  • openJev Verdict: The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
  • kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release.
  • GLiNER2 large: The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
  • GLiNER2: A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
  • GLiNER2.5 multi: The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
  • GLiNER2.5 small: The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • Certo v1: The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
  • classifier.dev (fast tier): Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.

Weighting: Intelligence : Calibration : Speed : Cost

Official default

Custom

The official JevBench Score weights the four axes 25 % each and takes their geometric mean. The other buttons are the earlier views (Balanced 33:33:33 and the three “Emphasis on” weightings, which leave Calibration out), recomputed the same way. Any of them is your view, recomputed in your browser from the published axis scores — not the published score.

Explore by task difficulty

All tasks is the published default. Choose a scope to see how the ranking changes by difficulty. Hard only uses all 220 hard-tier decisions and their measured Intelligence, Calibration, Speed and Cost.

Tier mapping: Easy = easy; Medium = standard. Easy scopes change Intelligence only. Hard only measures all four axes on the same hard-tier subset; systems without a hard-tier run are shown as partial and are not ranked.

What changed in the score

A system that is cheap and fast but barely better than guessing could rank high; intelligence is now measured above chance, and systems below half-way get a growing penalty. The tasks, Calibration, Speed, Cost and ranking eligibility are unchanged.

What the run says (JevBench Score)

  • Jev 1.13.0 (TypeSafe AI) leads with 74.4: Intelligence 85.7, Calibration 82.7, Speed 83.3, Cost 52.0 ($0.040 per 1,000 decisions).
  • classifier.dev scores 83.6 — higher than anything in the ranking — but is not ranked: it runs Jev (TypeSafe), so ranking it would put the same model in the list twice, once at the model's own price and once at the service's. It keeps every number it earned under Honorable mentions — services built on another entrant's model.
  • Open rebuilds of Jev appeared within days. The best of them, SemIf, formerly OpenJev (Qwen3.5-4B, TheoLeeCJ), is #2 at 73.1 — 1.3 points behind: more speed and a lower (estimated) price, less intelligence and calibration.
  • GPT-5.6 Luna (low reasoning effort) has the highest Intelligence (95.3) but places #14: its cost score is 28.5 ($0.242 per 1,000 decisions), and the geometric mean does not let accuracy buy that back.
  • Qwen3.8 27B, Needle 3, options as tools, Needle 3 did not answer every tier — each for the reason in its † note; they are shown below the ranking as partial runs, without a rank.

Axes, tiers, latency and cost

Sort by any column; values the run could not produce always sort last. Hover a cost for how it was priced, a latency for the endpoint. Names link to each project.

⏱ Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.

💲

$ per 1,000 decisions

, not per 1,000 tokens — one decision ≈ 950 input tokens.

Rank#SystemEndpoint
1by TypeSafe AIJev 1.13.074.485.782.783.352.0$0.040100.0%99.0%94.5%74.1%0.65 s rawp95 0.72 s rawproduction API
2by Theodore Lee (TheoLeeCJ)SemIfformerly OpenJev (Qwen3.5-4B, TheoLeeCJ73.179.072.683.759.5~$0.022 est.100.0%97.9%95.2%59.5%0.20 s raw→ 0.55 s adjustedp95 0.32 s raw → 0.78 sour RunPod GPU
3by Maisa (David Villalón)djev†Maisa, diffusion-gemma73.082.765.491.457.6$0.026 announced100.0%97.9%93.2%69.5%0.24 s rawp95 0.31 s rawproduction API
4by Eldan RingWinnow-12B Q8†71.282.072.082.352.9~$0.037 est.100.0%96.9%91.1%70.9%0.23 s raw→ 0.60 s adjustedp95 0.41 s raw → 0.98 sour RunPod GPU
5by kshetrajna12reflex 4B†70.380.175.268.059.7~$0.022 est.100.0%94.8%97.3%63.2%1.80 s raw→ 3.75 s adjustedp95 2.05 s raw → 4.26 sour RunPod GPU
6by hjmurmur (Octalab)jqv†Qwen3-32B zero-shot68.679.379.074.647.5~$0.056 est.100.0%95.8%92.5%64.5%0.75 s raw→ 1.64 s adjustedp95 0.97 s raw → 2.10 sour RunPod GPU
7by milliseconds.ai (Baptiste Laget)decision-machine-1†milliseconds.ai68.362.170.492.953.7$0.035100.0%76.0%89.7%46.8%0.17 s rawp95 0.30 s rawproduction API
8by Mapikadecider-35b-a3b†67.679.671.580.845.3~$0.067 est.100.0%96.9%91.1%65.5%0.29 s raw→ 0.73 s adjustedp95 0.49 s raw → 1.14 sour RunPod GPU
9by IkerMoelopen-alternative-jev†Qwen3.5-4B, IkerMoel67.064.063.283.559.6~$0.022 est.100.0%84.4%74.7%56.8%0.21 s raw→ 0.56 s adjustedp95 0.32 s raw → 0.80 sour RunPod GPU
10by mithalounisystem-one-openGemma 4 E2B LoRA on an L466.669.556.777.064.8~$0.015 est.100.0%93.8%87.7%49.1%0.65 s raw→ 1.30 s adjustedp95 0.77 s raw → 1.54 sauthor's demo server
11by razorback16 / CodivOpenJevDiffusionGemma 26B-A4B NVFP4, razorback1666.479.264.883.245.5~$0.066 est.100.0%95.8%91.1%65.5%0.24 s raw→ 0.63 s adjustedp95 0.31 s raw → 0.76 sour RunPod GPU
12by Featherless AISimpleJev Qwen3.8-27B†66.384.781.171.239.5~$0.104 est.100.0%96.9%93.2%75.0%1.01 s raw→ 2.03 s adjustedp95 1.88 s raw → 3.76 sauthor's demo server
13by ZeroEntropyZeroEntropy zerank-2†66.063.076.579.049.8$0.047100.0%79.2%88.4%47.3%0.13 s raw→ 0.40 s adjustedp95 1.50 s raw → 3.15 sour RunPod GPU
14by OpenAIGPT-5.6 Lunalow reasoning effort65.995.389.877.528.5$0.242100.0%97.9%96.6%94.5%0.97 s rawp95 1.82 s rawproduction API
15by ekzhangopenjev-sglangQwen3.6-35B-A3B on SGLang65.383.477.477.136.5~$0.131 est.100.0%95.8%95.2%71.4%0.68 s raw→ 1.36 s adjustedp95 0.73 s raw → 1.45 sauthor's demo server
16by QwenQwen3-Reranker-4B†63.864.067.078.749.2$0.050100.0%79.2%87.7%50.0%0.13 s raw→ 0.41 s adjustedp95 1.56 s raw → 3.27 sour RunPod GPU
17by kshetrajna12reflex-27b†Qwen3.8-27B63.385.886.267.532.3~$0.181 est.100.0%95.8%95.9%75.9%1.89 s raw→ 3.93 s adjustedp95 2.21 s raw → 4.57 sour RunPod GPU
18by Zhengxu YuLitJev†Qwen3.8-27B62.782.483.566.733.6~$0.163 est.100.0%97.9%88.4%73.2%2.03 s raw→ 4.20 s adjustedp95 2.46 s raw → 5.06 sour RunPod GPU
19by Jared Palmerkev 0.6B†research preview62.551.951.175.676.1~$0.0063 est.100.0%81.3%66.4%40.0%0.59 s raw→ 1.33 s adjustedp95 0.97 s raw → 2.09 sour RunPod GPU
20by Featherless AISimpleJev Qwen3.6-35B-A3B†62.579.567.175.038.1~$0.116 est.100.0%93.8%93.2%66.4%0.85 s raw→ 1.70 s adjustedp95 0.93 s raw → 1.86 sauthor's demo server
21by David Villalon / Maisadjev†thinking62.480.892.775.226.9~$0.274 est.95.8%99.0%80.1%77.7%0.43 s raw→ 1.00 s adjustedp95 1.45 s raw → 3.05 sour RunPod GPU
22by us (GitHub)jev-local†Qwen3.5-9B61.870.868.769.243.3~$0.077 est.100.0%84.4%89.0%59.1%1.05 s raw→ 2.24 s adjustedp95 2.62 s raw → 5.38 sour RunPod GPU
23by Mapikadecider-2b†61.761.246.683.261.0~$0.020 est.100.0%85.4%77.4%47.3%0.26 s raw→ 0.67 s adjustedp95 0.28 s raw → 0.72 sour RunPod GPU
24by Bespoke LabsBespoke Nimble 9B†60.577.965.378.733.4~$0.166 est.100.0%94.8%89.0%65.5%0.39 s raw→ 0.93 s adjustedp95 0.65 s raw → 1.46 sour RunPod GPU
25by GoogleGemini 3.1 Flash-Lite60.185.668.181.827.4$0.264100.0%99.0%93.2%75.0%0.76 s rawp95 0.88 s rawproduction API
26by razorback16OpenJev†thinking, BF1660.088.069.676.127.8~$0.255 est.100.0%100.0%94.5%78.2%0.46 s raw→ 1.08 s adjustedp95 1.08 s raw → 2.31 sour RunPod GPU
27by Jared Palmerkev 4B†research preview59.764.842.075.761.8~$0.019 est.100.0%91.7%85.6%42.3%0.55 s raw→ 1.25 s adjustedp95 0.99 s raw → 2.13 sour RunPod GPU
28by DeepSeekDeepSeek V4.1 Flashthinking default57.594.396.771.616.8$0.59498.6%99.0%93.2%95.0%1.42 s rawp95 4.89 s rawproduction API
29by Jared Palmerkev 8B†research preview56.469.444.274.944.0~$0.073 est.100.0%92.7%90.4%47.3%0.59 s raw→ 1.33 s adjustedp95 1.15 s raw → 2.45 sour RunPod GPU
30by Zefan Cai (@Zefan_Cai)Open-Jev 9B†Zefan Cai55.071.263.372.028.1~$0.249 est.100.0%90.6%81.5%60.9%0.75 s raw→ 1.66 s adjustedp95 1.81 s raw → 3.77 sour RunPod GPU
31by Sean Goedeckesystem-oneQwen3-8B, Sean Goedecke54.870.336.884.441.5~$0.089 est.100.0%90.6%91.8%50.0%0.17 s raw→ 0.48 s adjustedp95 0.30 s raw → 0.76 sour RunPod GPU
32by Logan Markewichjeff†Logan Markewich, GLiFormer 400M54.446.964.663.576.6~$0.0060 est.100.0%76.0%61.6%37.7%0.94 s raw→ 2.03 s adjustedp95 10.97 s raw → 22.09 sour CPU
33by Convai InnovationsLaya†Convai Innovations, ModernBERT-large 421M54.445.862.571.186.2~$0.0029 est.94.4%72.9%69.2%34.1%0.79 s raw→ 1.72 s adjustedp95 2.20 s raw → 4.54 sour CPU
34by Zefan Cai (@Zefan_Cai)Open-Jev 2B†Zefan Cai51.361.055.173.528.1~$0.249 est.100.0%79.2%88.4%42.7%0.66 s raw→ 1.48 s adjustedp95 1.45 s raw → 3.05 sour RunPod GPU
35by Deepan WadhwaOpenDecision†ModernBERT-large zero-shot40.640.856.179.975.3~$0.0066 est.87.5%62.5%71.2%33.2%0.34 s raw→ 0.83 s adjustedp95 0.54 s raw → 1.24 sour RunPod GPU
36by Hemant (heman10x)openJev Verdict 1.4†38.938.674.178.182.4~$0.0039 est.86.1%67.7%56.2%37.7%0.31 s raw→ 0.78 s adjustedp95 0.92 s raw → 2.00 sour CPU
37by Hemant (heman10x)openJev Verdict†heman10x, ModernBERT-base 151M38.139.851.376.783.1~$0.0037 est.86.1%65.6%61.0%38.2%0.28 s raw→ 0.71 s adjustedp95 1.45 s raw → 3.04 sour CPU
38by Jared Palmerkev 0.5B†33.238.247.477.076.1~$0.0063 est.95.8%52.1%71.2%30.9%0.43 s raw→ 1.01 s adjustedp95 0.92 s raw → 1.99 sour RunPod GPU
39by FastinoGLiNER2 large†29.640.124.361.773.3~$0.0077 est.98.6%62.5%61.0%36.4%1.10 s raw→ 2.34 s adjustedp95 14.49 s raw → 29.13 sour CPU
40by Aditya (isHeSatoshi)smalljev semantic-v9†27.435.158.979.857.9~$0.025 est.97.2%68.8%40.4%38.2%0.41 s raw→ 0.98 s adjustedp95 0.46 s raw → 1.07 sour RunPod GPU
41by FastinoGLiNER2†Fastino, gliner2.5-base24.035.623.771.883.1~$0.0037 est.97.2%66.7%45.9%36.4%0.31 s raw→ 0.78 s adjustedp95 4.15 s raw → 8.46 sour CPU
42by Kotoba Labsopen-jev-deberta-v3-largelocal CPU23.131.966.466.074.0~$0.0073 est.100.0%49.0%53.4%36.4%1.77 s raw→ 3.69 s adjustedp95 3.35 s raw → 6.85 sour CPU
43by FastinoGLiNER2.5 multi†Fastino, 287M16.627.756.167.882.4~$0.0039 est.90.3%51.0%43.8%37.7%0.43 s raw→ 1.01 s adjustedp95 8.18 s raw → 16.50 sour CPU
44by FastinoGLiNER2.5 small†Fastino, 74M13.825.647.277.882.4~$0.0039 est.83.3%47.9%50.0%33.2%0.11 s raw→ 0.38 s adjustedp95 2.10 s raw → 4.35 sour CPU
45by MixedbreadMixedbread mxbai-rerank-base-v2†0.86.783.187.567.9$0.01244.4%33.3%26.7%40.0%0.07 s raw→ 0.29 s adjustedp95 0.23 s raw → 0.62 sour RunPod GPU
46by BAAIBAAI bge-reranker-v2-m3†0.76.383.889.573.4$0.007743.1%36.5%8.9%36.8%0.03 s raw→ 0.22 s adjustedp95 0.18 s raw → 0.51 sour RunPod GPU
47by Alibaba-NLPAlibaba GTE Reranker ModernBERT-base†0.34.676.890.669.6$0.01033.3%39.6%30.1%33.6%0.05 s raw→ 0.25 s adjustedp95 0.10 s raw → 0.35 sour RunPod GPU
48by AltSlate LabsCerto v1†0.00.082.094.0100.0~$0.0010 est.27.8%30.2%21.9%31.8%0.02 s raw→ 0.19 s adjustedp95 0.03 s raw → 0.21 sour RunPod GPU
Honorable mentions — services built on another entrant's model — shown, not ranked: A service that runs another entrant's model is listed with all of its scores and axes, but is not ranked against the models.
by mrmps (@michael_chomsky)classifier.dev†fast tierhonorable mention · not ranked83.685.177.987.684.3~$0.0033 est.100.0%99.0%97.3%70.5%0.39 s rawp95 0.45 s rawproduction API
Partial runs — shown, not ranked: a tier attempted for fewer than 95 % of its decisions.
by Qwen / ChutesQwen3.8 27BChutes TEEpartial run · not ranked24.867.492.161.30.0~$2.669 est.98.6%99.0%95.3%21.4%5.75 s rawp95 12.97 s rawproduction API
by Cactus ComputeNeedle 3, options as toolspost-hoc adapter modepartial run · not ranked1.113.5none (label only)52.865.3~$0.014 est.66.7%31.3%34.2%—3.78 s raw→ 7.71 s adjustedp95 33.64 s raw → 67.42 sour CPU
by Cactus ComputeNeedle 3Cactus, 2-bit, local CPUpartial run · not ranked0.14.6none (label only)59.958.7~$0.024 est.47.2%16.7%31.5%7.7%1.69 s raw→ 3.52 s adjustedp95 14.36 s raw → 28.88 sour CPU

† Notes on 39 marked systems — how each was run

  • † djev: The measured endpoint was Maisa's hosted API in free preview; the cost uses its announced price ($0.035 per million input tokens, output free), and nothing was charged. The self-hostable djev-dev runtime is Apache-2.0 and applies a structured one-step inference method to Google's Apache-2.0 diffusiongemma-26B-A4B-it checkpoint; it adds no separately trained djev weights. Probabilities are djev's own (its docs call them experimental and uncalibrated).
  • † Winnow-12B Q8: The submitted Q8_0 GGUF ran through the pinned author's TypeSafe-compatible /v1/systemone server with 8,192 context, four resident decision branches, Q8 KV, and full GPU offload. The private training corpus was not released. The author's checksum-based audit reports zero exact public-item overlap, but that claim cannot be independently reproduced; our scan found no exact public state or instruction text in the released artifacts. Cost uses the $0.05/M-input hosted Gemma 3 12B reference, not free/100.
  • † reflex 4B: The author's reflex-serve: Qwen3.5-4B with the published LoRA and its per-primitive calibration file; the state is encoded once and each question read from the label logits. Run serially on our GPU; the author discloses that the 231 public items were used four times as a development gate.
  • † jqv: A stock Qwen3-32B with no decision training: the state is prefilled once, each question is an isolated branch and the answer is read from the option-letter logits, with one fitted temperature (3.02, 400 MMLU validation items). Re-run in v1.2.8 on our own GPU from the now-public serving code (Octalab-Inc/jqv 0189b67), so all 534 decisions including the held-out hard items were asked; this full run replaces the v1.2.7 partial row, which had been measured on the submitter's machine. Cost is the base model's public per-token tariff, not free.
  • † decision-machine-1: A closed-weights decision model behind a production API that serves TypeSafe's wire format, so the unchanged typesafe adapter ran it. Run on a free test key (30 requests a minute, 2.2 s between requests); the provider states the inference infrastructure is the same as for paid keys. Cost is the public paid tariff, $0.04 per million input tokens (output free), times the input tokens the API reported.
  • † decider-35b-a3b: The author's TypeSafe-compatible server and published FP8 weights, run serially on our H100 NVL. The exhaustive startup batch warmup was skipped; each required serial shape captured lazily before its measured request. Self-host latency receives the standard ×2 + 0.15 s adjustment. Cost uses the closest hosted 35B-A3B input tariff and is not the temporary rental charge.
  • † open-alternative-jev: With the options in reverse order (A. no, B. yes) the same model scored 21 % instead of 72 % on yes/no answer-judging items — small models are very sensitive to option order.
  • † SimpleJev Qwen3.8-27B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • † ZeroEntropy zerank-2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † Qwen3-Reranker-4B: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † reflex-27b: The frozen public Qwen3.8-27B checkpoint through reflex at the requested pinned commit, with two option orders averaged and temperature 1. No adapter or fitted calibration file. Run serially on our H100 NVL. Self-host latency receives the standard ×2 + 0.15 s adjustment; cost uses the exact base model's public hosted input tariff.
  • † LitJev: The author's reproduction of Jev's decision layer on an off-the-shelf model, in its default configuration: Qwen3.8-27B, scores read from the output head, no training and no calibration file (its README says probabilities are not calibrated by default). Run serially on our GPU through an SSH tunnel, because its server binds to localhost; the request still crosses the internet and gets the ×2 + 0.15 s adjustment.
  • † kev 0.6B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • † SimpleJev Qwen3.6-35B-A3B: Author's no-login shared demo, model id recorded verbatim, one request at a time at or below its 2 RPS limit. SimpleJev reads answer-token logits and returns the complete distribution; it does not generate an answer. Speed uses the public-demo x2 load adjustment; cost uses a hosted size-class input price and is not free/100.
  • † djev: Experimental full-generation path over the same DiffusionGemma checkpoint as djev-dev: thinking was enabled and the model could generate up to 8,192 tokens before returning its distribution. Current djev-dev itself hard-codes enable_thinking=false, diffusion_max_steps=1 and read_only=true, so this is not a switch in its published typed API. It is substantially slower/costlier, and 72/534 requests exhausted the output budget without a parseable distribution; those are failures. Cost uses measured tokens and a same-size hosted reference, not the H200 rental bill.
  • † jev-local: The author's local Jev-compatible server in its default full configuration: a frozen Qwen3.5-9B scores each option by its mean log-probability (one forward pass per option, no generation, no decision training). Run serially on our GPU. It re-reads the state once per option; if its reported token count covers one pass only, a per-token hosted price would be higher than this estimate.
  • † decider-2b: The author's TypeSafe-compatible server and published weights (Qwen3.5-2B-Base with a trained one-pass decision readout), run serially on our GPU. Self-host latency gets the standard ×2 + 0.15 s adjustment.
  • † Bespoke Nimble 9B: Re-run in v1.2.8 at Bespoke Labs' request after they raised the serving prompt limit from 2,048 to 8,192 tokens (bespokelabsai/nimble PR #4). Same recipe as the v1.1.3 run — the published LoRA merged into Qwen3.5-9B with the author's PEFT safe-merge, served with SGLang and the author's Jev-compatible API — now from current nimble main; the adapter weights are unchanged. Hard-tier accuracy rose from 43.6 % to 65.5 %, yet the score fell: the long hard items that used to fail at once are now answered and priced (so Cost fell), and this pod was in Canada while the v1.1.3 run's was in Sweden, so part of the lower Speed is network distance from our server in Germany. This complete run replaces the earlier row; its old score is kept in the artifact under superseded_rows.
  • † OpenJev: OpenJev's real typed-API thinking switch at think=512, using its own /v1/systemone server over BF16 DiffusionGemma. The thought is generated first, then native probability reads are taken after it. All 534 requests returned valid distributions. Cost counts the server's billed input and thought output tokens.
  • † kev 4B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • † kev 8B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. The author labels this checkpoint a research preview.
  • † Open-Jev 9B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • † jeff: Self-hosted from its GitHub repo with server defaults, on our CPU (the author recommends a GPU, e.g. an L4), through the same TypeSafe-compatible API as Jev.
  • † Laya: The English checkpoint (repo root), run on our CPU through its own `laya` package. Its budget is 512 tokens per question, so long hard-tier states are cut by the package itself.
  • † Open-Jev 2B: The author's pinned LoRA adapter, trained scalar decision head and calibration temperature, served by the author's Open-Jev server with prefix caching off, batch size 1 and 4,096-token limit. Serial requests were measured from Sandy over an SSH tunnel to the H100. Self-host latency receives the standing x2 + 0.15 s adjustment. Cost uses the exact Qwen3.5-9B hosted input tariff for 9B and the same conservative same-family proxy for the unlisted 2B; neither receives an automatic 100. Exact normalized comparison found no JevBench public task state or instruction in the 79,116-row public training projection.
  • † OpenDecision: A zero-shot NLI classifier behind a TypeSafe-compatible server, not a trained decision model: it scores each option as an entailment hypothesis with ModernBERT-large-zeroshot-v2.0. Its choice path runs several NLI passes over the same state, which the reported token count does not include, so a per-token hosted price would be higher than the estimate here. Pre-registered for our CPU in v1.2.7, run on our GPU because the CPU was far too slow.
  • † openJev Verdict 1.4: Same public weights as the earlier Verdict row, run through the author's fixed v1.4 engine. That engine auto-loads the calibrator for every option count, frames candidate labels as NLI sentences and uses a 512-token context budget. Run locally on our CPU, serially.
  • † openJev Verdict: The openJev-verdict-2.0 Hugging Face repo ships no weights; its config is byte-identical to heman10x/rlcd-modernbert-151m, whose published weights we ran with the author's engine. The 'verdict2-base' checkpoint behind the README's numbers is not downloadable yet (Git LFS 404); we will run it once it is.
  • † kev 0.5B: Self-hosted from the author's repository at commit 20fa626 through its native TypeSafe-compatible `/v1/systemone` server, BF16 on an RTX 3090; measured serially from Sandy over the internet. This is the v0.1 release.
  • † GLiNER2 large: The large checkpoint of Fastino's earlier GLiNER2 family, same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • † smalljev semantic-v9: The public semantic-v9 LoRA and native heads over MiniCPM5-2B-Base, through the mapping frozen before the run. It has a typed Python contract but no TypeSafe-compatible HTTP route. The released training recipe explicitly hill-climbed against JevBench's public shape and source families; this allowed public benchmark-directed development is disclosed. Cost is $0.04/M measured input tokens, not free/100.
  • † GLiNER2: A general schema classifier, not a Jev rebuild. The question goes in front of the text; the probabilities are GLiNER2's own single-label softmax over the labels, read out in full (mapping fixed before the run).
  • † GLiNER2.5 multi: The multilingual GLiNER2.5 checkpoint (287M), same family and same documented mapping as the GLiNER2 row. JevBench items are English only, so its multilingual training is not exercised here.
  • † GLiNER2.5 small: The small GLiNER2.5 checkpoint (74M), same family and same documented mapping as the GLiNER2 row: the question goes in front of the text and the probabilities are the model's own single-label softmax over the labels, read out in full. A general schema classifier, not a Jev rebuild.
  • † Mixedbread mxbai-rerank-base-v2: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † BAAI bge-reranker-v2-m3: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † Alibaba GTE Reranker ModernBERT-base: Neutral documented reranker adapter; instruction and no-instruction public calibration were run, then frozen before one held-out pass.
  • † Certo v1: The public Certo v1 checkpoint through the author's DecisionModel, serially on our rented GPU. The question instruction is prepended to the state because Certo exposes state + runtime options but no separate question field; the published 64-token state and 48-token option limits are unchanged. The model card says v1 does not yet transfer to arbitrary natural-language prose. Cost is an estimate from same-size hosted encoders times the checkpoint's retained input tokens, not free/100.
  • † classifier.dev: Its own benchmark page says the fast tier is Jev. Free for us; the price is its published Pro plan ($20/month for 200,000 fast classifications a day) at full use, $0.0033 per 1,000 decisions.

Honorable mentions — services built on another entrant's model

A service that runs another entrant's model is listed with all of its scores and axes, but is not ranked against the models.

classifier.devno rank

83.6 JevBench Score · Jev 1.13.0 (#1) scores 74.4

Runs on Jev (TypeSafe).

  • Intelligence85.1
  • Calibration77.9
  • Speed87.6
  • Cost84.3
  • $ per 1,000 decisions~$0.0033 est.

Runs on Jev (TypeSafe) — listed, not ranked. Ranking it against Jev would rank Jev's model against Jev's model, so from v1.2.4 it is an honorable mention instead of #1.

Why it is not ranked, what its price assumes, and what we found

classifier.dev is not its own model. Its own pages say so: "The fast tier is Jev, TypeSafe's decision model" (https://classifier.dev/benchmark, read 2026-09-20), and the API answers with "model": "jev-1.13.0" — the same model version this benchmark measures directly as Jev 1.13.0. What it adds is a price and, on its smart tier, an orchestration layer: "The smart tier is Jev plus a reasoning model re-asking only the answers Jev put under 0.7 confidence" — escalation on low confidence (a model cascade), not best-of-N, not self-consistency and not a committee. Its published escalation model is gemini-3.8-flash. Ranking it against Jev would rank Jev's model against Jev's model, so from v1.2.4 it is an honorable mention instead of #1.

Only the fast tier was measured. The smart tier's escalation was never run, so nothing here scores it.

Price. $0.0033 per 1,000 decisions is an estimate from the published flat-rate plan at full use: classifier.dev Pro is $20/month for 200,000 fast classifications a day (https://classifier.dev/pricing, read 2026-09-20), and one classification is one decision. Lower use costs more per decision — at a tenth of that allowance it is $0.033 per 1,000 — and the free tier (20,000 fast classifications a day), which is what our run used, costs nothing. Their pages do not say how the flat rate is funded, so we do not know their cost basis; the only figure they publish is what the model costs a caller: "The model behind the fast tier costs about $0.005 per thousand classifications and needs a TypeSafe key" (https://classifier.dev/pricing) — for their short single-sentence inputs, not for JevBench's whole questions.

Not a pass-through. On our set the fast tier scored 97.3 % on the judge tier against Jev's 94.5 %, and 70.5 % against 74.1 % on the hard tier. classifier.dev's own explanation for differences of this kind is batching ("The fast tier is Jev, packed a thousand to a request"); on their own two test sets they measured the same difference as noise.

A legitimate, well-documented product: free without an account, open source (https://github.com/mrmps/classifier-dev), by Michael Ryaboy (@michael_chomsky). Read 2026-09-20: classifier.dev · classifier.dev/benchmark · classifier.dev/pricing · classifier.dev/about

Which public tasks did each system get right?

This view shows public task outcomes only: 231 of 231 public tasks in the selected scope. Held-out and imported task text is not shipped.

Show 231 public task outcomes across 52 systems

TaskJev 1.13.0SemIfdjevWinnow-12B Q8reflex 4Bjqvdecision-machine-1decider-35b-a3bopen-alternative-jevsystem-one-openOpenJevSimpleJev Qwen3.8-27BZeroEntropy zerank-2GPT-5.6 Lunaopenjev-sglangQwen3-Reranker-4Breflex-27bLitJevkev 0.6BSimpleJev Qwen3.6-35B-A3Bdjevjev-localdecider-2bBespoke Nimble 9BGemini 3.1 Flash-LiteOpenJevkev 4BDeepSeek V4.1 Flashkev 8BOpen-Jev 9Bsystem-onejeffLayaOpen-Jev 2BOpenDecisionopenJev Verdict 1.4openJev Verdictkev 0.5BGLiNER2 largesmalljev semantic-v9GLiNER2open-jev-deberta-v3-largeGLiNER2.5 multiGLiNER2.5 smallMixedbread mxbai-rerank-base-v2BAAI bge-reranker-v2-m3Alibaba GTE Reranker ModernBERT-baseCerto v1classifier.devQwen3.8 27BNeedle 3, options as toolsNeedle 3
Easy · 48 of 72 decisions publicEasy · 48 of 72 public48/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4846/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4848/4846/4848/4842/4842/4841/4846/4848/4847/4847/4848/4844/4841/4822/4823/4817/4812/4848/4848/4832/4823/48
easy-intent-00intent-00choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××××✓✓✓×
easy-intent-01intent-01choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××××✓✓✓✓
easy-intent-02intent-02choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××××✓✓××
easy-intent-03intent-03choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓✓×
easy-intent-04intent-04choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓××
easy-intent-05intent-05choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××××✓✓✓×
easy-intent-06intent-06choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓×✓✓✓×
easy-intent-07intent-07choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓××✓✓✓✓
easy-intent-08intent-08choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××××✓✓✓×
easy-intent-09intent-09choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓✓××✓✓✓✓×
easy-intent-10intent-10choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓✓×
easy-intent-11intent-11choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓✓✓
easy-fact-00fact-00noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓✓
easy-fact-01fact-01noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×××✓✓✓×✓××✓×✓×✓✓✓✓
easy-fact-02fact-02noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓××✓✓✓✓
easy-fact-03fact-03noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓✓✓××✓×✓×✓✓✓×
easy-fact-04fact-04noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓✓✓×✓
easy-fact-05fact-05noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××××✓✓✓✓✓✓✓×✓×✓✓××
easy-fact-06fact-06noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓✓
easy-fact-07fact-07noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓✓✓✓✓✓×✓×✓✓××
easy-fact-08fact-08noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓✓
easy-fact-09fact-09noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓✓✓×✓✓×✓×✓✓××
easy-fact-10fact-10noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓××
easy-fact-11fact-11noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓✓✓✓✓✓×✓×✓✓××
easy-extraction-00extraction-00choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××××✓✓×✓
easy-extraction-01extraction-01choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓×✓
easy-extraction-02extraction-02choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓×✓
easy-extraction-03extraction-03choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓✓✓
easy-extraction-04extraction-04choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××××✓✓×✓
easy-extraction-05extraction-05choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓
easy-extraction-06extraction-06choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓×××✓✓✓✓
easy-extraction-07extraction-07choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓
easy-extraction-08extraction-08choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓
easy-extraction-09extraction-09choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓×✓
easy-extraction-10extraction-10choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓××✓✓×✓
easy-extraction-11extraction-11choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓✓✓
easy-tool_selection-00tool_selection-00choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓✓✓
easy-tool_selection-01tool_selection-01choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓✓×
easy-tool_selection-02tool_selection-02choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓×
easy-tool_selection-03tool_selection-03choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓!✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓×
easy-tool_selection-04tool_selection-04choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓✓✓×
easy-tool_selection-05tool_selection-05choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓×
easy-tool_selection-06tool_selection-06choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓××✓✓✓×
easy-tool_selection-07tool_selection-07choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓!✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓××✓✓✓✓
easy-tool_selection-08tool_selection-08choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓××
easy-tool_selection-09tool_selection-09choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓×
easy-tool_selection-10tool_selection-10choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓×
easy-tool_selection-11tool_selection-11choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓×
Medium (standard) · 72 of 96 decisions publicMedium (standard) · 72 of 96 public71/7271/7271/7269/7268/7269/7254/7270/7260/7267/7270/7270/7257/7270/7268/7254/7269/7271/7258/7267/7271/7260/7261/7267/7271/7272/7264/7271/7267/7265/7264/7254/7250/7255/7243/7250/7245/7235/7242/7249/7246/7231/7232/7230/7224/7226/7226/7224/7271/7271/7219/7212/72
original-policy-01-0original-policy-01-0noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓✓✓×✓×✓✓✓✓××
original-policy-01-1original-policy-01-1noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓!✓✓✓✓✓✓✓××✓✓✓✓✓××✓×✓×✓✓××
original-policy-02-0original-policy-02-0noul✓✓✓✓✓✓×✓✓✓✓✓×✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓××✓✓×✓✓✓✓××✓✓×✓
original-policy-02-1original-policy-02-1noul✓×✓✓×✓×✓×✓✓✓×✓✓××✓×✓✓××✓✓✓✓✓✓✓×✓✓✓×✓✓××✓××✓✓✓✓××✓✓××
original-policy-03-0original-policy-03-0noul✓✓✓✓✓✓×✓✓×✓✓✓✓×✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓×××✓✓××✓×××✓×✓××✓✓××
original-policy-03-1original-policy-03-1noul✓✓✓✓✓××✓×✓✓✓×✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓×✓✓××××××××✓××✓✓××
original-policy-04-0original-policy-04-0noul✓✓✓✓✓✓×✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓××✓✓✓✓✓××✓×✓×✓✓××
original-policy-04-1original-policy-04-1noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓××✓✓✓×✓××✓×✓✓✓✓××
original-policy-05-0original-policy-05-0noul✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓××✓✓✓✓✓××✓×✓✓✓✓××
original-policy-05-1original-policy-05-1noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓✓✓×✓×✓✓✓✓××
original-policy-06-0original-policy-06-0noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓××✓××✓✓×✓××✓✓××
original-policy-06-1original-policy-06-1noul×✓✓×✓✓✓✓✓✓✓×✓××✓×✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓×✓✓×✓✓×✓×××✓××
original-intent-01-0original-intent-01-0choice✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓×✓×✓××✓×✓×✓✓✓✓×
original-intent-01-1original-intent-01-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓×✓✓✓✓×✓✓✓✓×××✓✓✓×✓×✓××✓✓×✓
original-intent-02-0original-intent-02-0choice✓✓✓✓✓✓×✓×✓✓✓✓✓✓×✓✓×✓✓×✓✓✓✓✓✓×✓×✓××✓×××××××××××××✓✓✓×
original-intent-02-1original-intent-02-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓×××✓✓✓✓×✓××××✓✓✓×
original-intent-03-0original-intent-03-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓××××✓✓×××××✓✓××
original-intent-03-1original-intent-03-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓✓✓×
original-intent-04-0original-intent-04-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓✓✓×
original-intent-04-1original-intent-04-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓××✓✓×✓✓×✓××××✓✓✓×
original-intent-05-0original-intent-05-0choice✓✓✓✓✓✓×✓×✓✓✓×✓✓✓✓✓✓✓✓×✓×✓✓✓✓✓✓✓××××✓✓✓××✓×✓×✓×✓×✓✓××
original-intent-05-1original-intent-05-1choice✓✓✓✓✓✓×✓×✓✓✓×✓✓×✓✓✓✓✓×✓×✓✓✓✓✓✓✓××××✓✓✓✓×✓×✓×✓×✓×✓✓××
original-intent-06-0original-intent-06-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓××✓✓×✓×✓××✓××××✓✓✓×
original-intent-06-1original-intent-06-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓××✓××✓×✓×✓✓××××✓✓✓×
original-ordinal-01-0original-ordinal-01-0score✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓××✓××✓✓××
original-ordinal-01-1original-ordinal-01-1score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓××✓×✓✓✓××
original-ordinal-02-0original-ordinal-02-0score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×××××✓✓××
original-ordinal-02-1original-ordinal-02-1score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×××××✓✓××
original-ordinal-03-0original-ordinal-03-0score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓×✓×××✓✓✓××
original-ordinal-03-1original-ordinal-03-1score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓✓×✓✓×✓×××✓✓✓××
original-ordinal-04-0original-ordinal-04-0score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓×✓✓✓×✓×✓✓××
original-ordinal-04-1original-ordinal-04-1score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓×✓✓×✓×✓✓××
original-ordinal-05-0original-ordinal-05-0score✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓×✓!✓✓✓✓✓✓✓✓✓✓✓✓✓××××✓✓✓××××✓×✓✓✓××
original-ordinal-05-1original-ordinal-05-1score✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓×✓✓✓×✓✓✓✓✓✓✓✓✓×✓✓✓✓×✓✓✓×✓××✓×✓✓✓××
original-ordinal-06-0original-ordinal-06-0score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓✓✓××
original-ordinal-06-1original-ordinal-06-1score✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓××
original-extraction-01-0original-extraction-01-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓×✓✓✓××××××✓×××××××✓✓✓××
original-extraction-01-1original-extraction-01-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓×✓✓✓✓✓✓✓×××××✓×✓×××××✓✓✓×✓
original-extraction-02-0original-extraction-02-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓×✓✓✓×✓✓××
original-extraction-02-1original-extraction-02-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓×✓×✓✓✓✓✓×✓✓××
original-extraction-03-0original-extraction-03-0choice✓✓✓✓✓✓✓✓✓×✓✓×✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓×✓×✓×✓××××××✓✓×✓
original-extraction-03-1original-extraction-03-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓×✓××××✓✓××
original-extraction-04-0original-extraction-04-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓××✓✓✓×××××✓✓××
original-extraction-04-1original-extraction-04-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××✓×✓×✓××××✓✓✓✓
original-extraction-05-0original-extraction-05-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓×✓×✓✓✓✓✓×✓✓××
original-extraction-05-1original-extraction-05-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓×✓×✓✓✓✓✓×✓✓××
original-extraction-06-0original-extraction-06-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓✓××××✓✓✓✓
original-extraction-06-1original-extraction-06-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓××××××✓✓×✓
original-adequacy-01-0original-adequacy-01-0noul✓✓✓×✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓✓×✓✓×××✓✓✓×✓✓✓✓✓
original-adequacy-01-1original-adequacy-01-1noul✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓×××✓×✓✓×✓✓✓✓✓
original-adequacy-02-0original-adequacy-02-0noul✓✓✓✓✓✓✓✓×✓✓✓×✓✓✓✓✓×✓✓✓✓✓✓✓×✓×✓✓××✓×××✓×✓✓✓××××✓✓✓✓××
original-adequacy-02-1original-adequacy-02-1noul✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓✓✓×××✓✓✓✓✓✓
original-adequacy-03-0original-adequacy-03-0noul✓✓✓✓×××✓×✓✓✓×✓✓×✓✓××✓××××✓×✓×✓××××××××××✓×××××✓×✓✓××
original-adequacy-03-1original-adequacy-03-1noul✓✓×✓✓✓×✓×✓×✓×✓×××✓××✓×××✓✓×✓×✓××✓×××××××✓×××××✓×✓✓××
original-adequacy-04-0original-adequacy-04-0noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓××✓××××✓✓×✓✓✓××
original-adequacy-04-1original-adequacy-04-1noul✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓✓××✓××××✓✓×✓✓✓××
original-adequacy-05-0original-adequacy-05-0noul✓✓✓✓✓✓✓××✓✓✓×✓××✓✓××✓×××✓✓×✓××××✓××××✓×✓✓✓××××✓×✓✓××
original-adequacy-05-1original-adequacy-05-1noul✓✓✓✓×××✓×✓✓✓×✓✓✓✓✓×✓✓××✓✓✓×✓✓×××✓✓×××××✓✓×××××✓×✓✓××
original-adequacy-06-0original-adequacy-06-0noul✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓×××✓×✓✓×✓✓✓✓✓
original-adequacy-06-1original-adequacy-06-1noul✓✓✓✓✓✓✓✓✓✓✓×✓×✓✓✓✓✓×✓✓×✓✓✓✓✓✓✓✓✓×××✓✓×✓✓××✓✓✓✓×✓✓✓✓✓
original-routing-01-0original-routing-01-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓××✓✓✓×
original-routing-01-1original-routing-01-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×××××✓××××××××✓✓××
original-routing-02-0original-routing-02-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓××××✓×✓✓××
original-routing-02-1original-routing-02-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓×✓××××✓×✓✓××
original-routing-03-0original-routing-03-0choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓×✓✓✓✓✓✓✓×✓✓××
original-routing-03-1original-routing-03-1choice✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓×✓✓✓✓✓✓✓×✓✓××
original-routing-04-0original-routing-04-0choice✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓✓✓✓✓✓×✓✓××✓✓✓×✓✓✓×××××××✓✓✓×
original-routing-04-1original-routing-04-1choice✓✓✓✓×✓×✓×✓×✓✓✓✓✓✓×✓✓✓×✓✓✓✓✓✓✓××✓×××✓✓×✓×××××××××✓✓✓×
original-routing-05-0original-routing-05-0choice✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓×✓✓✓✓✓✓×✓✓×✓×✓✓××✓××××××××✓✓××
original-routing-05-1original-routing-05-1choice✓✓✓✓✓✓×✓××✓✓✓✓✓×✓✓✓✓✓××✓✓✓✓✓✓×✓✓××✓×××××××××××××✓✓××
original-routing-06-0original-routing-06-0choice✓✓✓✓✓✓×✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓✓××✓×✓×××××××✓✓!✓×
original-routing-06-1original-routing-06-1choice✓✓✓✓✓✓×

来源:Hacker News 热门(buzzing.cc 中文翻译)· benchmarkheaven.com