JevBench
3 storiesJevBench is a reproducible benchmark for typed decision models that return bounded choices and probabilities, scoring accuracy, latency and cost across 534 English decisions with configurable weights. Recent coverage includes Jeeves, a Jev-like model built on Qwen3.5-9B with LoRA, a pointer head, CISPO training and a block-4 diffusion drafter, which scored 0.889 on unseen test data and 0.935 on JevBench's public hard tier; earlier Jev led the board at 74.4, ahead of SemIf, djev, Winnow-12B Q8 and reflex 4B.
Related topics
A developer released Jeeves, a Jev-like decision model built on Qwen3.5-9B with LoRA, a pointer head, CISPO training, and a block-4 diffusion drafter. It scores 0.889 on unseen test data versus 0.822 for Kev-9B and 0.857 for Jev, and 0.935 on JevBench's public hard tier; latency on one H100 is about 0.3s without thinking and 3.3s median with it, with code, weights, and data open-sourced.
A reproducible benchmark for typed decision models measuring accuracy, latency and cost.
A developer launched JevBench, a reproducible benchmark for typed decision models that return bounded choices and probabilities, covering 534 English decisions and scoring accuracy, latency and cost with configurable weights. Jev leads at 74.4, followed by SemIf, djev, Winnow-12B Q8 and reflex 4B.
That is everything