JevBench

3 stories

JevBench is a reproducible benchmark for typed decision models that return bounded choices and probabilities, scoring accuracy, latency and cost across 534 English decisions with configurable weights. Recent coverage includes Jeeves, a Jev-like model built on Qwen3.5-9B with LoRA, a pointer head, CISPO training and a block-4 diffusion drafter, which scored 0.889 on unseen test data and 0.935 on JevBench's public hard tier; earlier Jev led the board at 74.4, ahead of SemIf, djev, Winnow-12B Q8 and reflex 4B.

Related topics

Jeeves: Reasoning Improves Jev-Like Decision Models

A developer released Jeeves, a Jev-like decision model built on Qwen3.5-9B with LoRA, a pointer head, CISPO training, and a block-4 diffusion drafter. It scores 0.889 on unseen test data versus 0.822 for Kev-9B and 0.857 for Jev, and 0.935 on JevBench's public hard tier; latency on one H100 is about 0.3s without thinking and 3.3s median with it, with code, weights, and data open-sourced.

Hacker News ·

JevBench

A reproducible benchmark for typed decision models measuring accuracy, latency and cost.

▲ 26 · Show HN · 1 comments ·
That is everything