Open-source Kev 4B tested side by side with Jev
A side-by-side test of Kev 4B (an Apache-2.0 Qwen3.5-4B fine-tune) and Jev on 362 fresh items found accuracy within 2 points on every task. Jev led on PAWS paraphrase detection (87.0% vs 74.5%) but counts ~257 extra input tokens per request, making short requests up to 12x costlier.