Developer fine-tunes Bev, a model wrong 98% of the time
A developer fine-tuned Qwen3.5-9B into Bev, a decision model that answered correctly only 1.9% of 324 held-out decisions while averaging 96% confidence. Built as a control case, it tests whether automated pipelines actually verify model correctness rather than just confidence.