Issue 2026-10-07 · Industry · 开源 · 研究

Developer fine-tunes Bev, a model wrong 98% of the time

A developer fine-tuned Qwen3.5-9B into Bev, a decision model that answered correctly only 1.9% of 324 held-out decisions while averaging 96% confidence. Built as a control case, it tests whether automated pipelines actually verify model correctness rather than just confidence.

r/LocalLLaMA5 d ago
Read original ↗