Logit penalty on Qwen models improves math accuracy
A user tested logit-bias penalties on words like "wait", "maybe" and "perhaps" across quantizations of Qwen3.5-4B in llama.cpp, running 50 random MATH-500 questions, building on a Meta paper on overthinking markers. The paper had not covered llama.cpp's quantization options.