Issue 2026-09-20 · Industry · 开源 · 研究
Community Urges Halt to FP4 Inference on Small Models
An r/LocalLLaMA post criticizes new inference engines that tout NVFP4/MXFP4 support, arguing FP4 quantization badly damages small dense models, causing hallucinations and reasoning errors, while large models can absorb minor precision loss.
Read original ↗