Community Urges Halt to FP4 Inference on Small Models
An r/LocalLLaMA post criticizes new inference engines that tout NVFP4/MXFP4 support, arguing FP4 quantization badly damages small dense models, causing hallucinations and reasoning errors, while large models can absorb minor precision loss.