Issue 2026-09-20 · Industry · 开源 · 研究
Qwen3.8 27B IQ3_XXS vs Bonsai Ternary PQ2 quantization test
A user benchmarked Qwen3.8 27B IQ3_XXS (10.18GiB) against Bonsai Ternary PQ2 (6.42GiB) on the same 16GB VRAM setup. Both completed 4/4 tasks, but Qwen ran 3.02x faster and used 3.10x fewer output tokens; Bonsai was slower with slightly worse results.
Read original ↗