Running Qwen3.8-Flash-Next on 12GB VRAM
On an RTX 4070 (12GB VRAM) with 64GB RAM, the author used quantization, n-gram SSD offloading, --fit settings, and MTP patches to raise Qwen3.8-Flash-Next generation from ~6 tok/s to nearly 20 tok/s, with PP around 300–350 tok/s, and documents the trade-offs and patches used.