Issue 2026-09-16 · Industry · 研究 · 模型

Offload most of Qwen3.8 KV cache to RAM

A vLLM user shared on r/LocalLLaMA a method to keep most of Qwen3.8-Flash-Next's KV cache in system RAM while using a barely-fitting VRAM quant, achieving 1M context on 3×3090 GPUs with ~80 tok/s at short context, ~60 tok/s at long context, and 3,701 tok/s at a 248k prefill; patches and the model are available on their Hugging Face page.

r/LocalLLaMA15 h ago
Read original ↗