Issue 2026-10-10 · Industry · 开源 · 研究

Qwen 3.8 Flash runs at ~21 tok/s on RTX 3060

A developer reports running a 68GB quantized Qwen 3.8 Flash Next MoE at about 21 tok/s on an RTX 3060 12GB with 16GB RAM, using expert-routing prediction to speed CPU/GPU offload with no gate pruning and bit-exact output. Warm cache reaches 24+ tok/s, versus roughly 1.4-2.1 tok/s with stock llama.cpp.

r/LocalLLaMA40 h ago
Read original ↗