Issue 2026-09-24 · Industry · 开源 · 模型发布 · 应用
Qwen3.8-Flash-Next Hits 65 Tokens/Second on 12GB VRAM
A developer built a custom inference engine, Strata, running Qwen3.8-Flash-Next on consumer hardware including a 12GB RTX 5070, achieving about 65 tokens/s output and 430 tokens/s prompt processing, a major gain over the earlier llama.cpp setup. The engine is open-sourced, one-click installable, and currently optimized for CUDA only.
Read original ↗