Issue 2026-09-30 · Industry · 模型发布 · 开源
Qwen3.8 runs locally at 51 t/s on 12GB laptop GPU
A user reports that the Strata inference engine with ISTA-DASLab's quantized Qwen3.8-Flash-Next hits 51 t/s generation and 1500 t/s prompt processing on a 12GB VRAM, 64GB RAM laptop, far outpacing stock llama.cpp. Strata currently supports only Nvidia and select GGUF models.
Read original ↗