Issue 2026-10-02 · Industry · 芯片 · 开源
Two 96GB Ascend Cards Run Qwen3.8 Flash-Next
A user built a local inference machine with two Huawei Atlas 300I Duo cards (96GB each, actually two 48GB chips per card). After vLLM, driver and operator optimizations, Qwen3.8 Flash-Next went from about 1 token/s to roughly 30 tok/s single-request and 61 tok/s at four-way concurrency, completing all 198 GPQA Diamond questions.
Read original ↗