Issue 2026-10-04 · Industry · 开源 · 产品
Developer releases Ninfer 4080 for 16GB GPUs with 100k context
A developer released Ninfer 4080, an inference solution that runs ISTA-DASLab-Qwen-3.8-27B-GSQ at 100k context on a 16GB RTX 4080, peaking at 2720 tok/s prefill and 262 tok/s generation. The project is open-sourced to bring better performance to 16GB-class GPUs than general-purpose engines.
Read original ↗