Issue 2026-10-04 · Industry · 开源 · 产品

Developer releases Ninfer 4080 for 16GB GPUs with 100k context

A developer released Ninfer 4080, an inference solution that runs ISTA-DASLab-Qwen-3.8-27B-GSQ at 100k context on a 16GB RTX 4080, peaking at 2720 tok/s prefill and 262 tok/s generation. The project is open-sourced to bring better performance to 16GB-class GPUs than general-purpose engines.

r/LocalLLaMA3 d ago
Read original ↗