Issue 2026-10-04 · Industry · 应用 · 开源

User Shares Strata Speedup for Local LLMs but 64GB RAM Constraint

A Reddit user reported that on a dual-GPU setup (R9700, RTX 5060 Ti) with 64GB DDR5 RAM, running a quantized Qwen3.8-Flash-Next via Strata boosted inference from about 21 to 60 tokens per second, but memory usage hit 96%, preventing simultaneous ComfyUI image/video workloads, highlighting high system RAM demands of local LLMs.

r/LocalLLaMA3 d ago
Read original ↗