User Shares Strata Speedup for Local LLMs but 64GB RAM Constraint
A Reddit user reported that on a dual-GPU setup (R9700, RTX 5060 Ti) with 64GB DDR5 RAM, running a quantized Qwen3.8-Flash-Next via Strata boosted inference from about 21 to 60 tokens per second, but memory usage hit 96%, preventing simultaneous ComfyUI image/video workloads, highlighting high system RAM demands of local LLMs.