Four RTX 3060 Ti GPUs Hit 120 t/s Local Inference With 262K Context
A user built a 4-GPU rig from power-limited 8GB RTX 3060 Ti cards and, using tensor parallelism with Exl3 and HyperQwen, reached roughly 120 t/s at 150K context in bf16; quantizing the KV cache extends context to 262K with little slowdown under concurrency.