Issue 2026-09-21 · Industry · 研究 · 开源

Qwen 27B runs autonomously for 3 weeks on one 3090

A user ran a quantized Qwen 27B locally on a single RTX 3090 for about 21 days with the goal of having the agent build its own CUDA inference engine. The run consumed roughly 230M tokens, spent about 83 hours in 699 compactions, and reached about half of llama.cpp's prefill speed while producing working kernels and benchmarks.

r/LocalLLaMA33 h ago
Read original ↗