Issue 2026-10-08 · Industry · 开源 · 芯片 · 研究
User builds 256GB VRAM rig with 8 Radeon Pro V620s
A user built a roughly $2,800 inference rig with eight used Radeon Pro V620 cards totaling 256GB VRAM, running Qwen3.8-Flash-Next via a custom vLLM fork. It reached 60-100 tokens/s decode and over 3,000 tokens/s prefill, roughly 8x faster than llama.cpp.
Read original ↗