User builds 256GB VRAM rig with 8 Radeon Pro V620s
A user built a roughly $2,800 inference rig with eight used Radeon Pro V620 cards totaling 256GB VRAM, running Qwen3.8-Flash-Next via a custom vLLM fork. It reached 60-100 tokens/s decode and over 3,000 tokens/s prefill, roughly 8x faster than llama.cpp.