Issue 2026-10-04 · Industry · 开源 · 芯片 · 研究

Developer Runs Qwen3.5 LLMs on Cheap FPGA Mining Hardware

A developer used ~$280 SQRL FK33 ex-mining FPGAs, with Claude assisting the RTL work, to run Qwen3.5-9B INT4 inference. Two cards at 75MHz produced ~3.2 tok/s generation, dropping to ~2.4 tok/s at longer context, with layer-by-layer validation against llama.cpp.

r/LocalLLaMA2 d ago
Read original ↗