Developer Runs Qwen3.5 LLMs on Cheap FPGA Mining Hardware
A developer used ~$280 SQRL FK33 ex-mining FPGAs, with Claude assisting the RTL work, to run Qwen3.5-9B INT4 inference. Two cards at 75MHz produced ~3.2 tok/s generation, dropping to ~2.4 tok/s at longer context, with layer-by-layer validation against llama.cpp.