Open-source engine Splash hits 144 tok/s on M5 Max
Open-source inference engine Inco Splash launched, claiming 144 tok/s for Qwen3.8-27B on an M5 Max MacBook Pro, up to 3x Ollama and 2x oMLX decode speed. It requires M3 or newer, macOS 26.4+, and 36GB RAM, works with Claude Code and Codex, and is available in LM Studio.