Issue 2026-09-19 · Industry · 开源 · 产品

Open-source engine Splash hits 144 tok/s on M5 Max

Open-source inference engine Inco Splash launched, claiming 144 tok/s for Qwen3.8-27B on an M5 Max MacBook Pro, up to 3x Ollama and 2x oMLX decode speed. It requires M3 or newer, macOS 26.4+, and 36GB RAM, works with Claude Code and Codex, and is available in LM Studio.

r/LocalLLaMA36 h ago
Read original ↗