Issue 2026-09-09 · Industry · 研究 · 开源
GLM-5.3-Flash optimized on M3 Ultra for 60 tps / 550 tps
An engineer optimized GLM-5.3-Flash on Apple’s M3 Ultra by fusing kernels and using parallel scans, boosting throughput from 21.6 to 37.4 t/s at 300k context and achieving up to ~60 t/s for SQL workloads at short context.
Read original ↗