Qwen3.8-Flash-Next supports 1M context on MLX-serve
Qwen3.8-Flash-Next was shown running up to ~760k context and claimed 1M-context support on MLX-serve using M5 Max with 8-bit KV cache, sustaining ~40–75 tok/s generation and requiring ~117GB peak memory.