Qwen3.8-Flash-Next Hits 1M Context on Strix Halo
Developer releases halogen 0.12.0, fixing degradation at long context depth. On a Ryzen AI Max+ 395 with 128GB RAM, decode at 1M tokens of context improved from 27.3 to 38.3 tok/s, with prefill taking about 18 minutes cold.