Issue 2026-10-07 · Industry · 开源
llama.cpp Adds GPU Cache for MoE Experts in Host Memory
Developer am17an submitted PR #29887 to llama.cpp adding a GPU cache for MoE experts kept in host memory, potentially delivering big speedups for Mixture-of-Experts models that don't fully fit in VRAM.
Read original ↗