Issue 2026-09-14 · Industry · 模型 · 研究
Hope for KVCache + Engram in Upcoming Models
A Reddit post urges applying DeepSeek-V4.1-Flash’s KVCache and Engram ideas to medium/large models so 30B-class models can run on GPUs in Q8/Q4. The post gives approximate VRAM and Engram estimates, suggests Engram may be ~1/3–1/2 of model size (10–15 GB for 30B), and argues 32GB VRAM should suffice for many optimized models.
Read original ↗