Issue 2026-09-24 · Industry · 模型发布 · 开源 · 研究

Hugging Face releases VLM speculative decoding drafter

Hugging Face releases VLM speculative decoding drafter

Hugging Face released an experimental DSpark draft model for its LFM2.5-VL-3B vision-language model, adding 280M parameters (8.9%) for up to 3.13x faster decoding on device and 2.66x on an H100 without changing output quality, with day-one support for llama.cpp, MLX-VLM and SGLang.

Hugging Face Blog2 d ago
Read original ↗