Qwen 3.8 27B
5 storiesQwen 3.8 27B is a locally runnable multimodal large model (often in q4xl format) with vision and MTP capabilities. Recent coverage highlights it shipping as the local model in Perplexity Portable on Windows with NVIDIA GPU acceleration, community experiments running it on Zima Board 2 + RTX 2000 ADA, and a PI agent pushing it through ~11M tokens for a 3D game design workload.
Related topics
Users on LocalLLaMA praise Muse Glimmer for natural, non-generic conversation, calling it a strong local chat model, while also experimenting with Qwen 3.8 27B for coding tasks.
UkisAI post‑trained Qwen 3.8 27B to penalize tokens tied to “overthinking,” using On‑Policy Distillation to cut thinking tokens by 58%, speed up 1.95×, and keep accuracy loss under 1%. The model is open‑sourced on Hugging Face and a free NVIDIA‑backed OpenAI‑compatible API is available (5 RPM limit).
Perplexity launched Portable Computer in its Windows app for NVIDIA GeForce/RTX PRO PCs, running local agents accelerated by NVIDIA GPUs. It ships with a local model (e.g. Qwen 3.8 27B), keeps sensitive data on-device, and can escalate tasks to cloud models with user permission.
A Reddit thread discusses using a Zima Board 2 (~$411) plus an RTX 2000 ADA (~$700) to run Qwen-3.8 27b, citing a Luke’s Dev Lab demo that achieved good token speed on Ollama powered by the Zima supply. The poster compares this setup to a Mac Mini M5 24GB and asks if cheaper new self-contained alternatives offer similar token performance.
A user pushed Qwen 3.8 27B (q4xl) locally with a PI agent to handle a large game design file, using ~120k context plus vision and MTP. The run read ~11M tokens and wrote ~3.2M tokens over about 12 hours, demonstrating local model capability for complex game design tasks.
That is everything