Issue 2026-09-25 · Industry · 模型发布 · 开源 · 研究
Benchmark: ThinkingCap vs Swift vs Qwen 3.8-27B
A third-party Aider evaluation found that ThinkingCap and Swift fine-tunes cut Qwen3.8-27B's median reasoning tokens from 12,547 to about 7,400, a roughly 40% reduction matching their model cards, with first-try pass rates of 27.1%-30.8% and retry pass rates of 75.7%-77.6%, showing no meaningful performance loss.
Read original ↗