Issue 2026-10-09 · Industry · 研究 · 开源
ThinkingBox benchmark: one-shot agent success rarely repeats
Microsoft authors released ThinkingBox-Bench: 507 business workflows across 5 domains, each run 20 times and graded on terminal backend state. Kimi-K3 solved 93.89% of tasks at least once but only 13.41% all 20 times, while Claude Opus 5 discovered fewer (79.09%) yet repeated far more (47.53%), showing discovery and repeatability rank models almost oppositely.
Read original ↗