Issue 2026-10-02 · Industry · 研究 · 产品
OpenAI Decisions API Confidence Unreliable

Tests show GPT-6 Luna, behind OpenAI's new Decisions API, is poorly calibrated: when it claimed 99% confidence across 3,600 reasoning problems, it was right only 68% of the time, highlighting AI confidence-calibration problems.
Read original ↗- OpenAI's Decisions API: a Jev clone for fast agent decisions
- jevals: replacing LLM judges with typed Jev decisions for agent evals
- Developer Says His Non-Autoregressive Decision Model Was Rebranded a Breakthrough
- AI decision models play Pac-Man: jev 1.13 tops leaderboard
- Musubi launches open-weight decision model PolicyLM-1.7B