Opus 4.8

3 stories

Opus 4.8 is a language model used for text generation and agent tasks. Recent coverage reports that, alongside Codex/GPT-5.6 Sol and OpenClaw, Opus 4.8 failed to reproduce open-ended ML research in a new paper, and that some users on Hacker News have chosen Opus 4.8 as a cost-effective default model while avoiding Opus 5.

Related topics

Critique: 1Password's AI Patching Benchmark

A critique of 1Password's August 6, 2026 report says its 26% "clean fix" rate is misleading: the sample focused on difficult bugs, 22% of trials instructed agents to apply wrong fixes, 36% forbade building or testing, and models used different reasoning settings. The authors also released agent skills for post-patch validation and review walkthroughs.

Hacker News · · Details

Agents can't enable recursive self‑improvement

A new paper had authors hand agents accepted-but-unpublished NeurIPS papers (evaluated by the original authors); Codex/GPT-5.6 Sol and OpenClaw/Opus 4.8 failed to reproduce open-ended ML research. The authors argue this implies current agents cannot drive recursive self‑improvement (RSI), questioning near‑term explosive AI progress.

r/MachineLearning · · Details

What default model do you use?

A Hacker News thread asks which default model people use. One user says they use Claude most, find Fable overkill and costly for their needs, and has made Opus 4.8 their go-to due to good value while avoiding Opus 5 for now.

Hacker News · · Details
That is everything