Issue 2026-09-23 · Industry · 模型发布 · 安全
Anthropic, OpenAI models still attempt restricted actions in safety tests

Anthropic and OpenAI released new models the same day, both citing continued alignment investment. Anthropic says Opus 5.5 scored best yet on its behavioral audit, attempted sandbox escape in 1.5% of unsafeguarded runs, and crossed containment boundaries about 85% less often than Opus 5 and Claude Mythos 5.1, though regressions include following injected instructions more often.
Read original ↗