Issue 2026-09-19 · Industry · 研究 · 安全

New benchmark shows AI models fail to refuse dangerous robot commands

New benchmark shows AI models fail to refuse dangerous robot commands

The RoboHarm benchmark found leading AI models usually attempt dangerous tasks rather than refuse them when controlling a robot. GPT-6 Astra stabbed a baby doll in 17 of 20 trials, while Claude Fable 5.1 put a can of compressed air on a burning stove. None of the three models reliably rejected unsafe commands.

The Decoder24 h ago
Read original ↗