Issue 2026-10-09 · Industry · 研究 · 安全

Reexamining AI Refusal: The Cost of Disobedience

Reexamining AI Refusal: The Cost of Disobedience

MIT Technology Review examines how AI refusal mechanisms evolved: early models readily generated harmful content, but reinforcement learning now makes them refuse vastly more prompts, at the cost of over-refusal. Anthropic's 2021 'harmless' principle became an industry norm.

MIT Technology Review2 d ago
Read original ↗