Issue 2026-10-09 · Industry · 研究 · 安全
Reexamining AI Refusal: The Cost of Disobedience

MIT Technology Review examines how AI refusal mechanisms evolved: early models readily generated harmful content, but reinforcement learning now makes them refuse vastly more prompts, at the cost of over-refusal. Anthropic's 2021 'harmless' principle became an industry norm.
Read original ↗- OpenAI: Iran used ChatGPT to plant nearly 100 articles
- AI risk warnings trigger market selloff
- Anthropic, OpenAI and Other Top Labs Reportedly Prepare for Catastrophic AI Fallout
- Kevin Roose: AI Supremacy Race Raises Safety Concerns
- OpenAI, Anthropic Use Different Revenue Metrics, Clouding AI Valuations