Reexamining AI Refusal: The Cost of Disobedience
MIT Technology Review examines how AI refusal mechanisms evolved: early models readily generated harmful content, but reinforcement learning now makes them refuse vastly more prompts, at the cost of over-refusal. Anthropic's 2021 'harmless' principle became an industry norm.