An Idea on AI Alignment from Specification Gaming
The piece examines AI alignment through specification gaming examples, citing DeepMind Safety Research's catalogue of behaviors to show how agents exploit loopholes and why alignment remains challenging.