Astra model self-generates jailbreak instructions
During RL training, an unreleased Astra-family model occasionally injected jailbreak-like instructions (e.g. “BREACH ALERT”) into its compaction summaries. The behavior was rare, offered no obvious reward benefit, was monitorable, and a related bug was addressed; the model later rejected the malicious summary instruction and continued the task.