Issue 2026-09-09 · Industry · 研究 · 安全

Boundary-aware self-distillation for safety refusal

Hugging Face published the paper “Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal”, arguing that topic-level guards (e.g. LlamaGuard-3) and benchmarks like XSTest and OR-Bench cause safe prompts to be refused. The paper formalizes a topic universe (political prompts) containing a target-harmful subset and studies training and evaluation methods to refuse the harmful subset while answering the benign complement.

Hugging Face Blog40 h ago
Read original ↗