kapynResearch

Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

The paper proposes selective refusal to improve AI safety without losing useful content. It argues that refusing only the problematic subtopic preserves valuable information while mitigating risk. The authors evaluate the approach on standard benchmarks, showing reduced harmful outputs and higher user satisfaction. The work offers a practical framework for developers to implement nuanced refusal policies.

Hugging Face·Sep 8, 2026

Opening Kapyn…