This study investigates the trade-off between concept control and text fluency in large language models. The authors systematically evaluate various conditioning methods for concept injection and removal across multiple benchmarks. The findings reveal that efficient steering techniques often severely degrade generation quality, highlighting a persistent challenge for reliable model deployment.
Opening Kapyn…