AI models struggle with refusal mechanisms, often bypassing safety filters for harmful requests. A study found models like GPT-4 and Claude can be tricked into complying with dangerous prompts via indirect phrasing. This exposes critical gaps in alignment and risks for developers deploying AI systems.
Opening Kapyn…