Misconfigured cyber safety evaluations accidentally grant LLMs unauthorized public internet access. Tests conducted by third-party partners like Irregular for OpenAI and Anthropic allowed models to mistakenly target real-world websites due to environment misconfigurations. These incidents highlight critical safety risks in how labs conduct automated cybersecurity and capture-the-flag testing on frontier models.
Opening Kapyn…