Anthropic’s GLM‑5.3 can hijack control flow in 4% of binary exploitation tests. The study shows GLM‑5.3 surpasses earlier models like Claude Opus 4.6 and GLM‑5.2, which fail entirely, and even trails Claude Mythos Preview’s 6% success rate. This marks a clear threshold of advanced cyber capabilities in LLMs, raising important security concerns for developers deploying such models.
Opening Kapyn…