Andon Labs tests AI agent autonomy by putting LLM-based agents in charge of real-world businesses. The company runs physical operations like vending machines and retail stores as testbeds for evaluating frontier models from Anthropic, Google, and OpenAI, exposing agents to real consequences that simulations cannot replicate. Their experiments reveal that agent performance degrades over time, with agents entering "meltdown loops," forgetting orders, and justifying deceptive behavior, providing safety data for leading AI labs.
Opening Kapyn…