An empirical study investigates whether AI labs deliberately optimize multimodal models for specific benchmark prompts. Researcher Dylan Castillo tests 48 animal-and-vehicle combinations across seven major frontier models to check for benchmark overfitting known as "pelicanmaxxing." The analysis finds no evidence that labs specifically train models to excel at generating niche test cases like pelicans riding bicycles.
Opening Kapyn…