~/AI BENCHMARK/do-ai-labs-over-optimize-models-for-the-pelican-riding-a-bicycle

Do AI Labs Over-Optimize Models for the "Pelican Riding a Bicycle" Benchmark?

Dylan Castillo conducted a structured evaluation of seven major AI models across 48 animal-vehicle prompt combinations to see if labs are over-optimizing for the popular "pelican riding a bicycle" benchmark. The study found no significant evidence of "pelicanmaxxing," as models did not perform better on this specific combination compared to other animals or vehicles. This analysis addresses concerns about benchmark contamination and gaming, where AI developers might train models on specific viral test prompts to artificially boost perceived performance. It reassures the community that informal benchmarks can still serve as relatively unbiased indicators of a model's general capabilities. The evaluation tested models like GPT-5.6 Terra, Claude Sonnet 5, and Gemini 3.5 Flash using 8 animals and 6 vehicles, with GPT-5.6 Luna and Gemini 3.1 Flash-Lite assisting in grading the outputs. While GLM-5.2 showed a slight performance boost on the exact pelican-bicycle prompt, the effect was too small to be statistically significant.

## BACKGROUND

The prompt "Generate an SVG of a pelican riding a bicycle" was popularized by developer Simon Willison as an informal benchmark to test the spatial reasoning and code-generation capabilities of AI models. In machine learning, "benchmark contamination" or "gaming" occurs when models are trained directly on test datasets, leading to high test scores that do not reflect real-world utility.

## REFERENCES

## KEYWORDS

#AI Benchmarking#Large Language Models#Model Evaluation#Machine Learning

$ subscribe --daily