Simon Willison Tests Frontier LLMs with an Absurd SVG Generation Prompt
Simon Willison ran an experiment testing frontier AI models—including Mistral Large 4, Claude Opus 5.5, GPT-6.1 Sol, and Gemini 3.8 Flash—by asking them to generate SVG code for an 'armadillo in fishnet tights jaywalking on Mars.' The test was directly inspired by a humorous Hacker News comment regarding AI benchmark saturation. As standard AI benchmarks suffer from ceiling effects, highly specific and novel visual prompts provide an informal way to evaluate a model's true reasoning and spatial synthesis abilities without dataset contamination. It highlights how different frontier LLMs handle complex XML code generation for highly surreal requests. Willison executed the experiment using his open-source command-line tool `llm` across the models at their default reasoning levels. The generated SVG outputs were rendered and compared using a custom online Markdown SVG renderer.
## BACKGROUND
AI benchmark saturation occurs when standardized evaluation metrics lose their effectiveness because frontier models reach near-perfect scores or absorb test sets into their training data. SVG (Scalable Vector Graphics) generation is widely used as an ad-hoc test for language models, requiring the LLM to output structured XML code that renders coherent visual scenes without real-time visual feedback.