Qwen Max Overtakes Claude Opus on Artificial Analysis Agentic Index
The Qwen Max model (referred to as Qwen 3.8 Max) has achieved the top ranking on the Artificial Analysis agentic index. It has surpassed Claude Opus (referred to as Opus 5) to become the highest-ranked model for agentic workflows. This milestone highlights a significant shift in LLM leaderboards, showcasing the rising competitiveness of Qwen models in complex, multi-step tasks. It provides developers with a new leading option for building autonomous AI agents that require tool use and planning. The Artificial Analysis Agentic Index evaluates models based on their ability to execute long-horizon knowledge work, use tools, and make autonomous decisions. The benchmark includes evaluations like AA-Briefcase, which tests models on realistic business workflows requiring deliverables like spreadsheets and presentations.
## BACKGROUND
Agentic AI refers to AI systems that can pursue goals, use tools, and take actions autonomously within human-defined constraints. Artificial Analysis is an independent platform that benchmarks AI models, and its Agentic Index focuses specifically on evaluating these complex, multi-step capabilities rather than simple text generation.