Moonshot AI's Kimi K3 Ranks Second on AA-Briefcase AI Agent Benchmark
Moonshot AI's Kimi K3 model scored 1543 on the AA-Briefcase AI agent benchmark, surpassing GPT-5.6 Sol (1501) and ranking second globally behind Claude Fable 5 (1574). This marks a significant improvement from its predecessor, Kimi K2.6, which scored 816 on the same evaluation. The benchmark results demonstrate Kimi K3's strong capability in handling complex, long-horizon business workflows, positioning Chinese AI models at the forefront of the global agentic AI landscape. Additionally, Microsoft's reported interest in testing Kimi K3 for Copilot highlights its potential to offer cost-effective, high-performance reasoning. Kimi K3 features 2.8 trillion parameters and a 1 million token context window, with pricing set at 2 RMB per million input tokens for cache hits (20 RMB for misses) and 100 RMB for output. The AA-Briefcase benchmark evaluates models on realistic, messy corporate tasks, including processing thousands of emails and Slack messages to generate Excel models and PPT presentations.
## BACKGROUND
AA-Briefcase is a frontier AI agent benchmark launched by Artificial Analysis in June 2026. Unlike traditional benchmarks that test short-form Q&A, it simulates multi-week, complex business projects to evaluate how well AI agents can handle messy, long-horizon knowledge work. Moonshot AI is a prominent Chinese AI startup known for its Kimi series of large language models, which emphasize long-context capabilities.