Tencent's Hy3 Ranks #5 Among Open-Weight Models in Agent Arena
Tencent's Hy3 open-weight model has achieved the #5 spot for open-weight models (#25 overall) in the Agent Arena benchmark. Additionally, it ranked as the #2 open-weight model (#16 overall) in the Frontend Code Arena. This achievement highlights the growing capabilities of open-weight models in handling complex, real-world tasks like tool-use and frontend coding. It demonstrates that open-weight alternatives are becoming increasingly competitive with proprietary models in agentic workflows. Hy3 demonstrated specific strengths in tool-use, particularly in recovering well from CLI and bash errors, which boosted its performance. The underlying model is Tencent Hunyuan 3, a 295-billion-parameter Mixture-of-Experts (MoE) model with 21 billion active parameters.
## BACKGROUND
Agent Arena, developed by Arena.ai, is a benchmark designed to evaluate AI agents on live, interactive tasks rather than static test questions. Tencent Hunyuan 3 (Hy3) is Tencent's latest flagship LLM, which has been integrated into various internal products like CodeBuddy and made available to developers via APIs.