~/LLM/bilibili-launches-ai-infinite-arena-leaderboard-for-creator-driven-llm-evaluation

Bilibili Launches 'AI Infinite Arena' Leaderboard for Creator-Driven LLM Evaluation

Chinese video platform Bilibili launched the "AI Infinite Arena," a real-time updated leaderboard where platform content creators evaluate over 100 large language models across real-world workflows, coding, reasoning, and creative tasks. In the initial rankings, GPT-6 Astra took the top spot after winning 10 creator evaluations, while domestic Chinese models secured three of the top five positions. While traditional LLM benchmarks rely on static test suites or blind pairwise voting, Bilibili's platform leverages video creators testing models in practical niche workflows. This reflects a growing trend of platform-driven crowd-testing designed to engage tech communities and help everyday users select AI tools for specific practical applications. Unlike standardized academic benchmarks like MMLU or LMSYS Chatbot Arena's blind pairwise voting, the platform places no restrictions on testing dimensions or topics, allowing creators to design their own evaluation prompts based on their professional fields. Models tested include major international and domestic systems such as OpenAI's ChatGPT, Anthropic's Claude, Google's Gemini, DeepSeek, Kimi, Doubao, Qwen, and Zhipu AI's GLM series.

## BACKGROUND

LLM evaluation typically relies on static academic benchmarks (such as MMLU) or crowdsourced arena-style platforms like LMSYS Chatbot Arena, which uses pairwise Elo ratings based on user preference votes. Zhipu AI's GLM family is among China's prominent open and foundational LLM series, competing directly with major Western models in various reasoning and domain-specific tasks.

## REFERENCES

## KEYWORDS

#LLM#AI Benchmarks#Bilibili#Model Evaluation

$ subscribe --daily

Bilibili Launches 'AI Infinite Arena' Leaderboard for Creator-Driven LLM Evaluation | Daily News