Arena Launches Factuality Leaderboard to Evaluate AI Model Accuracy and Hallucinations
Arena has announced a new factuality leaderboard that ranks AI models based on the factual accuracy of their responses alongside human preference. This new system aims to objectively evaluate model performance and guard against hallucinations. Traditional LLM evaluation leaderboards heavily rely on human preference, which can be biased by formatting, style, or length rather than factual correctness. Introducing a factuality-weighted metric helps users identify which models are truly reliable and accurate, rather than just articulate. The factuality score is calculated using a weighted combination of human preference and the correctness of claims in a model's response. Users can access this new ranking on the Text and Search Arena leaderboards via an opt-in toggle in the user interface.
## BACKGROUND
LMSYS Chatbot Arena is a widely used benchmarking platform that ranks large language models (LLMs) based on crowdsourced human preferences using Elo ratings. However, human evaluators often suffer from style bias, favoring longer or better-formatted responses even if they contain factual errors or hallucinations. To address this, Arena previously introduced style control methodologies and is now integrating direct factuality checks.