~/AI BENCHMARK/humor-arena-benchmarking-llms-on-joke-generation-using-an-automated-judge

Humor Arena: Benchmarking LLMs on Joke Generation Using an Automated Judge

Researchers introduced Humor Arena, a benchmark comparing 20 LLM model versions across 360 joke prompts using a fine-tuned open-source automated judge. Fable 5 achieved the highest estimated score, winning 66.8 points per 100 pairwise comparisons against other models. Evaluating subjective open-ended tasks like humor has traditionally been difficult for automated benchmarks focused on objective facts or code. Demonstrating an automated judge that correlates strongly with human humor preferences provides a scalable framework for assessing qualitative AI capabilities. The benchmark tested models on 360 frozen prompts requesting four jokes per prompt, with model identities hidden during scoring. The automated judge was audited against 1,400 human ratings from 50 people to verify that its evaluations closely align with human subjective judgment.

## BACKGROUND

LLM-as-a-Judge is an evaluation methodology where a specialized language model rates the outputs of other AI models based on specific rubrics. While conventional evaluation metrics struggle with open-ended creative writing, fine-tuning an automated judge on human preference data enables scalable, reliable scoring without requiring continuous manual human review.

## REFERENCES

## KEYWORDS

#AI Benchmarks#LLM Evaluation#Natural Language Processing#Humor Evaluation

$ subscribe --daily

Humor Arena: Benchmarking LLMs on Joke Generation Using an Automated Judge | Daily News