~/MULTIMODAL A/lmsys-arena-expands-preference-based-reward-modeling-to-multimodal-domains

LMSYS Arena Expands Preference-Based Reward Modeling to Multimodal Domains

LMSYS Arena has expanded its preference-based reward modeling beyond text to include vision, image generation, and code. They are training modality-specific reward models using millions of live Arena preference pairs, including over 3 million pairs for text-to-image generation. This expansion provides high-quality, large-scale human preference datasets for multimodal AI, which is crucial for Reinforcement Learning from Human Feedback (RLHF). It helps align multimodal models more closely with human preferences across diverse tasks beyond text. The text-to-image reward model was trained on over 3 million preference pairs and evaluated on the public Multimodal Rewardbench 2 (MMRB2) benchmark. MMRB2 is a unified benchmark developed by Meta Research for evaluating reward models on multimodal understanding and generation.

## BACKGROUND

Preference-based reward modeling is a machine learning technique that learns reward functions from human pairwise comparisons, often using Bradley-Terry models. These reward models act as "judges" to guide AI behavior during RLHF. MMRB2 (Multimodal Rewardbench 2) is a comprehensive benchmark designed to evaluate these reward models across interleaved text and multimodal inputs.

## REFERENCES

## KEYWORDS

#Multimodal AI#RLHF#Reward Modeling#AI Evaluation

$ subscribe --daily

LMSYS Arena Expands Preference-Based Reward Modeling to Multimodal Domains | Daily News