DeepSeek V4.1 Flash Narrows US-China Top AI Model Gap to 3% on LiveBench
According to a Bloomberg Intelligence report, the performance gap between top Chinese and US AI models on the LiveBench benchmark has narrowed to just 3%. This shift was largely driven by DeepSeek's release of its V4.1 Flash model, which scored 81.1 on LiveBench compared to Anthropic's leading score of 83.4. The narrowing gap highlights how Chinese AI labs are rapidly improving model capabilities through technical expertise and domestic hardware optimization despite US export restrictions. This ongoing convergence challenges the long-term sustainability of US technological dominance in artificial intelligence. DeepSeek V4.1 Flash ranked sixth globally on LiveBench, making it the highest-ranked Chinese model since the startup's R1 reasoning model breakthrough. However, despite closing the gap at the top, Chinese models still represent only 3 of the top 15 spots on the LiveBench leaderboard.
## BACKGROUND
LiveBench is an LLM benchmark designed to prevent test set contamination by evaluating AI models on complex, objective tasks that are regularly updated. Benchmarks are critical tools in the AI industry to objectively measure and compare model capabilities across reasoning, math, and coding. DeepSeek is a Chinese AI startup that gained global attention for developing high-performance models at lower training costs.