adtestbench

AI news · DeepSeek · models

DeepSeek cuts the US lead on LiveBench to about 3%

A Bloomberg Intelligence comparison puts China’s best model 2.3 points behind Anthropic’s top entry, and ahead on agentic coding.

By Alexandre S. , 05:30 UTC

DeepSeek’s September release has cut the gap between the best US and Chinese AI models to about 3% on the LiveBench leaderboard, in a comparison by Bloomberg Intelligence senior analyst Robert Lea, Implicator reports on Sunday 4 October 2026. In May his figure was nearer 9%, and 15% at the start of the year.

Retro-futurist illustration: a striped sunset over a grid horizon under a starry sky, with a microchip standing on the horizon.
Drawn by adtestbench from “DeepSeek Narrows US AI Benchmark Lead to 3% After September Release”,

Both at maximum effort, Anthropic’s Claude Fable 5.1 sits at 83.4 overall and DeepSeek V4.1 Flash at 81.1 in the 4 October snapshot. On agentic coding the order flips, with DeepSeek at 77.3 and Claude Fable 5.1 at 66.1. Lea sees Chinese developers taking more share on the back of it, and says China’s AI industry could remain unprofitable until 2030.

Crypto Briefing, summarising Bloomberg’s report, adds that a May 2026 evaluation by the US Center for AI Standards and Innovation (CAISI) found the best Chinese models six to eight months behind on complex reasoning and cyber tasks, and that DeepSeek V4-Pro costs about $3.96 per million output tokens.

For teams choosing a model, the single average hides a split: DeepSeek leads on agentic coding, while CAISI’s May evaluation had US models ahead on reasoning and cyber tasks.