DeepSeek V4.1 Flash Cut the US AI Lead to 3%. What LiveBench Measures
Bloomberg Intelligence put a number on it on Oct. 4: 2.3 points. That is the distance on LiveBench, an independent AI benchmark, between Anthropic’s Claude Fable 5.1 at 83.4 and DeepSeek’s V4.1 Flash at 81.1, which BI ranks sixth in the world. Senior analyst Robert Lea calls that a US lead of roughly 3%, the smallest US-China top-model gap BI has tracked, down from about 9% in May and about 15% earlier this year, Bloomberg reported, as summarized by AI Weekly and Implicator.
The figure travels well. Underneath it sit one specific test, one specific model setting and a bit of arithmetic. Here is how to read each of them.
Where the 3% figure comes from
The math is simple. Take the gap between the best US score and the best Chinese score, then divide by the US score: 2.3 divided by 83.4 is 2.8%, which Implicator notes was “rounded to roughly 3%.” Both numbers come from each model’s maximum reasoning-effort setting, the setting that spends the most computing per answer.

Three things follow. First, this compares the single best model on each side, not the average Chinese model against the average American one. Second, rank depends on how you count. Public mirrors of the leaderboard, such as BenchLM, list some models more than once at different effort settings, so the same 81.1 can show up lower than sixth. Third, the overall score hides category swings. On LiveBench’s agentic coding category, V4.1 Flash scored 77.3 against 66.1 for Claude Fable 5.1, according to Implicator’s reading of the Oct. 4 leaderboard data. A model that trails by 2.3 points overall leads by 11.2 on the category many developers care about most.
Lea’s read of the trend is that Chinese labs have improved their technical capabilities and tuned their models for domestic hardware, and that the progress raises questions about how well US chip export controls are working, per AI Weekly’s summary of the Bloomberg Intelligence note.
What LiveBench actually measures
LiveBench launched in June 2024 as a joint project of Abacus.AI, New York University, Nvidia, the University of Maryland and the University of Southern California. Its design answers one problem: models get trained on the internet, and the internet contains old benchmark questions. A model can ace a test it has already seen.
LiveBench’s fix, per its research paper, has two parts. Questions come from “recently-released math competitions, arXiv papers, news articles, and datasets,” and the team committed that “questions will be added and updated on a monthly basis.” Answers are “scored automatically according to objective ground-truth values” instead of being graded by another AI. That matters, because the paper found pass/fail judgments from GPT-4-Turbo had an error rate of up to 46%.
At launch the benchmark covered six categories, math, coding, reasoning, language, instruction following and data analysis, across 18 tasks. The board now also reports agentic coding, the category where DeepSeek pulled ahead.
A model that trails by 2.3 points overall leads by 11.2 on agentic coding.
What it does not measure is just as important. LiveBench says nothing about speed, cost, how a model behaves over a long working session, how it handles your company’s documents, or how it was safety-tested. It is a clean exam. It is not a job interview.
Why a ‘Flash’ model matters: price and open weights
At most labs “Flash” means the light model. Not here. DeepSeek’s Sept. 10 release notes describe V4.1 Flash as a mixture-of-experts model with a 552-billion-parameter backbone (763 billion counting every component, per its model card) that uses 8 billion active parameters to read input and 16 billion to write output. Only a small slice of the model works on each token, which is why it is cheap to serve. DeepSeek says its KV cache, the working memory a model keeps for each conversation, needs a quarter of the high-bandwidth memory and an eighth of the SSD storage of the previous generation, and that it beats the larger V4-Pro on its own tests.

The weights are public. The Hugging Face model card lists an MIT licence and a context window of 1 million tokens, and DeepSeek’s API price page caps output at 384,000 tokens. Then there is the price, set against the US model it nearly matched:
| Per 1M tokens | Input | Output |
|---|---|---|
| DeepSeek V4.1 Flash, peak | US$0.30 (about CA$0.42) | US$1.20 (about CA$1.67) |
| DeepSeek V4.1 Flash, off-peak | US$0.15 (about CA$0.21) | US$0.60 (about CA$0.83) |
| Claude Fable 5.1 | US$10 (about CA$13.90) | US$50 (about CA$69.50) |
DeepSeek’s peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, which is 9 p.m. to midnight and 2 a.m. to 6 a.m. in Toronto during daylight time. Anthropic lists Fable 5.1 at “$10 per million input tokens and $50 per million output tokens” on its launch page. At peak, DeepSeek’s output costs about one-fortieth as much. We broke down every lab’s tiers in our field guide to AI model names.
Open weights change who can use the model. On Sept. 16, Nvidia posted its own compressed version of V4.1 Flash on Hugging Face, quantized to its NVFP4 format and tuned for Blackwell GPUs such as the GB300. A company can run that build on hardware it controls without sending anything to DeepSeek’s servers. Don’t expect to run it on a laptop, though; this is data-centre software.
The caveats: benchmarks, chips and the bill
Start with the obvious one. A benchmark is a sample. LiveBench is better than most because its questions are fresh and its grading is mechanical, but a 2.3-point gap on one test is close enough that a different month’s questions could move it. Bloomberg Intelligence’s 15%, 9% and 3% series is also BI’s own tracking, and its method for the earlier points has not been published in the reports we could read.
The chips question is murkier. DeepSeek’s release notes do not say what hardware trained V4.1 Flash. What is on the record is direction: on Oct. 1, DeepSeek released tools for Huawei’s Ascend chips built with Huawei’s help, an effort aimed at Nvidia’s CUDA software lock-in. If Chinese labs can train and serve near-frontier models on domestic silicon, the export-control argument changes. That has not been shown yet.
Then the bill. China has more than 1,100 large language models competing on thin token margins, and Lea does not expect the sector to reach sustainable profitability before 2030, according to AI Weekly. A price of US$0.60 per million output tokens is great for buyers. It is a hard business.
What it means for you
- If you pick models for work: treat the 3% as a signal to test, not a reason to switch. Run your own tasks through V4.1 Flash and your current model, and compare cost per finished task, not per token.
- If data location matters: the hosted API runs on DeepSeek’s servers. The open weights, including Nvidia’s build, let you keep data on infrastructure you choose.
- If you follow the leaderboard: check which effort setting a score uses, and look at the category scores before the overall number.
One detail is easy to miss in the US-versus-China framing. Nvidia is one of the five institutions that built LiveBench, and it is also the company that published a copy of V4.1 Flash tuned for its newest chips. The model that narrowed the American lead ships, in at least one version, built for American silicon.
Catherine Crowe covers AI explained for prompt/power: the plain-English guides that break down how the technology works, what the jargon means and what it changes for everyday people. Originally from Canada, she writes from New Zealand.
Latest from prompt/power
- Ben Affleck Calls AI Job Fears ‘Propaganda.’ He Sold an AI Firm to NetflixOct 6
- Cohere’s North 2 Puts AI Agents on a Budget. Toronto Bets on BoringOct 6
- How to Stop ChatGPT, Claude, Gemini and Meta AI From Training on Your ChatsOct 6
- Musk Is a Trillionaire Again After SpaceX Stock Jumps 7.6%Oct 6
- Galaxy S26 and Pixel 10a Prices Jump in Canada: The New Price ListOct 6
Leave a Reply