Benchmarks
Where deepseek-v4-pro and deepseek-v4-flash
sit against the field, what the 0813 checkpoint changed, and the economics
argument the Chinese community calls the 斩杀线, the
kill line. The numbers below are DeepSeek's own launch-day figures unless
marked otherwise; read the caveats first.
- Vendor numbers, one harness. The table is DeepSeek's own agent-benchmark chart, published 2026-08-12 with the GA. A score is a model-and-harness result, not a model-only one, and no independent same-harness run of 0813 exists yet. Treat these as the claim, not the verdict.
- No GPT-5.6 column. DeepSeek's chart compares against Kimi K3, GLM-5.2, Claude Opus 4.8 and Fable 5 only. GPT figures elsewhere on this page are drawn from those vendors' own releases and are cross-vendor, so they are looser still.
- Kimi and GLM are single-sourced. Those two columns appear only on the extended variant of the chart; the shared columns are identical across every copy, so the numbers are consistent, but the Kimi and GLM rows rest on one source.
- Independent history says be careful. The one held-out check on the V4-Pro preview, from NIST/CAISI, put it closer to GPT-5 and roughly eight months behind the frontier, below its self-reported position. No equivalent 0813 evaluation exists yet.
The launch table
Higher is better. HLE is shown as without-tools / with-tools. Every figure is from DeepSeek's GA chart of 2026-08-12; a dash means the vendor did not report it.
| Benchmark | V4-Pro 0813 | V4-Flash 0731 | Kimi K3 | GLM-5.2 | Opus 4.8 | Fable 5 |
|---|---|---|---|---|---|---|
| Terminal-Bench 2.1 | 87.9 | 82.7 | 88.3 | 81.6 | 85.0 | 88.0 |
| DeepSWE | 62.7 | 54.4 | 67.5 | 46.2 | 58.0 | 70.0 |
| Toolathlon-Verified | 74.1 | 70.3 | 76.5 | 59.9 | 76.2 | 77.9 |
| CyberGym | 83.3 | 76.7 | 80.0 | – | 78.3 | 83.1 |
| NL2Repo | 61.5 | 54.2 | – | 48.9 | 69.7 | – |
| AutomationBench | 31.8 | 25.1 | 30.8 | 12.9 | 27.2 | 29.1 |
| DSBench-FullStack | 71.1 | 68.7 | 63.0 | 51.8 | 71.6 | 77.2 |
| DSBench-Hard | 67.2 | 59.6 | 73.7 | 54.5 | 71.7 | 68.3 |
| Agents' Last Exam | 25.7 | 25.2 | 24.5 | 23.9 | 25.7 | – |
| HLE (no / tools) | 42.7/60.0 | 37.8/51.5 | 43.5/56.0 | 40.5/54.7 | 49.8/57.9 | 53.3/63.0 |
The honest read of the row-by-row: V4-Pro is within a point of Fable 5 on Terminal-Bench and CyberGym, ahead of Opus 4.8 on Terminal-Bench, DeepSWE, CyberGym and AutomationBench, and behind Kimi K3 on Terminal-Bench, DeepSWE, Toolathlon and DSBench-Hard. It is in the cluster, not clear of it. On knowledge without tools (HLE) the closed models still lead. Flash trails Pro across the board but stays remarkably close for a fifth of the price, which is the whole point of the next two sections.
What GA changed
The model ID stayed deepseek-v4-pro and the
rate card did not move. What moved is the
checkpoint. Against the April preview, DeepSeek's own chart shows the gains
landing almost entirely in the agentic and SWE suites:
| Benchmark | Preview | GA 0813 | Delta |
|---|---|---|---|
| DeepSWE | 12.8 | 62.7 | +49.9 |
| DSBench-Hard | 31.1 | 67.2 | +36.1 |
| CyberGym | 52.7 | 83.3 | +30.6 |
| DSBench-FullStack | 41.8 | 71.1 | +29.3 |
| NL2Repo | 38.5 | 61.5 | +23.0 |
| Toolathlon-Verified | 55.9 | 74.1 | +18.2 |
| Terminal-Bench 2.1 | 72.1 | 87.9 | +15.8 |
A near five-fold jump on DeepSWE is not a new base model. Gains shaped like this come from agent post-training: better tool-error recovery, better context policy, reinforcement on the harness the benchmark runs in. That is real and it is useful, and it is also exactly the kind of gain that can be harness-specific, which is why the caveats above matter and why the preview numbers are the floor, not these.
The kill line (斩杀线)
The term comes from Chinese gaming: the 斩杀线 is the health threshold below which a target can be executed outright. Applied to models, the idea is that DeepSeek's price-to-capability ratio draws a line, and any model that is both weaker and more expensive falls below it and has no reason to be chosen. The lever is price, and the gap is not small:
| Per 1M tokens | v4-flash | v4-pro | GPT-5.6 Sol | pro is cheaper by |
|---|---|---|---|---|
| input, cache miss | $0.22 / $0.44 | $0.66 / $1.32 | $5.00 | 7.6x / 3.8x |
| input, cache hit | $0.007 / $0.014 | $0.022 / $0.044 | $0.50 | 23x / 11x |
| output | $0.66 / $1.32 | $1.98 / $3.96 | $30.00 | 15x / 7.6x |
DeepSeek cells read off-peak / peak, on the card in force since 2026-08-16 16:00 UTC; they are a conversion of the RMB card (pro: ¥4.5 / ¥0.15 / ¥13.5 per 1M off-peak, double at peak) at one consistent rate. GPT-5.6 Sol prices are from OpenAI's own listing. The pricing page has the full schedule.
These ratios were 11.5x / 138x / 34.5x on the flat card that ran until 2026-08-16. The cache-hit column is where the argument lived – a 138x edge is what made replayed context, parallel reviewers and long tool loops nearly free – and it is now 23x off-peak, 11x at peak. That is still a large advantage. It is no longer a different category, and any plan that was built on the old number should be re-costed rather than assumed.
The sober version matters as much as the slogan. The kill line is real for the middle of the market: a model that costs more than V4-Pro and scores below it on the table above is hard to justify, and that is most of the field. It is not real for the frontier. On the hardest multi-step work the strongest closed models still finish in fewer turns and need less steering, and per-attempt reliability is a thing you can measure in wall-clock and interventions, not just in dollars. DeepSeek does not have to win per attempt to win per dollar, and it does not have to win per dollar to lose the one task where getting it right the first time is the whole job.
What it means in practice
The economics only pay off if the workflow is built for them. Three moves, each of which this CLI is shaped to support:
- Route by role. Use
deepseek-v4-profor planning, ambiguous changes, security review and recovery; letdeepseek-v4-flashdo bounded implementations and parallel work.ds chat -m deepseek-v4-proand the Anthropic remap make the switch one flag. - Structure prompts for the cache. A cached input token
costs about 1/30th of an uncached one. Keep the system prompt, tool schemas,
repository map and durable instructions in an identical prefix and put the
volatile part last;
ds usagereports what the cache saved so you can see whether it is working. This is where the cache-hit column turns from a table cell into a bill – and since the repricing hit that column hardest, it is worth more attention now, not less. - Schedule what can be scheduled. The same call costs
half as much outside 01:00–04:00 and 06:00–10:00 UTC. Batch
evaluation, bulk review and overnight agent runs are exactly the workloads
that can move;
ds pricingsays which period you are in. - Measure successful-task cost, not per-token price. A cheaper model that retries five times can cost more than a dearer one that lands first. The ledger stores exact token counts per call, so you can price a whole task under any rate card, including the one that has not been announced yet.
Sources
The launch table and the preview deltas are transcribed from DeepSeek's official agent-benchmark chart, published on the Models & Pricing page and circulated on 2026-08-12; the extended variant carrying the Kimi K3 and GLM-5.2 columns was the widest copy available. Rate-card figures are DeepSeek's own and OpenAI's own. The independent-evaluation caveat refers to the NIST/CAISI review of the V4-Pro preview. Numbers change fast and vendor charts are vendor charts; cross-check a live leaderboard before betting on a single cell.