>thevibeworks/deepseek-cli

Benchmarks

Where deepseek-v4-pro and deepseek-v4-flash sit against the field, what the 0813 checkpoint changed, and the economics argument the Chinese community calls the 斩杀线, the kill line. The numbers below are DeepSeek's own launch-day figures unless marked otherwise; read the caveats first.

read this before quoting a number
  • Vendor numbers, one harness. The table is DeepSeek's own agent-benchmark chart, published 2026-08-12 with the GA. A score is a model-and-harness result, not a model-only one, and no independent same-harness run of 0813 exists yet. Treat these as the claim, not the verdict.
  • No GPT-5.6 column. DeepSeek's chart compares against Kimi K3, GLM-5.2, Claude Opus 4.8 and Fable 5 only. GPT figures elsewhere on this page are drawn from those vendors' own releases and are cross-vendor, so they are looser still.
  • Kimi and GLM are single-sourced. Those two columns appear only on the extended variant of the chart; the shared columns are identical across every copy, so the numbers are consistent, but the Kimi and GLM rows rest on one source.
  • Independent history says be careful. The one held-out check on the V4-Pro preview, from NIST/CAISI, put it closer to GPT-5 and roughly eight months behind the frontier, below its self-reported position. No equivalent 0813 evaluation exists yet.

The launch table

Higher is better. HLE is shown as without-tools / with-tools. Every figure is from DeepSeek's GA chart of 2026-08-12; a dash means the vendor did not report it.

Benchmark V4-Pro 0813 V4-Flash 0731 Kimi K3 GLM-5.2 Opus 4.8 Fable 5
Terminal-Bench 2.187.982.788.381.685.088.0
DeepSWE62.754.467.546.258.070.0
Toolathlon-Verified74.170.376.559.976.277.9
CyberGym83.376.780.078.383.1
NL2Repo61.554.248.969.7
AutomationBench31.825.130.812.927.229.1
DSBench-FullStack71.168.763.051.871.677.2
DSBench-Hard67.259.673.754.571.768.3
Agents' Last Exam25.725.224.523.925.7
HLE (no / tools)42.7/60.037.8/51.543.5/56.040.5/54.749.8/57.953.3/63.0

The honest read of the row-by-row: V4-Pro is within a point of Fable 5 on Terminal-Bench and CyberGym, ahead of Opus 4.8 on Terminal-Bench, DeepSWE, CyberGym and AutomationBench, and behind Kimi K3 on Terminal-Bench, DeepSWE, Toolathlon and DSBench-Hard. It is in the cluster, not clear of it. On knowledge without tools (HLE) the closed models still lead. Flash trails Pro across the board but stays remarkably close for a fifth of the price, which is the whole point of the next two sections.

What GA changed

The model ID stayed deepseek-v4-pro and the rate card did not move. What moved is the checkpoint. Against the April preview, DeepSeek's own chart shows the gains landing almost entirely in the agentic and SWE suites:

BenchmarkPreviewGA 0813Delta
DeepSWE12.862.7+49.9
DSBench-Hard31.167.2+36.1
CyberGym52.783.3+30.6
DSBench-FullStack41.871.1+29.3
NL2Repo38.561.5+23.0
Toolathlon-Verified55.974.1+18.2
Terminal-Bench 2.172.187.9+15.8

A near five-fold jump on DeepSWE is not a new base model. Gains shaped like this come from agent post-training: better tool-error recovery, better context policy, reinforcement on the harness the benchmark runs in. That is real and it is useful, and it is also exactly the kind of gain that can be harness-specific, which is why the caveats above matter and why the preview numbers are the floor, not these.

The kill line (斩杀线)

The term comes from Chinese gaming: the 斩杀线 is the health threshold below which a target can be executed outright. Applied to models, the idea is that DeepSeek's price-to-capability ratio draws a line, and any model that is both weaker and more expensive falls below it and has no reason to be chosen. The lever is price, and the gap is not small:

Per 1M tokensv4-flashv4-proGPT-5.6 Solpro is cheaper by
input, cache miss$0.22 / $0.44$0.66 / $1.32$5.007.6x / 3.8x
input, cache hit$0.007 / $0.014$0.022 / $0.044$0.5023x / 11x
output$0.66 / $1.32$1.98 / $3.96$30.0015x / 7.6x

DeepSeek cells read off-peak / peak, on the card in force since 2026-08-16 16:00 UTC; they are a conversion of the RMB card (pro: ¥4.5 / ¥0.15 / ¥13.5 per 1M off-peak, double at peak) at one consistent rate. GPT-5.6 Sol prices are from OpenAI's own listing. The pricing page has the full schedule.

the repricing moved this line

These ratios were 11.5x / 138x / 34.5x on the flat card that ran until 2026-08-16. The cache-hit column is where the argument lived – a 138x edge is what made replayed context, parallel reviewers and long tool loops nearly free – and it is now 23x off-peak, 11x at peak. That is still a large advantage. It is no longer a different category, and any plan that was built on the old number should be re-costed rather than assumed.

The sober version matters as much as the slogan. The kill line is real for the middle of the market: a model that costs more than V4-Pro and scores below it on the table above is hard to justify, and that is most of the field. It is not real for the frontier. On the hardest multi-step work the strongest closed models still finish in fewer turns and need less steering, and per-attempt reliability is a thing you can measure in wall-clock and interventions, not just in dollars. DeepSeek does not have to win per attempt to win per dollar, and it does not have to win per dollar to lose the one task where getting it right the first time is the whole job.

What it means in practice

The economics only pay off if the workflow is built for them. Three moves, each of which this CLI is shaped to support:

Sources

The launch table and the preview deltas are transcribed from DeepSeek's official agent-benchmark chart, published on the Models & Pricing page and circulated on 2026-08-12; the extended variant carrying the Kimi K3 and GLM-5.2 columns was the widest copy available. Rate-card figures are DeepSeek's own and OpenAI's own. The independent-evaluation caveat refers to the NIST/CAISI review of the V4-Pro preview. Numbers change fast and vendor charts are vendor charts; cross-check a live leaderboard before betting on a single cell.