DeepSeek V4 Pro 0813: Reading DeepSeek's Own Benchmark Table
DeepSeek quietly shipped DeepSeek-V4-Pro-0813 this week with no launch post, just an updated API pricing page and a benchmark table. Here's what that self-reported table actually shows against Preview, Opus 4.8, and Fable 5, and what it costs to run.
DeepSeek V4 Pro 0813: Reading DeepSeek's Own Benchmark Table
DeepSeek shipped the general-availability build of V4 Pro on 12 August 2026 with no blog post and no changelog entry of note. The only signal was the API pricing page quietly swapping the model version string from plain DeepSeek-V4-Pro to DeepSeek-V4-Pro-0813, alongside a benchmark table DeepSeek published itself comparing the new build against its own Preview and against Flash. That table is the actual subject here, not third-party chatter about it.
Worth being upfront about what this is: DeepSeek's own numbers, on DeepSeek's own harness. That's not a knock, every lab reports its own launch benchmarks this way, but it means these are vendor-reported scores rather than an independent, reproduced ranking. Treat the table as a strong signal of what changed between Preview and 0813, not as a settled verdict against every other model on the market.
What DeepSeek's Table Actually Shows
The headline in DeepSeek's own release numbers is the jump from V4 Pro Preview to the 0813 build, particularly on agentic and coding benchmarks:
- Terminal Bench 2.1: 72.1 → 87.9
- CyberGym: 52.7 → 83.3
- DeepSWE: 12.8 → 62.7
- AutomationBench (Public): 12.8 → 31.8
- DSBench-FullStack: 41.8 → 71.1
- DSBench-Hard: 31.1 → 67.2
That's not an incremental refresh. On DeepSWE specifically, the model went from barely functional to competitive in one release. DeepSeek's own table also places 0813 against Claude Opus 4.8 and Claude Fable 5 on the same benchmark set: it comes out ahead of Opus 4.8 on Terminal Bench 2.1, Cybergym, DeepSWE, and AutomationBench, roughly ties Opus on Agents' Last Exam, and sits behind Opus on HLE, NL2Repo, Toolathlon-Verified, and both DSBench tracks. Against Fable 5, DeepSeek trails on most shared benchmarks but the gap on Terminal Bench 2.1 is close to nothing (87.9 versus 88.0), and DeepSeek actually edges ahead on Cybergym and AutomationBench.
The honest reading of DeepSeek's own table is: not uniformly better than Opus or Fable, genuinely close on a specific cluster of agentic and coding benchmarks, and a substantial jump over its own Preview build across the board.
What It Costs
The pricing carried over unchanged from Preview into the 0813 GA build: $0.435 per million input tokens on a cache miss, $0.003625 per million on a cache hit, and $0.87 per million output tokens, against a 1,048,576-token context window and a maximum output of 384,000 tokens. Under the hood it's a mixture-of-experts model with 1.6 trillion total parameters and roughly 49 billion active per token, pretrained on more than 32 trillion tokens.
Put next to the benchmark table above, that's the actual story: a model landing within a few points of Fable 5 on Terminal Bench 2.1 and ahead of it on Cybergym and AutomationBench, at a fraction of what a frontier-tier API charges per token. DeepSeek has flagged that a price increase is coming at some point, without a date attached, so the numbers above are current as of the 0813 release rather than guaranteed to hold.
The Price War This Sits Inside
The 0813 release doesn't land in a vacuum. Frontier-tier API pricing has been in open freefall through 2026, and DeepSeek is a large part of why. One tracking site puts DeepSeek V4 Pro's output-token price at roughly 1/57th of Fable 5's, and V4 Flash undercuts even that: around $0.14 per million input tokens and $0.28 per million output, against $5 and $25-30 for the Opus and GPT-5.6-class flagships it competes with on several benchmarks. OpenRouter's own usage data shows the effect on developer behaviour directly: Chinese-origin models went from under 2% of token volume on the platform in late 2024 to over 50% by June 2026.
The Western labs have been repricing in response rather than holding the line. Anthropic cut Claude pricing by roughly two-thirds in a single announcement earlier this year and later shipped Opus 5 at half the price of Fable 5 - $5/$25 versus Fable's $10/$50 - for coding performance that comes close to matching it. OpenAI cut its GPT-5.6 Luna tier by 80% and Terra by 20% on 30 July, three weeks after the family's general availability launch, though the flagship Sol tier held its $5/$30 rate. Google positioned Gemini 3.6 Flash as cheaper per task than Kimi K3. And xAI's own marketing for Grok 4.6 leans on the same framing, pitching it as "half the price of other frontier models" against comparable output.
That's the pattern DeepSeek V4 Pro 0813 sits inside: a genuine race to the bottom on price at the frontier, with DeepSeek setting the floor and everyone else repricing downward to stay in the conversation. Worth noting the caveat that applies to any of this, including DeepSeek: sticker price isn't the full picture once caching, batching, and long-context surcharges are factored in, and a couple of labs (Anthropic among them) have so far held a premium tier on the argument that safety and precision justify it rather than chasing the floor.
The Caveats Worth Sitting With
A few things worth not glossing over before treating this as a settled result:
- It's a vendor-reported table. DeepSeek ran these numbers on its own harness against its own Preview build and the two comparison models. That's standard practice for a launch benchmark, but it's not the same as an independent, reproduced ranking, and agentic benchmarks in particular are sensitive to prompt harness, tool environment, and reasoning effort settings that aren't fully specified in the table.
- No vision support. V4 Pro 0813 is text-only. If your agentic workflow needs to read a screenshot or a diagram as part of the task, this isn't the model for that step.
- The release itself was unusually quiet. No blog post, no changelog entry, just an API pricing page and model version string update. Worth knowing if you're used to labs shipping benchmark write-ups alongside a release; here the table is effectively the announcement.
Data Handling, Since It's Not in the Benchmark Table
One thing DeepSeek's own launch material doesn't cover: data retention. DeepSeek's privacy policy permits training on prompt and completion data submitted through the official API, and the current OpenRouter listing for this model requires opting into "allow paid endpoints that train on request data" to use it at all, since no zero-data-retention endpoint exists for this release yet. That's not part of the benchmark story, but it's part of the model, and worth knowing before sending anything sensitive through it.
The Read
DeepSeek's own table shows a real jump from Preview to the 0813 build, a model that's competitive with Opus 4.8 on a specific cluster of agentic benchmarks and within a couple of points of Fable 5 on some of them, at a fraction of frontier-tier pricing. That's a genuinely strong result on DeepSeek's own numbers. It's also a vendor-reported table from a release that shipped without any independent write-up alongside it, so it's worth treating as a strong signal to go try the model yourself rather than a settled ranking.
References
- DeepSeek Ships V4 Pro as Its Flagship Model Leaves Preview - Unite.AI
- DeepSeek V4-Pro-0813 Appears in the Price List First - Digital Applied
- DeepSeek V4 Pro 0813: Opus 4.8 and Fable 5 Agent Benchmarks
- DeepSeek V4 Pro 0813 - API Pricing & Benchmarks - OpenRouter
- DeepSeek V4 Pro 0813 vs Fable 5: Fact-Check - Kingy AI
- Claude Fable - Anthropic
- Anthropic Launched Claude Opus 5 at Half the Price of Its Most Powerful AI Model - Quartz
- GPT-5.6: Frontier Intelligence That Scales With Your Ambition - OpenAI
- Why Anthropic, OpenAI, and DeepSeek Face a Pricing Collapse - Milk Road
- AI Price War 2026: Why the Real Winner Isn't DeepSeek or OpenAI - Memeburn
- The Frontier AI Price Wars Continue - Contrary Research
- The 2026 AI API Price War: Chinese Models vs OpenAI & Anthropic - Pickurai
Related reading
Jayden Lee
Founder of Proanalytica Technologies. Machine learning engineer and software developer based in Sydney, NSW. Helping Greater Sydney small businesses build better digital infrastructure.
Need help with your Sydney business?
From web design and WordPress maintenance to ServiceM8 setup and AI automation — we work with Greater Sydney SMBs.
Get in Touch