DeepSeek V4 Pro 0813: Benchmarks, Price Hike, Specs
DeepSeek moved its V4 Pro 0813 flagship to general availability with big agentic-coding gains and a peak-hour price hike up to 12x. What's verified and what isn't.
DeepSeek has pushed a new checkpoint of its flagship model into production. On Wednesday, August 13, 2026, the Chinese lab moved DeepSeek-V4-Pro-0813 from preview to general availability, updating its official API pricing page the day before to reflect the new version. The 0813 build does not change the model’s name or its openly published weights strategy, but it posts sharply higher scores on agentic-coding tests than the preview it replaces — and it arrives alongside a restructured price sheet that raises rates for the most demanding workloads.
The release lands weeks after DeepSeek’s broader V4 general-availability launch in July and its V4-Flash-0731 agent upgrade. The 0813 checkpoint is the Pro tier’s answer to that cadence: a refreshed flagship aimed squarely at the coding and agent workloads where DeepSeek has been trying to prove that an open-weight model can trade blows with closed Western systems.
What’s in the 0813 build
The underlying architecture is unchanged from the V4 Pro that shipped earlier this year. It remains a mixture-of-experts model with 1.6 trillion total parameters and roughly 49 billion active per token, pretrained on more than 32 trillion tokens. DeepSeek pairs that with a hybrid attention system designed to cut inference cost at long context lengths, and the model supports a 1,048,576-token (1M) context window with up to 384,000 output tokens. There is no vision support; V4 Pro remains a text-and-code model.
If the mixture-of-experts design is new to you, our explainer on what mixture of experts is covers why sparse activation lets a lab advertise trillion-scale parameter counts while keeping per-token compute — and therefore price — comparatively low. That efficiency is the whole reason DeepSeek can price a frontier-scale model the way it does.
The benchmark claims
The headline of the 0813 release is agentic coding, and the reported gains over the preview build are large. According to DeepSeek’s own model card and disclosures:
- DeepSWE rose from 12.8 to 62.7.
- CyberGym rose from 52.7 to 83.3.
- Terminal Bench 2.1 rose from 72.1 to 87.9.
On the widely watched aggregate tests, the 0813 card posts 80.6% on SWE-bench Verified, 90.1% on GPQA Diamond, and 93.5% on LiveCodeBench. DeepSeek framed the improvements as vendor-reported gains of up to 49.9 percentage points on individual agentic tasks.
Those are strong numbers on paper, and they would place V4 Pro at or near the top of the open-weight field for software-engineering work. But an important caveat travels with every one of them: none of the scores has been independently replicated. No third-party evaluator has yet reproduced the agentic-coding jumps or the aggregate figures, and the largest reported gains come from DeepSeek’s internal harness. The results are claims to be verified, not settled facts — a distinction that matters more for agentic benchmarks than for older static tests, because small differences in scaffolding, tool access, and retry policy can swing agentic scores by tens of points.
The pricing changes
Alongside the model, DeepSeek published a revised price sheet. As of the 0813 general-availability listing, V4 Pro is priced at $0.435 per million input tokens on a cache miss, $0.003625 per million on a cache hit, and $0.87 per million output tokens, all at the 1M-token context tier.
The more consequential change is scheduled to take effect at 00:00 Beijing time on August 17, 2026, when DeepSeek switches on peak and off-peak rates. Under that structure, V4 Pro rates rise by as much as 12 times the previous level during peak windows. That is a steep escalation from the 2× peak surcharge DeepSeek introduced at the July V4 general-availability launch, and it reframes how buyers should think about the model’s famous cheapness.
Peak-hour pricing is still unusual for a frontier API — closer to how cloud providers price spot capacity or how utilities price electricity — and a 12× peak multiplier is a strong signal that DeepSeek is managing a real capacity constraint rather than pricing purely to win share. The likely root cause is the same one that shadows every Chinese lab: limited access to the most advanced accelerators under U.S. export controls, which caps how much peak inference DeepSeek can serve at once. Rather than degrade latency for everyone, the company is metering demand by the clock.
How to read the release
The competitive logic of the 0813 checkpoint is consistent with DeepSeek’s strategy all year. The lab has leaned on aggressive pricing and openly licensed weights to court cost-sensitive, high-volume workloads — exactly the segment where a large price gap against U.S. flagships compounds fastest. The July GA established V4 as a stable, production-grade target; the 0813 build is the capability refresh meant to keep that target competitive as Western labs cut their own prices, a dynamic we traced in coverage of the OpenAI GPT-5.6 price cut and the AI price war.
The 0813 release also fits the broader pattern of Chinese labs shipping frontier-class systems and giving the weights away, part of the trend our reporting on open-source models closing the gap has followed through 2026. Each major Chinese open-weight release this year has arrived with the same two-sided story: a capability claim that pressures Western providers, and a set of benchmark numbers that the market waits to see reproduced.
What it means
The 0813 checkpoint is best understood as two moves bundled into one release: a capability push and a pricing reset, and they point in opposite directions for buyers.
On capability, the agentic-coding gains — if they hold up — would make V4 Pro a much stronger default for automated software-engineering and agent workflows than the preview it replaces. A jump from 12.8 to 62.7 on DeepSWE is not incremental; it is the difference between a model that mostly fails at multi-step engineering tasks and one that clears a majority of them. But the operative phrase is if they hold up. Until an independent evaluator reproduces the SWE-bench Verified, DeepSWE, and Terminal Bench figures under a documented harness, procurement teams should treat the numbers as a reason to run their own pilots, not as a reason to migrate production traffic on faith. Agentic benchmarks are unusually sensitive to how the test is wired, and vendor-run scores have a structural incentive to look their best.
On pricing, the 12× peak multiplier changes the calculus that made DeepSeek attractive in the first place. The base rates are still far below Western flagships, but a workload that runs during Beijing business hours could see its effective cost balloon. The clear arbitrage is scheduling: batch and asynchronous jobs move into cheap off-peak windows, while latency-sensitive traffic either pays the premium or routes to the MIT-licensed open weights on self-hosted hardware. Expect sophisticated users to split their traffic accordingly, and expect the peak surcharge to accelerate self-hosting among teams with the GPUs to run a 1.6-trillion-parameter MoE.
For the competitive picture, the release tightens the low-end squeeze on Western providers without dislodging the top proprietary models from the frontier of raw capability, and enterprises with data-residency or provenance concerns will still hesitate over a Chinese-origin model. The watch items from here are concrete: whether independent evaluations confirm the agentic-coding claims, how much the August 17 peak pricing pushes workloads onto self-hosted deployments, and whether DeepSeek’s capacity constraint — the thing a 12× surcharge implies — becomes a lasting ceiling on how far it can scale the Pro tier.
Tagged
Keep reading
Chisato · · 6 min read DeepSeek V4-Flash-0731: Benchmarks, Price, What Changed
DeepSeek's retrained V4-Flash-0731 beats its own flagship on nine agent benchmarks at the same $0.14/$0.28 price, with MIT-licensed weights on Hugging Face.
Chisato · · 5 min read DeepSeek V4 Release: Specs, Benchmarks, Peak Pricing
DeepSeek V4 graduates from preview to general availability with two open-weight MoE models, an 80.6% SWE-bench score, and new peak-hour API pricing.
Chisato · · 6 min read Kimi K3: Moonshot's 2.8T Open-Weight Model Explained
Moonshot AI's Kimi K3 is a 2.8-trillion-parameter open-weight model with a 1M-token context, ranking third on GDPval behind only Fable 5 and GPT-5.6.