DeepSeek shipped the production version of its flagship model this week with no blog post and no announcement. The tell was a table cell: the model version listed for deepseek-v4-pro on the API pricing page now reads DeepSeek-V4-Pro-0813, the general-availability release of a model that had been running as a preview since April, as Decrypt reported Tuesday.
Pricing Stays Aggressive
The API rates carry over unchanged: $0.435 per million input tokens on a cache miss, $0.003625 per million on a cache hit, and $0.87 per million output tokens, with a one-million-token context window and a maximum output of 384,000 tokens. Against Anthropic's Claude Fable 5 at $10 per million input and $50 per million output, the blended rate gap is roughly 46 times — about $30 versus $0.65. DeepSeek's own comparison table puts Claude Fable 5 ahead by an average of 5.3 percent across nine agent benchmarks, with DeepSeek winning two of them.
Quiet Rollout, Loud Signal
The 0813 build appeared on OpenRouter's model page on August 12, and DeepSeek's documentation now lists it behind the deepseek-v4-pro endpoint with a concurrency limit of 500 against 2,500 for the smaller Flash model. DeepSeek had signaled the move on July 31, when it graduated V4-Flash to official status and noted the Pro API was unchanged with an official release to follow soon. The Hugging Face model card still describes the V4 series as a preview, and DeepSeek has not published 0813 weights or a timeline for doing so.
One caveat for developers planning on the current rates: a notice on DeepSeek's pricing page warns that the company plans a significant overall increase in API pricing in the near future, with specifics to come by official notice. The listed V4 Pro rates hold for now, but teams budgeting around the current floor should treat it as temporary, the notice suggests.
Comments (0)
Log in or sign up to leave a comment.
No comments yet. Be the first to share your thoughts.