DeepSeek V4-Flash Public Beta Arrives for Coding Agents
DeepSeek has put the official V4-Flash API into public beta. The update keeps the same model name but adds a re-post-trained release aimed at agentic coding and tool use, while the app, web product and V4-Pro remain unchanged.
DeepSeek has moved the official V4-Flash API into public beta with an updated release it calls DeepSeek-V4-Flash-0731. The company says the update is aimed at stronger agent capabilities, while keeping the same API model name: deepseek-v4-flash.
For developers, this is a practical update rather than a brand-new consumer chatbot. DeepSeek says the change applies to the V4-Flash API; its V4-Pro API and its app and web models are unchanged.
What changed in V4-Flash
DeepSeek says V4-Flash-0731 keeps the same architecture and size as the earlier V4-Flash preview and was re-post-trained. The company has put the official API release into public beta, so existing API integrations can call the current version with the same model identifier.
The update adds native support for the Responses API format and is positioned for coding-agent workflows. DeepSeek also lists support for tool calls, JSON output, an Anthropic-compatible API format and both thinking and non-thinking modes.
According to DeepSeek’s current documentation, V4-Flash has a one-million-token context length and a maximum output limit of 384,000 tokens. Those are API capabilities, not a promise that every application will expose the same controls or limits.
Price is part of the story
DeepSeek’s API pricing page lists V4-Flash at $0.14 per million cache-miss input tokens, $0.0028 per million cache-hit input tokens and $0.28 per million output tokens. That is why the release is attracting attention: capable coding models are increasingly competing not only on benchmarks, but on the cost of running high-volume automated tasks.
Price alone is not a quality score. Developers still need to test reliability, privacy terms, rate limits, tool behaviour and failure handling in their own environment. But lower per-token costs can change which tasks are economical to automate or run at scale.
Benchmark claims need context
DeepSeek published a set of agent and coding benchmark results alongside the update, including Terminal Bench 2.1, NL2Repo and DeepSWE. Those figures are company-reported results, not an independent ranking, and some of the cited tests use DeepSeek’s own harness or internal datasets.
That does not make the release unimportant. It means teams should treat the figures as a reason to evaluate V4-Flash, rather than as a substitute for their own tests against their actual repositories, tools and constraints.
What has not changed
DeepSeek explicitly says the public-beta update upgrades the V4-Flash API only. V4-Pro has not been upgraded in this release, and neither have the company’s app and web models. Developers using the API should check their model selection, usage controls and billing assumptions before deploying the new release broadly.
Why it matters
AI coding is moving toward a market where model capability, tool integration and unit cost all matter together. DeepSeek’s V4-Flash beta gives developers another current option to benchmark for agentic work—and underscores how quickly API providers are updating models without necessarily changing the name an integration calls.
Sources
- DeepSeek API Change Log: V4-Flash public-beta update
- DeepSeek API Models & Pricing
- Axios: DeepSeek’s new bargain model and AI pricing pressure
Source: Axios
