Deepseek has pushed its budget AI model into a more competitive position with the release of V4 Flash "0731". The new version raises performance, lowers practical task cost, and keeps the same large-scale architecture as the earlier V4 Flash release.
The headline comparison is direct: according to the Artificial Analysis Intelligence Index, V4 Flash "0731" scores 50 points. That is ten more than the previous V4 Flash, which launched in April 2026, and just one point behind OpenAI's budget model GPT-5.6 Luna.
A Budget Model Moves Closer To GPT-5.6 Luna
The new Deepseek V4 Flash result matters because it narrows the gap with a competing budget model while changing the cost equation. GPT-5.6 Luna remains ahead by one point on the cited index, but Deepseek's new model is described as costing about 60 percent less per task.
That comparison is especially notable because it comes even after OpenAI's 80 percent price cut. In other words, the source article presents V4 Flash "0731" not simply as a cheaper model, but as a model that has become much closer in measured quality while keeping a large cost advantage.
For users comparing budget AI systems, the relationship between score and cost is the central point. A small benchmark gap can look different when the lower-scoring system costs materially less to run. The source does not say that Deepseek is better than GPT-5.6 Luna overall, but it does show that the two are now close on the cited index.
Why The Cost Gap Is So Large
The article identifies caching as a major reason for the lower cost per task. Deepseek offers a 98 percent cache discount, which is higher than the industry-standard 90 percent noted in the source.
Cache discounts matter because repeated or reusable context can become cheaper to process. The source does not provide a full pricing table, but it does state that this discount is a big reason Deepseek can land about 60 percent below GPT-5.6 Luna on cost per task.
The model also uses 12 percent fewer tokens than its predecessor. That is another practical efficiency gain: if a model can complete comparable work with fewer tokens, the total task cost can move down even without changing the model's headline capability.
- V4 Flash "0731" scores 50 points on the Artificial Analysis Intelligence Index.
- That is ten more than the previous V4 Flash from April 2026.
- It is one point behind GPT-5.6 Luna.
- It costs about 60 percent less per task.
- It uses 12 percent fewer tokens than its predecessor.
Where The Upgrade Shows Up
Deepseek's new version improves across every tested category compared with the previous V4 Flash. The largest gains are in agentic tasks, the area where models are expected to plan, execute, and adapt across more complex work.
The source highlights GDPval as one benchmark where the jump is clear. GDPval is described as a benchmark designed to test models on complex real-world office work. On that measure, V4 Flash "0731" rises from 1,189 to 1,559 Elo points.
That improvement is significant within the facts given because it points to progress beyond narrow prompt answering. The article frames the biggest movement around agentic tasks, which are closer to workflows where a model must handle several steps rather than produce a single isolated response.
The new model also hallucinates less often, according to the source. No exact hallucination rate is provided, so the important takeaway is directional: the upgrade is described as more reliable than the previous V4 Flash on that measure.
The Architecture Stays The Same
Although the benchmark results improved, the architecture has not changed. V4 Flash "0731" keeps 284 billion total parameters, 13 billion active, and a one-million-token context window.
That continuity matters because it suggests the upgrade is not being presented as a new model family with a different published architecture. The same broad setup now delivers better tested performance, lower token use, and reduced hallucination compared with the predecessor described in the source.
The one-million-token context window remains one of the key specifications. A long context window can support tasks involving large inputs, though the source does not provide examples of how Deepseek expects users to apply it.
Open Weights Add Another Dimension
The model weights are available under an MIT license on Hugging Face. That detail separates the release from models whose weights are not distributed in the same way.
Open weights can matter for evaluation, deployment choices, and experimentation, but the source only states the availability and license. The practical takeaway is that Deepseek is pairing a budget model upgrade with accessible weights, while also competing closely with GPT-5.6 Luna on the cited intelligence index.
For now, the reported picture is straightforward: Deepseek V4 Flash "0731" is a major upgrade over the April 2026 V4 Flash, scores nearly level with GPT-5.6 Luna, and does so at a much lower cost per task. The remaining difference is narrow on the index, while the pricing gap is large in the comparison provided.