09/11 2026
371

Text by Ai Ti
Edited by Shen Xiao
DeepSeek has once again rolled out a model update in the wee hours.
On September 8, keen-eyed users spotted a temporary model on the DeepSeek open platform: deepseek-v4.1-flash-expires-on-0910. Public reports indicate that this is an intermediate version of V4.1 Flash, made available for testing to a select group of users. The method for invoking this model is largely in line with the existing API, with billing temporarily based on V4 Flash and a concurrency limit of 20 per account.
The model's name speaks volumes.
This isn't your run-of-the-mill official version; it's a test version with a built-in “expiration date.” September 10 isn't a confirmed official release date but rather the cutoff time for this round of internal testing interfaces. In essence, DeepSeek has given testers less than 48 hours to evaluate whether this model update merits further exploration using real-world tasks.
The notice highlights that “V4.1 Flash outperforms V4 Pro across the board in terms of performance, cost, speed, and overall processing time.”
“Cost, speed, and overall processing time” are metrics that are more closely tied to the model's production efficiency than traditional Benchmark scores.
As of now, DeepSeek has yet to disclose the complete parameters, test results, or official pricing for V4.1 Flash on its official website, API update logs, or public model cards.

When DeepSeek released the V4 preview version in April this year, it clearly delineated the roles of the two versions.
V4-Pro is designed to tackle more complex reasoning and Agent tasks, while V4-Flash offers faster and more cost-effective services with a smaller parameter scale. Officially, V4-Pro boasts approximately 1.6 trillion parameters, with 49 billion activated; V4-Flash, on the other hand, has about 284 billion total parameters, with 13 billion activated.
Under this framework, Pro resembles a flagship model, while Flash functions more as a high-frequency workhorse model.
However, on July 31, when DeepSeek updated the official version of V4-Flash, it began to blur these lines. The company claimed that V4-Flash outperformed the V4-Pro preview version in multiple Code Agent tests, even though the V4-Pro API wasn't upgraded simultaneously at that time.
On August 13, the official version of V4-Pro was launched, with a focus on enhancing Agent capabilities and supporting more flexible thinking intensity and Responses API.
From a product development standpoint, DeepSeek continues to position Pro as the pinnacle of capability while gradually expanding Flash's reach into Pro's domain.
This is a pragmatic approach to model commercialization.
For developers and enterprises, the “strongest” model isn't the sole deciding factor. What truly influences procurement and invocation decisions are the cost per task, waiting time, stability of completion, and whether the model can be seamlessly integrated into production systems in batches.
If a Flash model can handle most daily coding, documentation, tool invocation, and simple Agent tasks, there's no need to invoke the more expensive Pro for every request. Even if it falls short in a few complex tasks, the overall usage cost may still be lower.
DeepSeek's official pricing reveals that the current off-peak input cache miss price for V4-Flash is 1.5 yuan per million Tokens, with an output price of 4.5 yuan; for V4-Pro, the corresponding prices are 4.5 yuan and 13.5 yuan, respectively, doubling during peak hours.
This means that as long as Flash approaches Pro's performance in enough tasks, it has the potential to become the default model.
What V4.1 Flash is truly competing for may not be the title of the “strongest model” but rather the “model that will be invoked by default for the most requests.”

The most captivating aspect of this internal test isn't the model's name but how DeepSeek has involved users in the model decision-making process.
Public information indicates that the notice asked users whether the intermediate version of V4.1 Flash could fully replace the current online V4 Pro. This question underscores that DeepSeek is testing not just model scores but the model's comprehensive performance in real-world scenarios.
Real-world environments are far more intricate than Benchmarks.
A model may excel in coding tasks but falter in long contexts, multi-round tool invocations, complex format outputs, or image understanding. It might produce high-quality single responses but experience speed degradation, increased Token consumption, or occasional failures to return structured results under high concurrency.
Therefore, this test of V4.1 Flash is at least verifying four key aspects:
First, can it handle more real Agent tasks? Second, is its native multimodality sufficiently stable? Third, does its response quality remain consistent while speed improves? Fourth, can it withstand large-scale API invocations?
If these test results hold up, the product boundaries between Flash and Pro will be redrawn.
In the past, Pro and Flash represented differences in capability tiers; in the future, they may represent differences in task types. Simple tasks, daily coding, and batch invocations will be assigned to Flash, while complex reasoning and high-difficulty Agent tasks will be reserved for Pro.
Furthermore, if Flash can indeed cover most Pro requests, DeepSeek might even reduce the default invocation ratio of Pro, freeing up more computational power for high-value tasks.
This isn't just a model upgrade; it's also a reallocation of costs and computational resources.
Currently, the information circulating in the market primarily comes from official user communication groups, test users, and retellings by tech media, rather than DeepSeek's official website press conferences or formal announcements.
Nevertheless, even if V4.1 Flash doesn't fully replace V4 Pro in the end, this brief internal test sends a clear message:
DeepSeek is shifting the model competition from “who has more parameters and higher rankings” to “who can complete more real tasks at a lower cost.”
By September 10, what the market truly anticipates isn't a declaration of “comprehensive superiority” but whether DeepSeek will turn this 48-hour test into the product rules for the next stage.