DeepSeek V4.1 Flash thumbnail showing open weights, MIT license and 1M context

DeepSeek V4.1 Flash Cuts Prices and Retires V4-Pro: What Changes for You

DeepSeek V4.1 Flash is the Chinese lab’s newest model, and it does something unusual: it’s cheaper than the flagship it replaces, yet DeepSeek says it scores higher. It launched on September 10, 2026, with open weights, image understanding and a 1 million token context window.

Here’s the short version, then the details that actually affect your bill and your workflow.

  • Released September 10, 2026, with weights on Hugging Face under an MIT license.
  • API output costs $0.60 per million tokens off-peak, and double that at peak.
  • Context window is 1 million tokens, with native image input.
  • V4-Pro API calls redirect to V4.1 Flash from September 14, 2026.

What DeepSeek V4.1 Flash actually changes

The headline move is a merger. DeepSeek retired its old top model, V4-Pro, and pointed everything at the Flash line. According to The Next Web, V4-Pro requests are rerouted to V4.1 Flash at Flash pricing starting 04:00 UTC on September 14.

If you built something on V4-Pro, you don’t have to do anything. Your app keeps working, but it now talks to a different model. Retest your prompts, because outputs may change.

The model is a mixture-of-experts design. It activates about 8 billion parameters when reading your prompt and 16 billion when writing the answer. Reports differ on the total size: SiliconANGLE and The Next Web say 552 billion, while ProPakistani says 748 billion. Treat the total as unconfirmed until DeepSeek’s model card settles it.

DeepSeek V4.1 Flash pricing, in plain numbers

DeepSeek charges less during off-peak hours and doubles the rate on weekday mornings (UTC). The table below uses the API rates reported at launch.

Token type (per million) Off-peak Peak
Cached input $0.003 $0.006
Uncached input $0.15 $0.30
Output $0.60 $1.20

That’s cheap next to Western frontier models. For a sense of the gap, see our breakdown of GPT-6.1 Sol pricing and benchmarks.

There’s a catch on trust. DeepSeek quadrupled its prices in August before this cut, so the cheap rate isn’t guaranteed to last. Don’t lock a product’s margins to it.

How good is it? Benchmarks to take with salt

DeepSeek’s own numbers are strong, but nobody independent has confirmed them yet. Per the reports above, it scores 74.2 on the DeepSWE coding test, against 74.0 for Claude Opus 5 and 62.7 for V4-Pro. On Terminal-Bench 2.1 it reports 90.6.

It isn’t ahead everywhere. The Next Web notes a Humanity’s Last Exam score of 36.8, well below Opus 5’s 56.3. In plain terms, it looks excellent at agent-style coding and weaker on hard general reasoning.

Vendor-reported scores often shift once outside testers get hold of a model. Wait for independent leaderboards before switching anything important.

Why the memory trick behind it matters

Long chats are expensive because the model has to keep a running memory of everything said, called the KV cache. DeepSeek says V4.1 Flash needs only 890 bytes per token for it, roughly a quarter of the earlier V4-Flash, according to SiliconANGLE.

Smaller memory per token means more users fit on the same GPU. That’s the real reason the price can drop without the company losing money, at least in theory.

It also explains why the cached-input price is tiny, at $0.003 per million tokens off-peak. If your app repeats a long system prompt or a big document, caching can cut costs far more than the headline rate suggests.

Open weights and long context: who benefits

The MIT license is the permissive kind. You can download the weights, modify them and run them yourself, including commercially. That matters for companies that can’t send data to a Chinese-hosted API.

Running it locally still takes serious hardware, so most people will use the API or a third-party host. Open weights simply give you the exit option. For another open-model contender, read our post on Reflection AI’s Beam.

The 1 million token window, with output up to 384,000 tokens according to ProPakistani, suits whole-codebase reviews and long document work. Image input handles screenshots and diagrams, up to 1344×1344 pixels.

Who should switch now, and who should wait

Developers running coding agents, bulk summaries or document extraction have the most to gain. Those jobs burn through tokens, so a lower rate shows up on the invoice fast.

If you handle regulated or customer data, slow down. Check where the hosted API processes requests and what your company policy allows. Self-hosting the open weights or using a trusted third-party host may be the safer route.

Everyone else can simply try the free web app and compare it with the assistant you already use. A ten-minute side-by-side on your own tasks tells you more than any leaderboard.

Frequently Asked Questions

Is DeepSeek V4.1 Flash free to use?

It’s live in DeepSeek’s web and mobile apps, and the API is paid per token. The weights are free to download under the MIT license.

What happened to DeepSeek V4-Pro?

It was retired. From September 14, 2026, API calls to V4-Pro are redirected to V4.1 Flash and billed at Flash rates.

Is DeepSeek V4.1 Flash better than Claude Opus 5?

On some coding tests, DeepSeek’s own figures put it level. On hard reasoning it trails by a wide margin, and the scores are unverified.

Can I run it on my own computer?

The weights are public, but a model this size needs server-class hardware. Most people will be better off using a hosted API.

If you write code with AI agents, test V4.1 Flash on a real task this week. It’s cheap enough that a trial costs pennies. Just keep your prompts portable, because a vendor that quadruples prices once can do it again.


Comments

Leave a Reply

Your email address will not be published. Required fields are marked *