DeepSeek's Thinking Upgrade: One Model, Two Minds

DeepSeek's Thinking Upgrade: One Model, Two Minds

No press conference. No demo video. Late Monday night, DeepSeek simply announced that its online models had been upgraded — and that the new version ships two modes in one model: a thinking mode and a non-thinking mode, available to users right now.

Most people will read that as a small changelog line. It is not. It is the quiet end of a two-week story in which the smallest model in DeepSeek's new lineup did something unusual: it retired the company's own flagship.

The Smallest Model That Killed the Flagship

On September 10, DeepSeek released V4.1 Flash. In the release notes, the company made a claim that rarely appears in a model launch: V4.1 Flash, it said, surpasses V4 Pro on performance, cost, speed, and total time — and V4 Pro would be retired.

A Flash model. The "cheap, fast" tier. Outperforming the Pro tier across the board, then watching the Pro tier get shut down.

The benchmark sheet backs the claim. GPQA Diamond: 90.9. Codeforces rating: 3471. MathArena Apex: 65.6. DeepSWE v1.1: 74.2. Terminal-Bench 2.1: 90.6. CyberGym: 88.1. Vision-heavy agent benchmarks like Chartography and BabyVision land at 78.9 and 89.6 with tools — because this is the first Flash with native multimodal visual understanding. It does not just read text; it looks at images, charts, and interfaces and acts on what it sees.

All of it on a new asymmetric architecture built for a higher capability ceiling, faster inference, higher throughput — and lower prices. Old Flash models were retired too, with their names temporarily routed to V4.1 Flash so nobody's code breaks.

DeepSeek V4.1 Flash — smallest model outranks and retires the flagship

One Model, Two Minds

Now the upgrade lands at the product level: the same model offers thinking and non-thinking behavior, and ordinary users — the people who just open the app and type — can switch between them.

Thinking mode is the reasoning path: it plans, checks, and works through the problem before answering. Non-thinking mode is the fast path: straight answer, low latency, ideal for the daily bulk of requests that do not need a chain of thought.

This is not a gimmick split. DeepSeek has been building toward it for months. Since the V4 generation went GA in August, thinking effort has been adjustable in three levels — low, high, and max — so a task can choose how long it thinks before it speaks. And with V3.2, released in late September, DeepSeek became the first open model to fold thinking into tool calling itself: the model can reason about which tool to use, and why, while it is using it — earning gold in IMO, CMO, ICPC, and IOI along the way.

Put together, the direction is clear: thinking used to be a premium toggle for power users. Now it is on by default for everyone, with an off switch for the cases that do not need it.

DeepSeek thinking and non-thinking modes — winding thought path versus straight fast path

What It Actually Changes for You

For users: the app you already have now answers questions it previously fumbled — the ones that need a model to stop, think, and reconsider before replying.

For developers: prices went down with the V4.1 Flash release, the old model names still resolve, and one API call now controls how much thinking a task gets.

For the open-source world: this is a reminder of how DeepSeek keeps redefining what "cheap" means. Two weeks ago, the company's smallest model outran its flagship and retired it. This week, thinking mode — the feature rivals sell as a premium tier — became a default in a free chat app.

The changelog line was short. The story under it is two weeks of a company making its flagship obsolete on purpose, then handing the upgrade to every user for free.

References:

[1] DeepSeek release notes — DeepSeek-V4.1-Flash release and API changes, api-docs.deepseek.com/updates, 2026-09-10.

[2] DeepSeek official news — DeepSeek-V4.1 Flash: stronger, faster, more accessible, deepseek.com/news, 2026-09-10.

[3] DeepSeek V4 technical documentation — model card and architecture, fe-static.deepseek.com, 2026-04-27.

[4] DeepSeek-V3.2 official release notes — thinking in tool calling, api-docs.deepseek.com, 2025-12-01.

[5] All illustrations are AI-generated.