No press conference. No demo video. Late Monday night, DeepSeek simply announced that its online models had been upgraded — and that the new version ships two modes in one model: a thinking mode and a non-thinking mode, available to users right now.
Most people will read that as a small changelog line. It is not. It is the quiet end of a two-week story in which the smallest model in DeepSeek's new lineup did something unusual: it retired the company's own flagship.
The Smallest Model That Killed the Flagship
On September 10, DeepSeek released V4.1 Flash. In the release notes, the company made a claim that rarely appears in a model launch: V4.1 Flash, it said, surpasses V4 Pro on performance, cost, speed, and total time — and V4 Pro would be retired.
A Flash model. The "cheap, fast" tier. Outperforming the Pro tier across the board, then watching the Pro tier get shut down.
The benchmark sheet backs the claim. GPQA Diamond: 90.9. Codeforces rating: 3471. MathArena Apex: 65.6. DeepSWE v1.1: 74.2. Terminal-Bench 2.1: 90.6. CyberGym: 88.1. Vision-heavy agent benchmarks like Chartography and BabyVision land at 78.9 and 89.6 with tools — because this is the first Flash with native multimodal visual understanding. It does not just read text; it looks at images, charts, and interfaces and acts on what it sees.
All of it on a new asymmetric architecture built for a higher capability ceiling, faster inference, higher throughput — and lower prices. Old Flash models were retired too, with their names temporarily routed to V4.1 Flash so nobody's code breaks.

One Model, Two Minds
Now the upgrade lands at the product level: the same model offers thinking and non-thinking behavior, and ordinary users — the people who just open the app and type — can switch between them.
Thinking mode is the reasoning path: it plans, checks, and works through the problem before answering. Non-thinking mode is the fast path: straight answer, low latency, ideal for the daily bulk of requests that do not need a chain of thought.
This is not a gimmick split. DeepSeek has been building toward it for months. Since the V4 generation went GA in August, thinking effort has been adjustable in three levels — low, high, and max — so a task can choose how long it thinks before it speaks. And with V3.2, released in late September, DeepSeek became the first open model to fold thinking into tool calling itself: the model can reason about which tool to use, and why, while it is using it — earning gold in IMO, CMO, ICPC, and IOI along the way.

