What Is DeepSeek R1? Explained

2026-08-04
What is DeepSeek R1? Learn how the DeepThink reasoning model works, how it compares to o1 and o3, and how to access R1 through the API or free app.
DeepSeek R1 is the model that made DeepSeek famous. When it launched in January 2025, it didn't just match OpenAI's o1 on math and coding benchmarks. It did it as an open-source model with a published training recipe, something no one had done before at that level. Here's what R1 is, how DeepThink works, and why it matters.
What R1 Actually Is
R1 is a reasoning model. Unlike standard language models that generate text directly, R1 produces a chain of thought before giving its final answer. You ask a question. The model writes out its step-by-step reasoning. After working through the problem, it delivers the answer.
The core innovation in R1 was the training method. DeepSeek used reinforcement learning (RL) as the primary training mechanism rather than supervised fine-tuning on human-written reasoning examples. The model learned to reason by trying different approaches, getting rewarded when it arrived at correct answers, and adjusting its strategy. This approach, called Group Relative Policy Optimization (GRPO), proved that you don't need thousands of hand-written chain-of-thought examples to teach a model to reason. You just need a clear reward signal.
The R1 research paper was published on arXiv and later appeared as a cover article in Nature in September 2025, a first for a Chinese AI lab.
How DeepThink Works in Practice
In the DeepSeek app, R1 is called DeepThink. When you turn it on and ask a question, the interface splits. A thinking panel appears on the left (or top, on mobile) showing the model's internal reasoning. On the right (or below), the final answer appears once the thinking is complete.
The thinking panel reads like a log of a very focused person working through a problem. You'll see lines like "The user is asking about probability. Let me define the sample space first." Then it lists possibilities. Then "Wait, I need to check if events are independent." Then it recalculates. Then "That doesn't match, let me try a different approach."
For math, coding, and logic puzzles, watching the reasoning is educational. You can see exactly where the model went right or wrong. For everyday questions like "what's the weather," DeepThink is overkill. The model will still produce a chain of thought, but it'll mostly say "This is a straightforward factual query."
R1 vs o1 and o3
The natural comparison is with OpenAI's reasoning models. Here's the honest assessment:
Accuracy: OpenAI's o3 (the latest reasoning model) still leads on the hardest benchmarks. On competition math (AIME) and PhD-level science questions (GPQA), o3 scores a few points higher.
Transparency: R1 wins decisively here. o1 and o3 hide the raw reasoning chain and show only a summary. R1 shows everything. For developers debugging model behavior or researchers studying reasoning, this is a big deal.
Cost: R1 is essentially free through the chat app. OpenAI's reasoning models require at least a Plus subscription ($20/month), and heavy use needs Pro ($200/month). The R1 API pricing is also dramatically cheaper: roughly $0.55/million input tokens and $2.19/million output tokens for R1-level reasoning, versus OpenAI's $15/million input and $60/million output for o1.
Availability: R1 is open-source under the MIT license. You can download the full model weights and run them locally. The o1 and o3 models are closed. You can only access them through OpenAI's API.
DeepSeek R1 API
Developers can access R1-level reasoning through the DeepSeek API. As of mid-2026, the R1 model has been merged into the V4 architecture. You access reasoning mode through the deepseek-v4-pro endpoint with the thinking parameter enabled.
The older deepseek-reasoner endpoint still works until July 2026 but will be deprecated. All new development should target the V4 endpoints.
API pricing (V4-Flash with reasoning):
- Input: $0.28/million tokens
- Output (including reasoning tokens): $1.10/million tokens
- Reasoning tokens are billed at the same rate as output tokens
- A typical reasoning-heavy query might generate 2,000-8,000 thinking tokens before producing the final answer
How to Download R1 DeepSeek
The R1 model weights are available on Hugging Face and ModelScope under the MIT license. Several options:
Full R1 (671B): The original model. Needs serious hardware, multiple A100/H100 GPUs.
R1-Distill versions: DeepSeek released distilled versions trained on R1 outputs. The 32B and 14B versions run on consumer hardware:
- R1-Distill-Qwen-32B: Runs on a single RTX 3090/4090 with 4-bit quantization
- R1-Distill-Qwen-14B: Runs on most modern GPUs with 8GB+ VRAM
- R1-Distill-Llama-8B: Runs on nearly any modern GPU
The distilled versions don't match full R1 on hard benchmarks, but they retain the chain-of-thought reasoning style and perform well for most practical tasks.
The app is the easiest way. Download the DeepSeek app, sign up for free, and enable DeepThink with one tap. No setup, no hardware, no configuration.
What R1 Changed
R1's impact went beyond benchmark scores. It proved three things the industry wasn't sure about:
- First, that reinforcement learning alone could produce strong reasoning. Before R1, the dominant approach was to collect human-written chain-of-thought examples and fine-tune on them. R1 showed you could skip that step and let the model figure out reasoning on its own.
- Second, that open-source could compete with closed-source on reasoning. Before R1, the best reasoning models (o1) were all closed. R1 was the first open-source model to match them, and the research paper gave anyone the recipe to replicate the results.
- Third, that cost efficiency was possible without cutting corners. R1's training cost was a fraction of what OpenAI spent on o1, achieved through architectural innovations rather than just throwing more GPUs at the problem.
The Nature cover and the Nvidia stock drop were side effects. The real story is that R1 reset expectations about what a small team with limited hardware could build.