DeepSeek R1 vs V3: Models, Free API, V4 and Limits
I’ll be upfront about something. I haven’t moved any of my own work to DeepSeek. I’ve stuck with the tools I already know. But I’ve followed this company closely since R1 dropped and briefly made Silicon Valley panic, and honestly, it’s hard not to.
The problem is the naming. R1, V3, V3 0324, R1 0528, V4 Flash, V4.1 Flash, V4 Pro. Plus a pile of community spin-offs on top. It’s a mess, and I say that as someone who reads model cards for fun.
So here’s my attempt to untangle it. What each model actually is, how to run DeepSeek for free, what the API costs right now, why you keep hitting “server is busy,” and where V4 stands.
DeepSeek R1 vs V3: What Actually Differs
Here’s the short version. V3 is the general chat model. R1 is the reasoning model.
V3 answers straight away. It’s good at writing summaries and everyday questions. R1 “thinks” first, writing out a long chain of reasoning before it answers. That makes it slower but much better at math, logic and code.
So when people search DeepSeek R1 vs V3, the real question is usually “do I need the thinking or not?” For quick stuff, V3 was the better pick. For hard problems, R1.
Now, the twist. That split doesn’t really exist anymore. DeepSeek merged the two ideas, first with hybrid thinking in V3.1, and now V4 runs both as modes of one model. The old API names, deepseek-chat and deepseek-reasoner, were retired in July 2026. Sure, people still compare DeepSeek V3 vs R1 out of habit. But on DeepSeek’s own API, it’s now one model with a thinking switch.
If you’re weighing DeepSeek against the big names, my roundup of ChatGPT alternatives covers where it fits.
The Model Versions, Decoded
Those numbers on the end are dates. That’s the trick to reading them.
V3 0324 and R1 0528
DeepSeek V3 0324 is the March 2025 refresh of V3. It noticeably improved coding and reasoning without a new name. R1 0528 is the May 2025 update to R1. It hallucinated less and thought for longer on hard problems.
Both still matter, because open weights never disappear. A lot of free hosting and local setups still run these two.
MoE and Sparse Attention
DeepSeek MoE is the architecture behind almost everything they make. MoE stands for mixture of experts. The model is huge, but only a small slice of it switches on for each word. V3 has 671 billion parameters, yet only about 37 billion work at a time. That’s why it runs so cheaply for its size.
DeepSeek sparse attention came later, in V3.2. In plain terms, the model learns to skip parts of a long document that don’t matter for the current word. Long chats get cheaper as a result. V4 builds on the same idea.
The Paper Everyone Cites
In September 2025, R1 became the first major language model to pass peer review in Nature. The paper’s title is literally “DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning.” It also put R1’s reasoning training cost at about $294,000. That figure surprised a lot of people. Me included.
How to Run DeepSeek for Free
This is the part most people actually want. There are two routes: someone else’s servers or your own machine.
OpenRouter
The OpenRouter DeepSeek API is the easiest free option I know of. You sign up, grab a key, and pick a model ending in “:free.” The one everybody uses is deepseek/deepseek-r1-0528:free. It’s still listed as free, with a context window of about 164,000 tokens.
The catch is limits. Free models get around 50 requests a day, or about 1,000 if you’ve ever bought $10 in credits. Free listings also change. Older IDs like deepseek/deepseek-chat-v3-0324:free have come and gone, so check the model list before you build anything around one.
Running It Locally
If you’d rather keep everything on your own computer, the smaller “distilled” versions are the way in. These are compact models trained to copy R1’s reasoning style.
In Ollama, deepseek-r1:8b runs on a decent gaming PC. The deepseek-r1:14b tag needs more memory but reasons better. In LM Studio, the command lms get deepseek/deepseek-r1-0528-qwen3-8b downloads the newer 8B distill.
People with serious hardware go bigger. On Hugging Face, DeepSeek V3 0324 GGUF files let you run compressed versions of the full model. Teams serving lots of users tend to run DeepSeek R1 0528 vLLM setups instead, since vLLM handles many requests at once far better.
To be fair, a distilled 8B model is not “real” R1. It’s a small model doing an impression. Useful, sometimes surprisingly good, but not the same thing.
Chimeras, OCR, and Other Spin-Offs
Open weights mean anyone can remix DeepSeek’s models. Some of the remixes got popular on their own.
R1T Chimera
R1T Chimera comes from TNG Technology Consulting, a German firm. They stitched together pieces of R1, R1 0528, and V3 0324 without any extra training. The result reasons nearly as well as R1 0528 but writes far shorter answers, so it feels faster.
You’ll also see people call it the DeepSeek Chimera. Same family. It’s MIT licensed, and it’s not recommended for tool use.
DeepSeek-OCR 2
DeepSeek-OCR 2 is a small model that reads documents. It arrived in January 2026. Its trick is reading a page in a sensible order, like a person would, rather than scanning strictly left to right.
If you’re searching DeepSeek OCR 2 for scanning receipts or PDFs, it’s genuinely good at tables and messy layouts. It reportedly scored about 91% on a major document benchmark.
Coder V2 Lite and OpenClaw
DeepSeek-Coder-V2-Lite is an older coding model from 2024. It’s small enough to run locally, and some people still use it for autocomplete.
OpenClaw DeepSeek setups are newer. OpenClaw is a popular open-source agent framework, and it made V4 Flash its default model the week V4 launched. That says a lot about cost, honestly.
DeepSeek API Price
DeepSeek’s whole pitch has always been price. It’s still cheap, though not quite the bargain it was.
Here’s the DeepSeek API price as of late September 2026, per million tokens at peak hours:
- V4.1 Flash: about $0.30 for input and $1.20 for output. This is the default, everyday model.
- V4 Pro: about $1.32 for input and $3.96 for output. This is the big one, for harder work.
- Cached input: a tiny fraction of those rates, for text the model has already seen.
The newer twist is peak pricing. Since August, off-peak hours cost half. Peak hours fall in the UTC morning on weekdays, which lands in the Indian morning and early afternoon. So timing a big batch job actually saves money.
My mild gripe? The pricing has changed several times this year. Check DeepSeek’s own pricing page before you budget anything. Don’t trust a number in an article from three months ago. Including this one, eventually.
DeepSeek Limits and Common Errors
Does DeepSeek have a limit? Yes, a few, and they show up as error messages that aren’t very helpful.
“The server is busy. Please try again later.”
This is the famous one. It means DeepSeek’s servers are overloaded, not that you did anything wrong. It got really bad after R1 went viral in early 2025, and it still pops up during busy hours.
The fixes are boring. Wait a few minutes. Try off-peak hours. Or use the same model through OpenRouter or another host, which skips DeepSeek’s own servers entirely.
“Length limit reached. Please start a new chat.”
The full message reads, “DeepSeek length limit reached. Please start a new chat.” Every conversation has a maximum size. Once you hit it, that chat is done.
My workaround is simple. Ask DeepSeek to summarize the conversation first, then paste that summary into a new chat. You lose some detail, sure. But you keep the thread going.
Is DeepSeek Down?
Is Deep Seek down, or is it you? DeepSeek runs its own status page, so check that first. If the site works but the app doesn’t, it’s probably the app. If both fail, wait it out.
DeepSeek V4: Release Date and Paper
V4 took longer than anyone expected. Rumors started in late 2025, and the date kept slipping.
The actual DeepSeek V4 release date came in stages:
- April 24, 2026: V4 Pro and V4 Flash arrived as previews, with a 1-million-token context window.
- July 31, 2026: the official V4 Flash landed, with better agent skills.
- Mid-August 2026: V4 Pro left preview.
- September 2026: V4.1 Flash replaced V4 Flash on the API.
V4 Pro is massive, at 1.6 trillion parameters with about 49 billion active. V4 Flash is 284 billion, with about 13 billion active. Both are open weight.
The DeepSeek V4 paper explains the main idea. It’s about compressing long text and skipping what doesn’t matter, so a million tokens doesn’t cost a fortune. Same spirit as sparse attention, pushed further.
Independent testers rank V4 Pro near the top of open models, though a bit below DeepSeek’s own claims. That’s normal. Every lab grades its own homework generously.
Where DeepSeek Goes From Here
I expect DeepSeek to keep doing what it does best: releasing strong models at prices that make everyone else look expensive. A V4.1 Pro already has a name. No date yet, of course.
Will I switch? Maybe not for daily work. But for anyone building on a budget, ignoring DeepSeek is just leaving money on the table. Keep an eye on the rest of the AI startups worth watching too, because this race is far from settled.