Writing 11 min read
DeepSeek and the world — why the open, cheap model became my daily driver
From Claude's rate-limit walls to Opus 5's hallucinations and a $1.5 billion book settlement — the month that moved my daily AI work to DeepSeek, with the receipts.
Zain Haroon — Full-stack Engineer
For a long time I was a Claude loyalist. Not the casual kind — the kind who recommends it to everyone, pays for the subscription, and treats the other models as also-rans. It took about a month of watching the ecosystem fall apart around me — rate limits, hallucinations, and one very expensive model that kept apologising instead of working — to admit I'd been wrong about what I was paying for. This is that month, written up properly: what happened, what I tried instead, what everything cost, and why a $5 DeepSeek API top-up is now my daily driver.
The subscription that taught me about limits
My timeline with Claude started on Codex. I used Codex before Claude, and when I moved I didn't look back — then I tried Codex 5.5 again later and its "thinking" still didn't have the reasoning quality Claude could do. So Claude became my reference point: the best model I'd used, right up until GPT 5.6 came along and took that crown.
But the subscription kept finding new ways to be the bottleneck. The $20 plan sounds generous until you actually read the fine print of how the limits work: a five-hour rate-limit window that you're supposed to work inside. Imagine half of that window gone within the first hour, at maximum usage — that stopped being hypothetical for me. I'd burn half my five-hour allowance in under an hour just chatting — not even coding, just using the model to reason through a decision like whether to buy a mouse or a controller. A shopping decision, and it was costing me work budget.
And the ecosystem around it was just as fragile. We watched models get pulled and tools get banned out from under us — Fable was the one everyone noticed — and Claude tightening its rate limits as a regular event, not an emergency one. It became clear that I didn't own my tooling; I was renting it, on increasingly bad terms.
Opus 5 — the expensive model that makes things up
Then Opus 5 launched, and at first it looked like everything I'd asked for. A flagship that was visibly optimised around token usage — fewer tokens, tighter outputs, the promise of more work per dollar. The benchmarks agreed with it; Artificial Analysis has Opus 5 as the leader in agentic knowledge work. I was excited. It didn't last.
Opus 5 was the worst model I have ever used for verification. It hallucinated with total confidence, and the failure pattern was infuriating: ask it to re-do something and it would respond with "Oh, I'm sorry — let me re-create that for you," and produce another confidently wrong version. And again. And again. It never stopped to check its own work — it just apologised and re-fabricated.
The moment that killed it for me: I asked it to create a deliberately defective test file — something broken on purpose, so I could use it to test whether other models could spot and fix the bug. That's a simple, well-scoped ask. Opus 5 hallucinated the file: invented structure, invented the defect, invented the fix. And here's the part that should embarrass a flagship: the open-source model that costs a fraction of Opus 5 did it correctly, first try. When I pointed this out, Opus 5's answer was to say it would rewrite the test file again.
That combination — expensive, confident, wrong — is the worst a model can be. And it's what pushed me to actually go looking for alternatives instead of staying loyal out of habit.
The alternatives tour
Once I started looking, the landscape was crowded. I came across three serious contenders: Kimi K3, Kimi 2.7 Code, and the DeepSeek V4 family — Flash and Pro are both first-party DeepSeek models, and the two tiers turned out to matter more than I expected. I tried them in that order, roughly.
Kimi 2.7 Code is a smaller model, positioned a step below Sonnet. It's genuinely good for most tasks — as long as the task isn't UI. Give it interface work and the result is unmistakably AI slop: that flat, generic, everyone's-crappy-dashboard look that you can spot from across the room.
Kimi K3 fixed exactly that. I gave it the design 2.7 Code had produced, and K3's version actually looked designed. It's the open-weights frontier on Artificial Analysis for a reason. But the upgrade had a cost: the experiment ran through $7 of my OpenRouter balance. Not a huge number, but suddenly I was paying retail for the privilege of comparison shopping.
Then I tried the OpenCode Go plan, which gave me access to a mix of models including GLM 5.2 and Grok. The quality was fine — that wasn't the problem. The problem was that one session of serious use ate 70% of my limits and 40% of my monthly allowance. I went from shopping for models to rationing tokens in the span of a weekend. There was a model I'd earmarked to try — Muse Spark, which scores a point above the new Flash on the same index — and I never even got the chance.
What OpenRouter taught me
OpenRouter is the easiest way to try everything at once, and that's also exactly where it fails. I topped up $15 to keep the tour going and learned three lessons in one evening:
- Provider-level rate limits are severe. The model can be fine while the provider you got routed to is a wall. OpenRouter's limits are their own thing, separate from what the model's original API would give you.
- Some providers are just slow. Same model, two providers, one of them takes four times as long and you can't always tell before you've paid.
- Auto-switching picks the expensive lane. DeepSeek V4 Pro is cheap on DeepSeek's own API. Through OpenRouter, the auto-router would land me on Fireworks — where the same model is meaningfully more expensive than buying direct. The "just works" aggregator was quietly marking up the one thing I was trying to get cheap.
The aggregator is great for a weekend of exploration. It's not where you want to live when you actually care about the bill.
DeepSeek: the daily driver
I'd saved the best for last by accident. DeepSeek has been the real find of this whole exercise — and yes, I know how that reads; the open-source option that beats the paid flagships is a very 2026 sentence. But the numbers are the numbers.
Even DeepSeek's Flash model is the best value I've found in this comparison, full stop. For the price they charge, the thinking it does is genuinely good — it shows its chain of thought, you can watch it consider a wrong answer and correct itself, which means you can audit its work. And its validation is solid: it checks its output instead of apologising and re-creating it. The contrast with Opus 5 could not be sharper.
DeepSeek's update schedule is the other half of the story. A few days ago, the 0731 build of V4 Flash landed, and it was the biggest single-model improvement I've seen this year: a 10-point jump on the Artificial Analysis Intelligence Index, from 40 to 50. For context on that number: Claude Sonnet 5 at max effort scores 53. So DeepSeek's small model now sits three points below a flagship that costs fourteen times as much per input token and thirty-six times as much per output token — and it left its own Pro sibling six points behind in the process. The 0731 update also put Flash within a point of GPT-5.6 Luna and GLM-5.2 at max, and level with Gemini 3.6 Flash. And the part that should make Opus 5's PR team nervous: the measured hallucination rate dropped twelve points, and Artificial Analysis attributes that improvement purely to fewer hallucinations, not better guesses.
Then there's the receipt. I put $5 into the DeepSeek API, and two days in, this was the state of things:
{
"top_up": "$5.00",
"elapsed": "2 days",
"cost_used": "$1.24",
"api_requests": 522,
"tokens": 67,515,910
} $1.24 across 67.5 million tokens — under two cents per million, blended across cached and uncached traffic. The 98% cache-hit discount on DeepSeek's first-party API is doing a lot of that work, and it's a discount Anthropic doesn't come close to matching. Five dollars of DeepSeek is not a top-up; it's a lifestyle. I have genuinely stopped thinking about cost while using it, which is more than I can say for any subscription I've had this year.
Pricing at a glance
Here's the full map of what I used and what each pricing model actually cost me, because "generous subscription, brutal API" turns out to be the industry pattern. Where I have real numbers, I'm using them:
- Claude Sonnet 5 — $20/month subscription. The five-hour rate window drained in under an hour on chat-only reasoning. API: $2.00 per 1M input, $10.00 per 1M output, $0.20 cache hit (90% off). Scores 53 on the Intelligence Index.
- Codex / GPT 5.6 — the standard $20/month tier. GPT 5.6 was genuinely the best model I'd used at the time; GPT-5.6 Luna scores 51. Codex 5.5's thinking still trailed Claude's reasoning quality.
- Opus 5 — the most expensive model in this comparison, per token, on a flagship that advertised token optimisation. That optimisation bought fewer tokens, not better ones: the worst hallucination record of the lot, in my testing.
- Kimi 2.7 Code / Kimi K3 — pay-as-you-go on OpenRouter. 2.7 Code is near-Sonnet for non-UI work; K3 fixed the UI and leads the open-weights field at 57 — but the experiment alone burned ~$7 of balance.
- GLM 5.2 / Grok — bundled quotas on the OpenCode Go plan. GLM-5.2 scores 51 at max; quality was fine. One session consumed 70% of my limits and 40% of the monthly allowance. Quota models are the trap: the meter runs even when you're exploring.
- OpenRouter — $15 top-up, with provider-level rate limits, slow providers, and an auto-router that could land on Fireworks charging more for DeepSeek V4 Pro than the original API does.
- DeepSeek V4 Flash / Pro — both first-party DeepSeek models, direct API. Official pricing, per 1M tokens: Flash 0731 charges $0.14 input (cache miss), $0.28 output, and $0.0028 on cache hits — a 98% discount off input. Pro charges $0.435 input, $0.87 output, $0.003625 cache hit. My actual spend: a $5 top-up, $1.24 used across 67.5M tokens in two days — a blended rate that only makes sense because most of those tokens rode the cache-hit price. (DeepSeek has announced a peak/off-peak policy that will double prices during Beijing peak hours — worth watching if you're on a global schedule.)
The data question
The pricing story is one half of why I left. The other half is trust, and it's the half that doesn't show up in any benchmark. The flagship labs are spending their money on lawsuits — and the biggest one landed two weeks before this post: on July 20, 2026, a US federal judge approved Anthropic's $1.5 billion settlement with book authors over Claude's training data. The suit alleged Anthropic had trained on roughly half a million pirated books — 482,460 titles, at about $3,000 a work. The money is not the anomaly; it's the pattern. Between 2023 and 2024, more than fifty copyright suits were filed against AI companies; the 2026 trackers count over seventy active — the New York Times against OpenAI, Universal Music against Anthropic, Getty against Stability, a Meta complaint alleging millions of pirated books and articles in its own training set.
The legal picture is genuinely messy, and it's worth being precise about it: some courts have found that training on copyrighted material is fair use, and it was the piracy — downloading a pirated library, not the training itself — that drove the Anthropic payout. But that distinction is cold comfort when you're the one whose data is in the pile. And the regulation is catching up: the EU AI Act now fines general-purpose AI providers up to €15 million or 3% of global turnover for failing to even publish what they trained on.
Here's where the open-source question stops being theoretical. DeepSeek trained on the internet too — open weights don't make the training clean, and I'm not pretending they do. Nobody in this comparison is innocent on data.
The difference is what happens after the training. A proprietary model is a black box: you can't inspect the weights, can't audit what went in, can't run it yourself, and your usage data goes into a walled garden that charges you subscriptions, throttles you with rate limits, and sells you API tokens at flagship prices. An open model lets you look — you can audit its weights, self-host it, leave whenever you want. If the industry is going to train on everything anyway, the model that lets you check its work is the one whose behavior I want to reward with my usage. Transparency is the only governance that actually scales.
The verdict
Three patterns came out of this month, and they're all worth stating plainly.
First: the subscription is a foot in the door, not a deal. The companies that feel generous in their $20 plans make it back — and more — in API pricing, where the per-token rates are the worst of anyone. Fourteen times the input price buys you three index points, and thirty-six times the output price buys you one model's worth of apologies. "Cheap subscription, expensive API" is not a pricing strategy, it's an acquisition strategy.
Second: expensive and confident is the worst combination a model can have, and it's exactly what Opus 5 turned out to be. The open-source model that costs a fraction of its price did the task it failed, verified its output, and — when it wasn't sure — said so. That's the entire argument in one sentence.
Third: the data question is not going to get quieter. The billion-dollar settlements and the seventy-odd active lawsuits are the cost of doing business for the closed labs, and that cost is built into the subscriptions and the API rates you're already paying for. An open model doesn't get to claim innocence — but it gets to be examined, which is more than any of the black boxes can say. If they're all going to use the world's data anyway, I'd rather pay the one I can inspect.
DeepSeek is my daily driver now. Not because it's open source, not because of benchmarks — though the 0731 numbers don't hurt — but because after a month of model shopping, it's the one that didn't make me pay for its confidence.