If you run an agency, the deepseek harness vs claude code question is really a margin question.
One of these agents built a full site in 11 minutes for about 5 cents.
The other took over 30 minutes and produced the build you would actually put in front of a client.
We ran both on the same brief on the same afternoon, and the result changes what you should be paying for.
The agency verdict
Deliverables run on Claude Code.
Everything upstream of the deliverable runs on DeepSeek Harness.
That single split is where the margin is, and the rest of this article is the evidence for it.
The brief we used
One prompt, no context, no follow-ups β deliberately like a cold client brief.
A 3D animated accountancy website, plus a simple game.
DeepSeek Harness ran DeepSeek V4 Pro. Claude Code ran Claude Opus 5 on high.
Frontier model on both sides, so this is a fair read on what each stack does out of the box.
What happened on the clock
DeepSeek Harness finished in 11 minutes.
Claude Code was past 30 minutes and still going.
For a production team, that is the difference between five iterations in an hour and one.
If your bottleneck is throughput rather than quality, that alone is worth an afternoon of setup.
What happened on the page
The Claude build had animation that followed the mouse, so it felt like a designed site rather than a decorated template.
It also referenced details nobody typed, including Northwest England, because it carried context from earlier work.
For an agency, that is the whole value of a tool that has been living inside your projects.
The DeepSeek build came out animated and cartoony, and it bolted a cartoon game onto an accountancy brief.
That is a tone failure, and tone failures are the ones clients notice first.
Neither build was ready to send. Both needed more context and more rounds.
Claude simply started closer to the finish line.
The numbers
| Metric | DeepSeek Harness | Claude Code |
|---|---|---|
| Time to finished build | 11 minutes | 30+ minutes, still running |
| Tokens used | 483,000 | 48,000 at 20 minutes |
| Relative price per token | ~57x cheaper | Baseline |
| Cost of this run | About 5 cents | Materially higher |
| Used unprompted client context | No | Yes |
| Score | 7 / 10 | 9 / 10 |
| Stage | v0.1 developer preview | Mature |
Old way vs new way for agencies
| Old way | New way |
|---|---|
| One premium agent does every job on the account. | Cheap engine handles drafts and scaffolds, premium engine handles deliverables. |
| Token spend scales linearly with client count. | Most of the volume moves to an engine costing pennies per run. |
| Every new tool means retraining the team. | An orchestrator holds the process; the engine underneath is a config change. |
| Vendor pricing sets your gross margin. | You keep an exit option, so pricing changes are survivable. |
The token burn an agency should model
DeepSeek burned 483,000 tokens on a single one-page site.
Claude was at 48,000 at the 20-minute mark on the same brief.
Ten times the tokens for the less finished result.
The reason it still wins on cost is price: roughly a 57th per token, so the run landed around 5 cents.
Model that across a production team and the saving is real, but only for work where a rougher first pass is acceptable.
Also model the risk. This is early-stage pricing on a v0.1 preview, and verbose plus normally-priced is a very different proposition.
π₯ Want this mapped to your agency?
Book a free AI SEO strategy session and we will look at where your delivery hours actually go: go.juliangoldie.com/strategy-session. More on how we work: goldie.agency.
Why it hit 105,000 stars in two days
DeepSeek Harness collected 105,000 GitHub stars in roughly two days, one of the fastest open source launches anyone has tracked.
Not because it out-built Claude. It did not.
Because it is free, open and model-agnostic β it runs on your machine and you choose the brain inside it.
You can plug a free model in, and OpenCode works as a free brain, which takes the running cost to zero.
For an agency, that is an exit option from vendor pricing, and exit options are what keep margins predictable.
How to roll it out without disrupting delivery
Install it through the agent you already use. I did not set it up by hand β Claude installed, tested and wired it in.
Start on internal work only. Scaffolds, first drafts, repetitive variations, internal tooling.
Write the routing rule down. Client-facing goes premium, internal goes cheap. Make it a rule, not a judgement call per task.
Measure on a real job. Toy prompts hide verbosity; token burn only shows up at production scale.
Keep the premium engine wired in. Our test showed precisely where the cheap one falls short, and it falls short where clients look first.
The v0.1 reality check
This is a developer preview. Things will change fast and things will break.
That is a reason to keep it away from deliverables this quarter, not a reason to ignore it.
I scored it a 7 out of 10 on trajectory; Kasra Dash scored it lower on what is on screen today, and he scored Claude Code a 9.
Both readings are fair.
What matters for an agency is that a credible free alternative now exists, and that alone changes your negotiating position.
Where the margin actually comes from
Most agencies think about AI cost as a subscription line.
The useful way to think about it is cost per attempt.
At around 5 cents and 11 minutes, an attempt on the cheap engine is close to free in both money and time.
That means your team can fail four more times per hour on internal work without anyone noticing on the P&L.
The premium engine is not competing with that, and it should not be asked to.
It is competing on the last mile β the version that carries your name to the client.
Split the workflow on that line and you get most of the saving with none of the reputational risk.
Blur the line and you will eventually send a client something with a cartoon bolted onto it.
What the plugin design means for your stack
The harness is open source and every part of it is swappable, including the model.
For an agency that is not a technical curiosity, it is procurement leverage.
You can run DeepSeek V4 Pro today, plug a free brain in tomorrow, and swap again when something better lands β without rebuilding your process.
OpenCode works as a free brain inside it, which is the zero-cost entry point if you want to trial it before committing budget.
That is why 105,000 people starred the repository in about two days.
They are not voting on build quality. They are voting on not being locked in.
What we would not do yet
We would not put a v0.1 developer preview anywhere near a client deadline.
We would not assume today’s pricing is permanent, because a very verbose model at normal prices is a different product entirely.
And we would not migrate a team’s habits to a tool that is two days old.
What we would do β and have done β is wire it in behind an orchestrator, point internal work at it, and measure for a month.
That gets you the upside with a reversible decision, which is the only kind of decision worth making about tools this new.
Frequently asked questions
Which agent should an agency use?
Claude Code for deliverables, DeepSeek Harness for internal volume work.
How much can we save?
Roughly 57 times cheaper per token, offset by about ten times more tokens used. Our build cost around 5 cents.
Is it stable enough for client work?
Not yet β it is a v0.1 developer preview. Keep it upstream of the deliverable.
How long did each take?
Eleven minutes versus over 30 on the identical prompt.
Do we have to choose?
No. Route per job with an orchestrator and use both.
About Julian
I’m Julian Goldie, founder of Goldie Agency, a 7-figure SEO and link building agency with a 70+ team.
I have 400K+ YouTube subscribers, 163K X followers and 29K+ Udemy students, and I wrote Link Building Mastery.
We test these tools on live client delivery before we recommend them, which is why this piece has numbers rather than opinions.
Also on our network
The same test from other angles: juliangoldie.com, juliangoldie.co.uk and goldstarlinks.com.
πΊ Video notes + links to the tools π
π₯ Learn how I make these videos π
π Get a FREE AI Course + Community + 1,000 AI Agents π
For agencies the deepseek harness vs claude code decision is not either-or β it is which engine touches the deliverable.