AI News Today
← All episodes
Episode 113 · August 12, 2026 · 12:28

Pi Agent + DeepSeek AI Just Beat Claude Code?

PI Agent & DeepSeek Beat Claude Code: The Harness Multiplier

Discover why the setup around your AI model determines over 50% of its success and how PI Agent outperformed Claude Code in recent benchmarks. Learn how the "Harness Multiplier" allows simpler setups to beat complex ones and how to benchmark your own AI pairings for maximum reliability.

Full transcript

Pi Agent plus DeepSeek just beat Claude Code and Hermes Agent and it proves something huge. Imagine making your AI way more powerful without touching the AI at all. That's exactly what just happened in a public test. So the same AI, same tasks, but one setup passed 20 out of 30 whilst the others only passed 15.

That means the setup you pick is quietly deciding half your results and most people don't even know it. In this video I'll show you the winning setup, why the simple option beat the heavy one and the one mistake that's holding almost everyone back. Get this right and your AI becomes more reliable, faster and works harder for you starting today. Let's get into it.

Composio, they built tooling for AI agents and they just ran a big public test. They took one model, DeepSeek V4 Flash and pushed it through eight different agent harnesses on 30 hard agentic tasks. Tasks where the agent has to go to do things, use tools and finish the job on its own. Quick thing before you get the numbers because this word matters, a harness.

Simple way to think about it, the model is the engine, the harness is a car, you drop that engine into. So Claude Code is a car, Hermes is a car, PyAgent is a car, same engine, different car, completely different drive. The harness decides how the model plans, how it uses tools, how it handles mistakes and when it stops. That's a whole test, one agent, eight cars, which car wins.

And let me run you through the numbers PyAgent passed 20 out of 30 tasks. DeepAgents, the harness built by LongChain, sorry, LongChain passed 16 out of 30. Hermes agent passed 15 out of 30 and PrimeAgent passed 15 out of its 24 valid runs. So six of its runs actually got thrown out because the grader literally couldn't score them.

We'll come back to why in a second because the reason is a trap a lot of people fall into. Speed told a similar story, PyAgent finished tasks in about 132 seconds median, Hermes took about 175 seconds, DeepAgents about 187 and PrimeAgent about 242. So the harness that passed the most tasks was also the fastest of these four and the lightest to run. Worth being fair here, in the overall rankings across all eight harnesses, Claude Code and OpenCode were actually quicker than Py.

So Py didn't win everything, but on pass, rate and efficiency together it came out on top. Now zoom out because this is the part that actually matters. Across all eight harnesses the exact same model scored anywhere from 47% to 67% task success depending on the harness. Same model, same 30 tasks, a 20 point swing in how often it succeeds.

So the AI never got smarter and it never got worse, only the wrapper changed. I call this the harness multiplier. The tool you wrap around your AI multiplies the result up or down. If you pick the right harness and the same model becomes more reliable and faster at the same time.

Pick the wrong one and you get more failures and slower runs with the exact same intelligence underneath. Most people have never even thought about this. They argue about which model's best, GPT versus Claude versus DeepSea versus Gemini, and meanwhile the harness is quietly deciding half the outcome. Composio's own conclusion said it plainly.

Benchmark the model harness pair you actually use, not the model in isolation, and that one sentence changes how you set up everything. And this is exactly why we build the Agent OS the way we did inside the AI platform. The Agent OS lets you plug in all your favorite agent harnesses, Claude Code, Hermes, OpenCore, into one system with shared memory so you can run the same task through different setups and see which pairing actually works for your business. That's the harness multiplier put to work.

You get the full zip file, a 30-day roadmap to set up, video tutorials, and daily updates as we keep improving, plus four weekly coaching calls where you ask questions about your own agent setup live, which model, which harness for your exact tasks. There's over 3,700 business owners inside there and plenty of them are running Hermes and testing these exact pairings right now. Link in the comments description or just go to theairprofitborm.com to get access. So back to the test and there's a detail in here that explains why the light option won.

And it goes against everything most people assume about AI. So Pi ran with almost nothing added. A fresh vanilla install. The only thing bolted on was the MCP server plugin and that's the standard connector that lets an agent plug into outside tools like a universal adapter.

So there was no custom setup, no tuning, no special configuration, straight out of the box and it passed most tasks. Now if you compare that to Prime Agent, Prime ran the heaviest sessions of all eight harnesses. Some sessions hit 3.5 million tokens and 33 tool calls. Now picture an agent writing itself a to-do list the length of a phone book before it even starts a job.

Sessions so big that the grading system timed out just trying to score them. Two runs couldn't be graded because of that. Four more never recorded at all. Six runs gone in total and the runs that did count are still only matched by Hermes on passes whilst taking nearly twice as long as Pi.

So the pattern is right there in the data. The heavy to-do everything setup choked on its own weight. The light simple setup passed the most tasks in the least time. More layers did not mean better results in this test.

It meant the opposite. That flips the old way of thinking on its head. The old way was grab the biggest model, stack on every plugin, every extension, every fancy layer and assume more equals better. The new way equals something new and this test proves it.

You just pick a fast model, put it in a clean light harness and test a pair on your actual tasks. DeepSea V4 Flash is a flash model of course. It was built for speed, not for winning intelligence contests and in the right harness it passed two out of every three agentic tasks. The setup carried it further than raw brain power ever could on its own and there's a simple reason that light setups keep winning these kind of tests.

Every extra layer you add is another place for the agent to get lost. Every extra tool is another decision it has to make. Every giant instruction file is more noise it has to read through before it acts. So a clean harness gives the model a short path from here's a task to task done.

A bloated one sends it wandering and that's why the vanilla install beat the heavy weight. It's a shorter path with fewer wrong turns. Now let me hit the three beliefs that stop people from acting on this because I see them constantly. Belief number one is I need the most powerful top tier models to get good results.

This test just showed the opposite. So the model was fixed, it never changed and results still swung from 47% success to 67% success based purely on the harness. So if you've been waiting to get access to some flagship model before starting with agents you've been waiting for the wrong thing. A lightweight model in the right harness passed more tasks here than most people would expect from any model.

The setup mattered more than the raw model itself. Belief number two is this is a technical benchmark. It doesn't apply to me, I'm not a coder, but if you flip that around if you're not technical this applies to you more than anyone. A developer can dig into a broken setup and fix it, you can't.

So your choice of harness is the result and here's the good news. Buried in this test the winning setup was the vanilla one. Fresh install, one standard plugin, nothing custom. The thing that won is the thing a non-technical person would end up with anyway.

So you don't need to out-engineer anyone. You need to pick well and keep it simple. Simple won on the scoreboard. And belief number three, the AI agents are unreliable so I'll just wait until they get better.

But if you look at the spread again 47% to 67% same model, same day, same tasks. So when someone says agents don't work, the honest question back is which pairing did you try? Which harness did you try? Because some pairings clearly work far more reliably than others right now.

The people getting results today aren't waiting for better AI, they're testing pairings and keeping the ones that pass. Waiting doesn't improve your setup, testing does. And every month you wait the people who test pull further ahead with the same tools you already have access to. There's one more lesson highlighted in this benchmark that changes how you should read every AI headline from now on.

When you see a score like for example this model passed x percent of tasks, that number is incomplete without knowing the harness it ran. This exact test proved the same model can swing 20 points either way based on the wrapper. So a leaderboard that only names a model is telling you only half the story. From now on when a new model drops and everyone's sharing benchmark charts, your first question should be what harness?

Because that answer matters more than the model name at the top of the chart. So what should you actually do with this? Simple, build your own mini version of this test. Composio use 30 tasks, you need three.

So step one, write down three real tasks from your actual week. Things you already do by hand, maybe it's sorting new leads into hot and cold, maybe it's writing the first draft of your weekly email, maybe it's pulling details from a spreadsheet that you have and just organizing it properly. The benchmark only matters as well if it's related to your work because that's the work you're trying to automate. Step number two, run those same three tasks through three different setups.

Two harnesses, same model, same instructions, word for word. Don't change anything else. If you change two things at once, you learn nothing. This is the exact discipline that Composio used.

So one variable, everything else locked. And step number three, score them the same way you did and the way Composio did. So three columns. Did it finish your job, pass or fail?

How long did it take? How much cleanup did you have to do afterwards? The last one is the honest column that most people skip. An agent that finishes beneath like 20 minutes of fixing didn't really pass.

And step number four, keep the winner, retest when new versions ship. Because the rankings in this space don't sit still. A harness update can flip these results next month. Let's say for example Hermes releases a new update, that could make it better.

So does Pi, so does Cloud Code. The people win aren't the ones who picked right once, they're the ones with a quick way to recheck. Do that and you'll know something about your own business that most companies genuinely do not know. Which exact pairing performs best based on your work?

One more honest note, because I don't want to oversell Pi here. This test was one model. DeepSeek V4 Flash on one set of 30 tasks. Composio was clear that different harnesses suit different models.

Pi won for this model. That does not mean Pi wins for every single model. Cloud Code was quicker overall. Hermes has its own strengths and it's built for a much wider range of agentic work than pure task running.

The lesson isn't everyone should switch to Pi. The lesson is the pairing is the product, so test yours. And that's really where this leaves us. The gap between businesses right now isn't access to AI.

Everyone has access. Everyone can run the same models. This test literally used one model for everything and the gap is that a small group of people test the setups and most people don't. One group knows the numbers, pass rate, time per task, cleanup needed.

The other group is guessing and usually getting half the reliability they could be getting without knowing it. And if you want to be in the group that knows, that's exactly what we do inside the AI Profit Border. You get daily tutorials with step-by-step videos including how to set up agent harnesses like Hermes and run models like DeepSeek through them, plus how to swap models inside the AgentOS so you can test pairings on your own business tasks just like the benchmark did. The AgentOS already integrates multiple models, so trying a new pairing takes minutes not days.

And you get four weekly coaching calls where you can bring your exact setup, your model, your harness, your tasks and get it reviewed live. Plus you get a 30-day roadmap so you're not guessing what to do next and a prompt library for the tasks you actually automate. So for example, lead follow-up, content drafts, customer replies and a member map so you can find business owners near you who are running these exact tools and compare notes. 3,700 plus members inside there, support around the clock because there's always someone online, link in the comments description or just go to the AIProfitBorder.com because here's the truth this benchmark widely proves.

The AI didn't get smarter between the best result and the worst result. The human choices around it change. Which harness, which setup, whether anyone bothered to measure. Same model and one setup succeeded two thirds of the time whilst another failed more than half the time and ran twice as slow.

The intelligence was never the bottleneck, the setup was and the setup is about your control. See you in the next one.

More episodes

Browse all episodes →