AI News Today
← All episodes
Episode 1 · March 23, 2026 · 19:52

AI Agents Fail 97.5% of Real Jobs, Here's the Fix…

Why 97.5% of AI Agents Fail (and How to Join the 5% Who Succeed)A recent study reveals that AI agents fail 97.5% of real-world freelance tasks, but this failure masks a massive opportunity for businesses. Learn how to bridge the 'last mile' gap and build a 'relay race' system that combines human judgment with AI speed for maximum profit.00:00 - The 97.5% Failure Rate02:10 - The Last Mile Problem05:45 - Reasoning vs. Delivering08:09 - Narrow vs. Complex Tasks10:22 - 3 Technical Reasons Agents Fail13:32 - The Relay Race Strategy

Full transcript

AI agents fail 97.5% of real jobs. The best AI agents in the world, the ones built by the smartest teams with the most money, the best technology, the most hype, they fail 97.5% of the time when you give them real work. Not test work, not demo work, not show this to investors work, real work, the kind people get paid for. The kind you need done in your business, 97.5%.

That number actually comes from a real study. It's called the Remote Labor Index. It was built by the Center for AI and Scale AI, and they took 240 real projects from a work, the freelance platform where clients pay to get work done. And things like building a video game, for example, creating architectural drawings, making 3D models, designing products, editing audio, writing code, not simple projects.

And the average project took a human about 29 hours to complete. The median value was $200, and all 240 projects together represented over 6,000 hours of work, worth more than $140,000. And then they gave all of that work to the best AI agents on the planet. They actually tested out six models.

So they tested Manus, Grok, 4, Claude, Sonic 4.5, GPT-5, ChatGPT agent, and Gemini 2.5 Pro. Now the highest performing agent, Manus, achieved an automation rate of 2.5%. All other models actually performed worse, ranging from 0.8% to 2.1%. That means 97.5% failure rates.

The best AI agent on Earth, given real paid work, failed at it 97.5% of the time. Now, I know what you're thinking. You've been watching AI do incredible things. You've seen it write emails.

You've seen it build websites in minutes. You've seen all those YouTube videos, maybe even mine, talking about how AI is gonna change everything. Now I'm standing here telling you it fails 97.5% of the time. How do both of those things seem to be true?

That's exactly what we're gonna talk about today. Because the answer to that question is the most important question you can understand right now, if you wanna make money with AI. Not someday, right now, in 2026. Because here's the thing, the failure rate is real, and the opportunity is also real.

And understanding the gap between those two things, that gap is where all the money is. So let's get into it. First, let's understand what the actual study measured. What this actually meant was that the way the test worked was researchers took 358 freelancers with verified Upwork accounts, professionals who had, on average, completed 89 jobs each, and earned $23,364 on the platform.

They paid those freelancers about $15 to $200 to hand over examples of completed real work that had already been paid for by real clients. So every single project in this test was work that a real human had already done, a real client had already paid for, a real professional had already delivered. And that completed human work became the gold standard, the passing grade. Then they gave the same project brief with the same files, the same instructions to an AI agent, and they asked a simple question.

Would a reasonable client accept this AI's work as good enough? And the projects required agents to understand and produce multi-file deliverables, including dozens of unique file types, documents, audio, 3D models, CAD files, and more. This is the key word, complex. Not write a paragraph about can cats eat biscuits, not summarize this article, not give me five ideas for a marketing campaign, like actual multi-step deliverables.

And here where it gets interesting because the AI didn't usually fail because it didn't say, I don't know how to do this. It actually failed in much more human ways. In the analysis of failed projects, there were failures around three main problems. Number one was quality.

About 45% of failures where agents produced amateur or low-grade work was found, right? So that was 45% of the failure rate. Then about 36% of failures came down to submissions where there were truncated videos or missing files or partially completed work. And then 15% of failures were found because of inconsistencies.

Now think about that for a second. The AI wasn't failing because it had no idea what it was doing. It was failing because it would get most of the way there and then not finish a file. It would produce something that looked like the right thing, but was done at the quality level of a first-year student when the client needed a professional.

It was failing at the last mile. And that phase, that phrase, last mile, is the whole story here. AI is unbelievably good at the first 90% of the tasks, the brainstorming, the research, the first draft, the rough version, the skeleton. But that last 10%, the polish, the file format, does this actually meet what the client asked for in the brief, the quality bar, the paying customer expects, that's where it falls down.

And in professional work, that last 10% is everything, right? You can have a beautiful building plan that's 90% done, but if the foundation drawings are missing, nobody can build the building, right? You can have a video that's almost finished, but if the last three minutes are cut off, the client can't use it. And the last mile is what separates a draft from a deliverable.

Now, here's where most people stop reading. They see AI fails 97.5% and they think, okay, AI's overheard, I knew, it's not ready, I'll wait. I want to talk to you directly if you're that person for a second, right? Because that conclusion is exactly wrong and it's gonna cost you.

Here's what the data actually says if you read it carefully. So for example, Claude Opus 4.5 hit 80% on SWE Bench Verified and that's up from 33% a year ago, right? And you've got, for example, when it comes to PhD level scientific reason, Claude Opus 4.6 scored 91%, which exceeds human experts at 69.7.7% by over 21 points. So let me put those two things by side.

You've got AI failing 97.5% of complex paid freelance work and AI is beating PhD level humans by 21 points on scientific reasons. Now, both of these things are true at the same time. How? Because these tests are measuring different things.

The RLI benchmark, the freelance work test, is measuring whether AI can fully replace a human professional doing multi-hour, multi-file, client-ready creative and technical work entirely on its own with zero human involvement. The other benchmarks are measuring whether AI knows things, can reason, can solve problems. There's a massive difference between knowing the answer and delivering the finished product. Think about a brilliant student who knows every answer in the head, but can't write a clean essay.

They'll ace the multiple choice, but they'll bomb the essay. And AI right now is that student, right? The logic is there. The reason is that the raw material is extraordinary, but the end-to-end autonomous delivery, getting from here's a brief to here's a finished product, a client would pay for without a human in the loop, that's still being built.

And here's what you need to understand. Being built doesn't mean not useful. In early 2024, frontier models could sustain autonomous work for about five minutes. By February 2026, Claude Opus 4.6 crossed a full-day work at 14.5 hours, doubling every 123 days, right?

So that is not a trend line and it's increasing exponentially. At that rate, and I'm not saying this to scare you, I'm saying this to tell you because the maths is real, autonomous tasks will arrive by late 2026. Autonomous tasks will arrive by mid-2027. The AI that fails 97.5% of jobs today is getting faster, better, longer, more capable on a curve that doesn't look like a straight line.

It looks like a ramp that suddenly becomes a wall. We're at the bottom of that wall right now. So the people who are waiting and telling themselves, oh, I'll learn AI when it's ready, are doing the equivalent of someone in 2007 saying, I'll learn about the iPhone when apps are actually good. By the time you decide it's ready, the people who started learning now will have a two-year headstart.

And in AI, two years isn't just an advantage, it's a different category of business entirely. So that's basically how it works. And the other thing is as well, the, you know, when it comes to this, right, the study says AI agents fail 97% of the time when they try to fully replace a human professional working alone. But nobody told you that was the only way to use AI.

That's like saying cars are useless because they can't fly. Nobody said cars are supposed to fly, right? Take Air India's virtual assistant as one example, automatically handles 97% of over 4 million customer queries, saving millions in support costs. Not a flashy demo, but a high impact, quietly revolutionary system.

97% success rate for Air India, 2.5% success rate on complex freelance work. Same technology, completely different use case. The difference is that Air India built the agent for a narrow, well-defined repetitive task. You know, for example, customize a question about a flight and you can just answer it with AI.

Same type of question thousands of times a day with clear success criteria. Whereas the freelance work benchmark was the opposite. It's different projects every time, it's different files, different client expectations, different formats, different creative standards. So when you pick the right job for AI, it's actually extraordinary.

But when you pick the wrong job, it falls apart. And right now at this moment, most businesses are picking the wrong job. They're watching demos of AI doing incredible things and thinking, I'll just give a AI agent a huge complicated task and it will do it from start to finish. And then they're shocked when it doesn't work.

And then they come to exactly the wrong conclusion. AI doesn't work. Instead of the right one, which is, I gave it the wrong job. Now, also what's interesting here is that research shows leadership decisions determine outcomes too here.

So for example, 73% fail projects, lack clear executive alignment on success metrics. 68% under invest in data governance foundation. 61% treat AI as IT projects rather than business transformation. And 56% lose active C-suite sponsorship within six months.

Why? Because they're treating AI like a magic box. For example, point out a problem, press a button, get a solution. And that's not how this works.

And this is the most important part. There's a small group of businesses that figured it out. So for example, MIT's Project Nanda report found that only 5% of integrated pilots generate millions in profit, whilst the rest never reach profit and loss impact. So 5% of companies are making millions and 95% are making money.

The 5% aren't smarter than everyone else. They're not bigger. They're not better funded. They just learned how AI actually works, not how the demos make it look like it works.

And there's nothing stopping you from being in that 5%. Let's talk about the three reasons AI agents actually fail. Because if you understand the real reason, you can avoid it. The first reason is what the experts call dumb rag, right?

And I know that sounds like a weird phrase. Let me explain what it means in plain English. Rag stands for retrieval augmented generation. It basically means you give AI a bunch of information to search through, and then it uses that information to answer questions or complete tasks.

The way most businesses do it is they take all their company documents, their manuals, their reports, their emails, their spreadsheets, dump all of it into the AI all at once. It's like pouring a whole library into a box and then they ask the AI questions. So for example, the AI searches through the pile and tries to find the right answer. The problem with that is that agents produce truncated context, misidentified documents, and logic that breaks when edge cases appear.

So bad memory management means that AI is working with garbage in and garbage out. It's like asking someone to find a specific book in a room where all the books have been thrown on the floor in a pile. The information is there, but the organization is so bad, they can't find it reliable. Now, the fix is simple.

Don't dump everything in. Be selective. Give the AI exactly the information it needs for the specific task and nothing else. Narrow the job, narrow the information, get better results.

The second reason AI agents fail is bad connections. In the enterprise, you don't control Salesforce's API, right? You definitely don't control your customer's 5,000 custom fields and undocumented workflows. It's like giving a new hire the server room keys without documentation.

Something will break. Most businesses have dozens of software tools, their CRM, their email, their project management, their accounting software, their customer database. So when AI tries to work across all of these tools, pulling data from here, pushing data over there and updates over here, it runs into a maze of broken connections, missing permissions, the software that doesn't talk to other software. And every time one connection breaks, the whole chain breaks, right?

What is a fix? Well, start with one system, one connection, one workflow, prove it works there, then add another. Don't try to connect everything at once. So we've talked about the first two things, limited context, and also just being selective with which apps you connect.

The third reason, this is a big one, is that most people build demos instead of systems, right? So most AI agent initiatives, whenever designed to scale. Technical teams stand up demos using frameworks that are quick to start, but impressive to watch and easy to showcase. The problem is they fall apart when real world requirements show up.

So for example, security reviews, compliance checks, identity management, audit trails, integration with enterprise systems and long running exception heavy workflows. A demo is not a system, my friends. A demo is a magic trick. It works perfectly on stage in controlled conditions.

We've all seen that with OpenAI and all these other tools that come out, right? It looks amazing on stage, and then when you try it, you're like, not as good as I expected, right? A system has to work on Tuesday morning when the data is weird, the connection drops, and someone uploaded a file in the wrong format. That gap between demo and system is where almost all the money is being lost.

And crossing that gap is a skill, it's learnable, it's teachable, but you can't cross it by watching demos. Now, let me tell you what actually works because the 5% who are making money with AI automation right now, they're not doing anything magical. They just figured out a simple but powerful principle. Don't try to replace a whole human.

Automate one step at a time. Think about your business. Think about the work you or your team does every single day. Most of it isn't complicated, creative, judgment-heavy work.

Most of it is just answering the same questions over and over, sorting information, following up and that sort of thing. And the work, the repetitive rule-based predictable work, or AI is extraordinary at it right now. Not 2.5% success rate, closer to 90% or 95% of the right setup. 83% of business leaders actually expect AI agents to outperform humans in repetitive rule-based tasks.

And they're right, that's not hype. That's where the tool actually works. So let me give you a practical example of this. Uber Freight used AI to cut empty miles by 10 to 15%, move $20 billion in freight, and reduce support wait times from five minutes to 30 seconds.

And that's not AI replacing a human, that's AI doing the first 80% of the work so fast that the human only needs to handle the tricky exceptions. Intelligence-infused demands forecast cut lead times by 22% and reduce expedited shipments by 27%, boosting supplier level accuracy by 35 to 42%. And these aren't hypothetical. These are from Uber Freight themselves.

These are real numbers, right? So you can automate the pieces that are repetitive and keep humans in the loop for the pieces that require judgment. Human plus AI, not AI instead of human. That's a mental model that I would walk away with.

And I would also think of this as a relay race because in a relay race, you don't have one runner that does the whole thing. You break it into parts. Each runner does their leg. The baton passes cleanly and the team wins.

Right now, a lot of businesses are trying to make AI run the entire relay race alone. All four legs start to finish, no handoffs. And it fails because that's not what the technology does well. But the businesses making money right now, they've set up a relay.

AI runs certain legs really fast, really reliably. And then it hands the baton to a human for the parts that need human judgment. Let me give you some practical examples. Let's say, for example, an email comes in.

AI reads it, categorizes it, pulls a relevant customer address and drafts a response. Then a human reviews and hits send in 10 seconds instead of three minutes. That's how you use it. That's how you relay between these different tools.

And the fact that AI fails 7.5% of fully autonomous complex work, that's not really bad news for you. It's good news. And here's why. Because it means the people who know how to build the relay, who know how to design which steps AI handles and which steps humans handle, who know how to set up the connections, feed it the right information and build a system that actually works on a Monday morning, those people are worth a lot of money right now.

And there's almost none of them. So MIT's warning is blunt. The next 18 months will determine which side of the divide your company lands on. Enterprises are locking in vendor relationships.

And the more a system learns your data, the harder it becomes to switch. On top of that, every week without an agentic AI strategy, it's not just a lost opportunity, it's growing technical and organizational debt, right? So every week you don't learn this, every week that you don't apply this is a week where someone else is learning it and applying it, right? So they're going in one direction and you're going in the other, and that creates a huge technical debt, a huge opportunity waste, right?

Here's for example, like people who learned Facebook ads in 2012, they built agencies that dominated, right? The people who learned SEO in 2005, they're crushing it with SEO now, right? So it's one of those things where if you learn it early, you're gonna win and get ahead. If you leave this, then you're gonna fall behind.

And that's what it's all about. And if you're watching this and you're a business owner or a freelancer or a creator or someone who wants to build something, this is the moment I want you to really hear. The study says AI fails 97.5% of jobs. That's true for a fully autonomous AI running on its own without anyone who knows what they're doing to set it up correctly.

But the people in the AI Profit Boarding, they're learning exactly how to set it up. The AI Profit Boarding has 2,600 members learning AI automation right now, right? Real automation, not theory, step-by-step video tutorials, 30-day roadmaps to take you from no idea what you're doing to having actual working systems in your business, four weekly coaching calls where you can ask for help and step-by-step roadmaps. Plus there's always someone online so you can get help whenever you want.

So the gap between the 95% of businesses wasting money on AI and the 5% making with it, that gap is knowledge. That's it. The technology is the same. The tools are the same.

The difference is knowing how to use them. Feel free to go to the AI Profit Boarding to check it out. Now, analysts expect the AI agent market to grow from $5.4 billion in 2024 to between $50 billion and $105 billion by 2030 to 2034, with 40% of enterprise apps, including task-specific AI agents by the end of 2026. That's a 20x increase in under 10 years.

And the money is not just in building AI, it's knowing how to deploy it, how to automate it, how to build workflows. And I want to zoom out for a second and talk about this because this is just the beginning, right? These AIs are only going to get better. They're only going to improve.

And the way that I would look at it is this, right? Step number one is just stop thinking about AI as a single thing, that ever works or it doesn't. Start thinking it as a toolkit with different tools for different jobs. A hammer is not a screwdriver.

AI that writes first drafts is not the same as AI that runs fully autonomously on multi-step workflows. Know which tool you're using and what it's actually designed for. Step number two, I would find the repetitive work in your life or your business, not the creative-like, judgment-heavy, relationship-driven work, the stuff you do the same way every single time. That's your first AI automation target.

It could be answering the same customer questions. It could be writing content, whatever it is. That's where you start. And then step number three, to make sure that your tasks don't fall in that 97%, build the relay, not the solo runner.

Don't try to automate an entire complex process from start to finish on day one. Find one specific step in your workflow, automate that step and keep a human handling everything else. Once that works reliably, add another step, slowly, carefully, test every stage. And then step number four is learn from people who are actually doing it in production, not theorists, not researchers, not demo builders.

People who are running real automations in the business. That's what we actually built the AI Profit Boarding for, and that's what can help you learn and get resolved to this. Link in the comments description, or just go to the AIProfitBoarding.com. Thanks for watching, and I'll see you in the next one.

Cheers, bye-bye.

More episodes

Browse all episodes →