Karpathy’s AI Auto Research: 126 Experiments While You SleepDiscover how Andrej Karpathy’s 'Auto Research' project is revolutionizing efficiency by allowing AI agents to run hundreds of experiments autonomously. Learn how this loop delivered massive gains for Shopify and why automated testing is the next frontier for business growth.00:00 - The 126 Experiment Breakthrough00:45 - Who is Andrej Karpathy?01:09 - Auto Research Explained03:03 - Shopify’s 53% Performance Jump04:18 - 36,000 Marketing Experiments Per Year05:53 - The 630-Line Kapi Loop07:12 - How to Implement the 3-Step Pattern09:14 - The Future of Machine-Speed Progress
Full transcript
Kapafi's AI auto research just shocked the world. So Andre Kapafi went to sleep one night in early March 2026 and whilst he slept an AI ran 126 experiments, not one or two, not 10, 126 all on its own testing things, trying stuff, keeping what worked, throwing out what didn't over and over again whilst he was unconscious in his bed. By morning the AI had made his system 11% faster, not a human, not a team, an AI agent running loops whilst one man slept. That is what auto research is.
And once you understand what that actually means, not just for labs and researchers but for you running a business or a brand right now, I think it's going to hit you harder than anything I've covered on this channel. So let's get straight into it. Who is Andre Kapafi? Super short version, he helped start OpenAI, he ran AI at Tesla, he coined the term vibe coding, he has 1.9 million followers on X and when this man posts something the entire world stops and pays attention.
On March 9th he posted a project called auto research. The goal is in his own words engineer your agents to make the fastest research progress indefinitely and without any of your own involvement. That post got 8.6 million views in two days. Why did it explode like that?
Because what he built isn't just a cool tech demo, it's a preview of how everything gets done from here. Let me explain what auto research actually is in plain English right. So you know when you're trying to improve something, your ads copy, your email subject lines, your pricing page, you try one thing, wait, check the results, try another thing, wait, check again. That process can take days, sometimes weeks.
A good marketing team might run like 30 tests in a whole year. The best ones maybe push 52. Auto research runs 12 experiments per hour, about 100 whilst you sleep. That's the whole idea, you set a goal, you tell the AI what better looks like and then you walk away.
The AI tries things, checks if they work, keeps the wins, throws out the losses and then tries again all night every night without you touching anything. Think about what that means for a second right. The bottlenecking improvement has always been human time. You can only test so many ideas a day, you have a business to run, clients to talk to, content to make, emails to answer.
You get maybe five or ten tests a week if you're really disciplined. Auto research does that before breakfast. Now here's what made people lose their minds in one overnight run. McCarthy's agent completed 126 experiments driving measurable performance improvements across the board.
He then left it running for two full days on a more complex setup. The agent ran about 700 experiments and found roughly 20 improvements that all stacked on top of each other, dropping his system's training time from 2.02 hours to 1.8 hours. An 11 efficiency gain on a project he believed was already well optimized. Here's the crazy part, the agent caught oversights and attention scaling and regularization that McCarthy himself had missed after two decades of working on this stuff.
Let that sink in. One of the most respected researchers on the planet, 20 years of experience and the AI found things he missed. That is not a minor footnote, that's the whole story. Now I want to tell you about what happened when Toby Luker tried it right and this is where it gets real.
So Toby is the founder and CEO of Shopify, one of the biggest e-commerce platforms on earth. Over a million businesses run their stores on Shopify. He heard about auto research and decided to point it at one of Shopify's internal systems, the engine that runs every single storefront on the platform. His agent found a 53% speed up in performance, 61% fewer resources used through 93 automated changes.
Not a team of engineers working for months, an AI agent running overnight and it didn't stop there. Luker also used auto research to optimize an internal AI model, telling the agent to improve both quality and speed. After one overnight run of 37 experiments, the agent delivered a 19% performance gain and the optimized model was smaller, 0.8 billion parameter model outperforming the previous 1.6 billion parameter model. Smaller and better, less and more, the AI got smarter by simply getting simpler.
That sentence right there is the one I keep coming back to because it goes against everything our intuition tells us. We assume bigger means better, more means more, but the AI ran the experiment we never would have run because they seemed backwards and discovered the truth by doing that. That's what happens when you remove the human bias from the testing loop. Now here's where it gets very interesting for people like you.
The tech world looked at auto research and saw a machine learning tool and that's fine, but other people looked at it and saw something much bigger. So Eric Seager, the founder of Single Grain, a major marketing and ad agency posted on X saying most marketing teams run around 30 experiments a year, maybe 52 if they're really good, a new landing page or a new ad creative, a new subject line. That's what passes for data-driven marketing today, but in his view the next generation of marketing systems will run like 36,500 or more experiments per year. 30 versus 36,000.
Same time period, same resources, different approach. That's not a small upgrade, that is not just getting better, that is a completely different game being played on the same field and the same loop applies to anything that you can measure. It could be your email open rates, could be your ad click rates, could be your content hooks, your product pricing, even your checkout page or your headlines. Anything where you can say this result was better than that result, that can be auto-researched.
Here's an example that's directly relevant to creators and agencies. So Auto Voice Evals applied the same loop to optimizing AI prompts. 20 automated iterations improved a scheduling agent's success rate from 25% to 100% and the final prompt was shorter, not longer, from a quarter success rate to perfect. 20 rounds of testing that a human would never have had the patience to run manually.
An e-commerce brand could point this loop at their product description prompts, an agency could point it at the outreach templates, a content creator could point it at the hook formats. The mechanism is the same but you define what better looks like, you let it run, you wake up the results, right? And now I know what some of you are thinking. You're thinking this sounds cool but it's probably for tech companies with huge budgets and engineering teams.
But it's the wrong frame. The entire auto-research repository is 630 lines of code. The github star count hit 42,000, it's climbing right now. Fortune magazine called it the Kapafi loop.
It runs on a single GPU. You can rent GPU time by the hour for a few dollars and the barrier to using something like this is not money, it's just awareness. It's knowing the pattern actually exists. That's exactly what I'm talking about.
In just 17 hours of running, AI agents using the auto-research pattern independently rediscovered machine learning milestones like RMS norm and tied embeddings that took human researchers at Google Brain and OpenAI nearly eight years to formalize. Eight years of human progress, 17 hours. Now I want to be straight with you here. Kapafi himself was clear that his original setup was a relatively small project, right?
630 lines of code, one GPU, a contained system. He acknowledged that scaling this to a frontier AI lab with millions of lines of code is a lot more complex. But he also said it's just engineering and it's going to work. So it's not magic yet at the larger scale.
It's not replacing OpenAI's whole research team tomorrow. But that's not the point. The point is the pattern, the loop. And that loop is already working right now on real business systems today.
Let me talk about what the pattern actually is because once you see it, you can't unsee it. The loop has three steps. Step number one, try something. Step number two, measure whether it works.
And step number three, keep it or throw it out and try something else. That's it. That's the whole thing. The reason it's powerful isn't the three steps.
Humans have always done three steps. The reason it's powerful is because of the speed and the iteration and the volume and the fact that AI doesn't get tired. It doesn't lose interest. It doesn't get bored.
It doesn't get distracted by a Slack message, right? It doesn't decide to skip testing on a Friday afternoon because it's going out on the weekend, right? Kapafi described the goal as making the AI handle the full cycle. So hypothesis, experiment, evaluation, next iteration.
Without needing human input between rounds, you become the person who sets the direction that AI runs the laps. Think about how your business actually works today. You probably have things you know you should test, right? Your email subject lines, your lead magnet headline, the first line of your social media posts, your pricing page layout, your proposal format.
Testing these would help. But when do you actually do it? Probably never, right? Because you're too busy.
Because it takes too long. Because by the time you set it up, some other priority appears. What if you could just describe what you want improved and then just walk away? That's what's on the table here right now.
The auto research pattern works on any domain where you can run something and measure result, right? So the key insight is that you don't need to be an AI researcher to apply it. You just need a clear metric, a tool that can run the experiment, and something to optimize. So for a freelancer, that might be which version of my pitch email gets more replies.
For an agency owner, it might be which version of my onboarding sequence retains clients longer. For an e-commerce store, it might be which product page layout drives more add to cart. The experiments run, the AI measures, the AI picks a winner, and the AI runs the next round. Here's something else Kapafi said that stuck with me.
He said the goal of auto research isn't to replace like a single PhD researcher. The goal is to emulate an entire research community of them. A swarm of agents collaborating, exploring different directions in parallel, promoting the best ideas upward. Not one person, a whole community working on your stuff whilst you sleep.
Let me zoom out and show you what this means for the speed of AI development itself. This is a part people aren't quite taking and talking about enough, right? Every major AI lab in the world, OpenAI, Google DeepMind, Anthropic, Meta, Mistral, is now either using auto research style loops or will be very soon. Kapafi himself said all LLM frontier labs will do this.
It's a final boss battle. Think about what that means for the pace of AI improvement right now. We've seen this with Minimax, which actually iterated on itself and used improvement loops to iterate 100 times and actually made itself get 30% better with the latest release of Minimax M2.7. So AI can get better when human researchers run experiments, analyze results, write papers, have meetings, debate what to try next, and that loop runs at human speed, which is pretty slow.
A few experiments a day, a few breakthroughs a year, but when AI agents start running the experiments on the AI itself, that loop runs at machine speed. A team at Skypilot gave the auto research agent access to 16 GPUs on a cluster and let it run in parallel. Over eight hours, it submitted around 910 experiments and drove measurable improvement. With one GPU, it was stuck testing one idea at a time.
With 16, it ran grids of 10 to 13 experiments per wave, catching interactions between variables that sequential testing would have missed entirely. One GPU, 12 experiments per hour, 16 GPUs over 100 per hour simultaneously, and it gets smarter about what to test next as it goes. The acceleration is real, the curve is steep, and here's the thing I want you to sit with. The businesses that figure out how to apply this loop to their own operations, not just to AI training, but to their marketing, their offers, their content, their processes, those businesses are going to be running a fundamentally different operation than the ones that don't.
This isn't hype, the Shopify example is real, 53% faster rendering across every storefront, 19% model improvement overnight. These are real companies with real numbers, and if you're on an agency and you run 30 tests a year on your client campaigns, and then your competitor starts running like 3,000, you are going to feel that gap. They're going to have a competitive advantage. So at this point, I want to talk to you about something that most people miss, because when they hear developments like this, their instinct is, for example, to say that's interesting and just move on.
And then six months later, they wonder why things feel harder than they used to, why clients are harder to find, why see old approaches don't work as well. The people winning in AI right now are not the ones that are the most technical. They're the ones who understand the pattern, the facets, and then find a way to actually implement it. For example, an agency owner who figures out how to run automated testing loops on the client deliverables doesn't need a data science degree.
They need to understand the concepts, they need to know what to measure, and they need to take action before the competitors do. That's exactly what we do inside the AI Profit Boardroom. So we give you tutorials, tips, techniques, 30-day roadmaps, and video coaching, so you can get help and support whenever you need it. It's an awesome community, and if you want to join that, link in the comments description, or go to the AIProfitBoardroom.com.
More episodes