OpenAI GPT 5.4 Released: The World's Smartest AI for Real Work?
OpenAI’s new GPT 5.4 model marks a massive leap in professional AI performance, introducing native computer use and deep Excel integration. Learn how it stacks up against competitors in coding, math, and reasoning to see if this truly is the world's smartest AI for real-world tasks.
00:00 - Intro: GPT 5.4 is Here
01:10 - Native Computer Use Capabilities
03:59 - Dominating Professional Benchmarks
07:23 - Cracking Frontier Math
10:40 - GPT 5.4 vs. Gemini and Claude
14:24 - The New Excel Integration
18:11 - Cyber Security & Safety Risks
20:49 - Pricing and Market Impact
Full transcript
OpenAI just built the world's smartest AI. So OpenAI just released GPT 5.4 and they're calling it the most efficient and capable frontier model for professional work. And honestly, depending on how you define the smartest, that might have a real claim here, but it's complicated because GPT 5.4 is not the best at everything. It wins in some areas that frankly matter more for your daily life, more than any AI model ever has done.
And we'll get into that in a minute. It also loses in other areas where you might expect it to dominate. So let's actually break down what this model did, where it wins, where it loses, and why it matters, because the answer to is this the world's smartest AI is actually more interesting than a simple yes or no. Let me start off with the stuff that blew my mind.
So on March the 5th, 2026, OpenAI released GPT 5.4 in three versions, GPT 4.0 5.4 Thinking, which is the standard version for chat GPT users, GPT 5.4, the API version, and GPT 5.4 Pro, which is the maximum performance version for pro and enterprise subscribers. And the headline feature that I think most people are gonna underestimate is computer use. GPT 5.4 is the first general OpenAI model with native computer use capabilities. And what that means is that this model can look at your screen, see what's on it, understand it, and then actually operate the computer, move the mouse, click buttons, type into field, switch applications, fill out forms, navigate websites, complete entire multi-step workflows across different pieces of software by itself without you touching anything.
Now on the OS world benchmark, which measures how well something can navigate a real world desktop environment using screenshots and keyboard and mouse actions, GPT 5.4 scored 75%. The human expert baseline on this exact same benchmark is 72.4%, and the previous model, GPT 5.2, scored 47.3%. So we went from 47 to 75 in a single model generation. And the AI now scores higher than the human experts who were specifically selected and paid to do this test perfectly.
That is not an incremental improvement. That's crossing the line. And I've been looking at some of the real world results. There's a company called Mainstay that works with about 30,000 property tax portals.
These are gnarly government websites with dense layouts, tiny buttons, and complicated forms. They pointed GPT 5.4 at this portal's first attempt success rate was 95%, by the third attempt, 100%. And it completed the sessions about three times faster than previous computer use models whilst using around 70% fewer tokens. So faster, cheaper, more accurate, all three at the same time.
Another company called TinyFlash that builds web agents for live dynamic websites, not sandbox demos, real websites that change, they called GPT 5.4 by far the best open AI model they've tested. They specifically called out improvements in click accuracy on cluttered screens, long trajectory reasoning across hundreds of steps, faster evaluation cycles, and the fact that the model maintains spatial memory across steps, meaning it remembers where things are on the screen even hundreds of actions later. Now, this is the part where I think most people aren't connecting the dots. Think about how many jobs in the world are essentially just someone sitting at a screen or a computer looking and clicking buttons to move information from one place to another.
Data entry, scheduling, bookkeeping, customer service, expense reporting, HR onboarding, insurance claims processing, all those jobs were protected by one simple fact, AI couldn't see the screen and click the buttons, and that barrier is now gone. I'm not saying those jobs will all disappear tomorrow, that's not how technology disruption typically works, but the technical barrier that was protecting them has evaporated and that changes the conversation fundamentally. So let me get into the benchmarks I think actually matter. These aren't the standard saturated benchmarks that we look at every single time, these are new ones that measure real professional work.
The first one is the APEX Agents Benchmark. This was built by Mercor and it launched in January 2026. It contains 408 tasks across simulated work environments, real professionals, investment bankers, consultants, lawyers. They spent five to 10 days building each scenario with actual files, tools, and software.
The models have to complete the tasks the same way a junior employee would, single shot, no asking for help, financial modeling, slide decks, legal memos, market analysis. When this benchmark launched in January 2026, the best score from any AI model in the world was 24%, GPT 5.4 just scored 52%, first model ever to cross a 52% benchmark. That's roughly double the score in about six to eight weeks. And to be fair, the CEO of Mercor, Brendan Foody, said when the benchmark launched that right now it's fair to say it's like an intern that gets it right a quarter of the time.
Well, now it's getting it right half the time and it's improving at a pace that should make you pay attention. Now, the second benchmark is GPT VAL, so GDP VAL. This is OpenAI's own internal benchmark and I want to be transparent about that. They made it themselves, so take it with that context.
But here's what it tests. It measures AI performance against real world knowledge workers across 44 occupations in the nine industries that contribute the most to GS and US GDP. Sales presentations, spreadsheets, scheduling, legal work, financial analysis, real deliverables of real people produce at real jobs. GPT 5.4 matched or exceeded human professionals in 83% of those comparisons.
That's up to 70.9% on GPT 5.2. Now, the important caveat, these are well-defined single-shot tasks. The model gets one attempt at a clearly specified task. Real jobs involve back and forth, iteration, context, relationships, judgment.
So 83% on clean tasks doesn't mean that AI can do 83% of your job, but the execution layer, the sit down and produce a deliverable layer, that is getting dominated. And here's a stat that really jumped out at me. OpenAI says the models complete these tasks roughly 100 times faster and 100 times cheaper than a human expert. 100 times faster and 100 times cheaper.
Even if you cut those numbers in half for real world conditions, the economics are staggering. Harvey, the AI legal setup, tested GPT 5.4 on their big law benchmark and it scored 91%. They said it's better at structuring complex transactional analysis, maintaining accuracy across lengthy contracts and delivering the detail that legal practitioners need. And on OpenAI's internal investment banking benchmark, the kind of work that junior analysts at Goldman and Morgan Stanley do, like building three statement models for Fortune 500 companies with correct formatting and citations, GPT 5.4 scored 87.3%.
For context, that was 68.4% on GPT 5.2 and 43% on the original GPT 5. So we've gone from 44 to 87% on real investment banking tasks in less than a year. Now, let me talk about maths because this is where it gets genuinely surprising. GPT 5.4 set a new record on frontier math.
For those who don't know, frontier math is a benchmark of extremely challenging maths benchmarks and problems designed by professional mathematicians. These require research level thinking, not textbook problems, genuinely novel, hard problems. When the benchmark first launched, top models scored around 2%. GPT 5.4 Pro scored 50% on tiers one through three and 38% on tier four, which is the absolute hardest tier.
Now, I want to be really precise about this because there's important context. Epoch AI, which runs the benchmark independently, noted that OpenAI funded frontier math and has exclusive access to all 290 problems in tiers one through three, plus solutions to 237 of those problems and 28 of the 48 tier four problems. That said, Epoch tested the held out separate set properly. That said, Epoch tested the held out set separately, the problems OpenAI did not have access to, and the differences between the held out and non-held out scores were not statistically significant.
So the model appears to genuinely be solving these, not just memorizing. Now, here's the moment that got everyone's attention. GPT 5.4 Pro solved one tier four problem that no model had ever solved before. However, and this is important for accuracy, Epoch AI's preliminary analysis found that the model appeared to have found a research preprint from 2011 that let it shortcut much of the intended work.
The problem author was not aware of this preprint, so it's not quite as magical as it might sound. The model found a clever shortcut rather than deriving a novel solution from scratch. But there was a separate moment that was genuinely remarkable. When they ran GPT 5.4 with extra high reasoning 10 times on tier four, it achieved a pass at 10 score of 38%.
And in one of those runs, it solved another problem that no model had solved before. This was by a mathematician named Bartosz Naskerecki. And his reaction was incredible. He called it his personal move 37, which is a reference to AlphaGo's famous move in 2016 that shocked every Go expert on the planet.
He said the solution was very nice, clean and feels almost human. This was a problem he had personally curated for about 20 years, two decades of work and then the AI cracked it. Now to be fair, GPT 5.0 was tested on genuinely open math problems, problems where no human knows the answer yet. It did not solve any of them.
It made some novel observations on one problem, but the author described those as relatively uninteresting. So the model can solve extremely hard problems that have known solutions, but it cannot yet make mathematical discoveries. That distinction matters. Ok so now let me talk about where GPT 5.4 is not the smartest, because this is where the honest picture gets really interesting.
On abstract reasoning, Google's Gemini 3.1 Pro actually beats GPT 5.4. On GPQ-A Diamond, which is a graduate level science benchmark, Gemini scores 94.3% vs GPT 5.4's 92.8%. On ARC-AGI-2, which tests fluid intelligence and novel pattern recognition, Gemini leads at 77.1% vs GPT 5.4's 73.3%. Now GPT 5.4 does close that gap significantly at 83.3% on ARC-AGI-2, but that costs $30 per million input tokens and $180 per million output tokens.
Gemini 3.1 Pro does its reasoning at $2 input and $12 output. That's a massive price difference for similar or better reasoning performance. On pure coding, called Opus 4.6 from Anthropic still leads. Opus holds the highest SWE bench verified score at 80.8% and on expert level visual reasoning, it scores 85.1% on MMU Pro.
If you're doing heavy software engineering, debugging complex codebases or building multi-agent systems, Opus 4.6 is still the model to beat. And here's something really interesting from the Artificial Intelligence Index. GPT 5.4 with extra high reasoning and Gemini 3.1 Pro preview are literally tied at 57 on their composite intelligence score. Called Opus 4.6, it's just behind at 53.
These models are not in different leagues anymore. They are effectively tied, each winning in different categories. I also found something buried in the technical report that I think is worth mentioning. On a benchmark called GPOPQA, which tests genuinely novel hard engineering problems, the kind of stuff that took an actual team at OpenAI over a day to solve, GPT 5.4 thinking scored 4%.
That's actually worse than GPT 5.2 and 5.3 codecs, which scored 8%. On the hardest novel engineering problems, this model actually regressed. OpenAI put this in their own report, they didn't hide it, but I do think it's worth noting. On accuracy, OpenAI says GPT 5.4 is 33% less likely to make individual factual errors and 18% less likely to have any errors in a full response compared to GPT 5.2.
But when it's wrong, it tends to be confidently wrong. About 89% of the time, it's errors come wrapped in confident sounding language. That's a real problem because humans naturally trust confident answers. So is this the world's smartest AI?
Here's my honest take. It depends entirely on what you mean by the smartest. If smartest means the best at raw abstract reasoning, the answer is no. Gemini 3.1 Pro wins that.
If smartest means the best at coding, the answer is no. Claude Opus 4.6 wins that. If smartest means best at actually doing professional work, the kind of work that contributes to GDP, the kind of work that real humans do at real jobs every single day, then yes, GPT 5.4 has the strongest claim to that title by a significant margin. And I think that second definition, the doing real work definition, is actually what matters to most people.
Because most people are not professional mathematicians solving novel theorems. Most people are not software engineers debugging complex code bases. Most people are knowledge workers who sit at computers and produce deliverables. And that's exactly where GPT 5.4 dominates.
Now let me talk about the Excel integration because this is where the rubber meets the road for millions of people. OpenAI launched ChatGPT for Excel in beta on March the 5th. This is not just a feature. This is ChatGPT embedded directly inside your Excel spreadsheet as a sidebar panel.
You type in plain English what you want. Build me a three statement financial model using the data in this workbook. And it does it. It reads your existing sheets.
It understands how your formulas connect across tabs. It builds new formulas, creates new sheets, formats everything, explains what it did. And before it changes anything, it asks your permission. It's available right now in the US, Canada and Australia for ChatGPT business, enterprise, pro and plus subscribers.
Google Sheets version is coming soon. For enterprise and education workspaces, it's off by default and admins have to enable it. But the Excel integration is not even the biggest part. OpenAI simultaneously launched integrations with nine financial model providers.
Moody's, Dow Jones, Factiva, Factset, S&P Global, LSEG, the London Stock Exchange, MSCI, Third Bridge, Dalupa, MT Newswire. What that means is a finance professional can now sit in ChatGPT, pull live market data and Moody's and S&P Global, feed it into an Excel mode built by GPT 5.4, run scenario analysis, generate an investment memo with citations and sources, and export it all as a PDF, all in one place, all in a fraction of the time. And Fortune Magazine pointed out that this puts OpenAI in direct competition with Anthropic in the entire space. Anthropic launched Claude for financial services last year and then released Clowork, which lets AI work with desktop applications.
When those Clowork plugins launched, it triggered a sell-off across SaaS stocks. The marketplace panicked because investors realized AI tools might start making some judicial software companies obsolete. OpenAI is now going even more aggressive with the same playbook. ChatGPT inside Excel, real-time financial data integrations, native computer use.
They're not building a better chatbot, they're building and trying to replace chunks of the enterprise software stack. Now I want to talk about the context window because this is a huge upgrade that most people are glossing over. GPT 5.4 supports up to a million tokens in codecs and the API, that's roughly 750,000 words, about 10-12 full-length novels, or an entire year of financial reports for a major company. Previously, the biggest limitation was that AI could only work with small chunks, give it one document, fine, but ask it to reason across 100 documents and understand how they all connect, and it would lose the thread.
GPT 5.4 can hold all of that in context. Though it's worth noting that on long-context retrieval benchmarks, performance does decline at extreme ends. On the MRCR needle retrieval benchmark, it's strong through 128 tokens at 86% but drops to 36.6% at the 512k-million token range. So the million token window exists, but the model's ability to find specific needles in that haystack degrades at the extremes.
There's also something called tool search that's more technical but really important for developers. Previously, when you wanted AI to use external tools, you had to describe every single tool up front. GPT 5.4 gets a lightweight list and looks up tool details only when needed. This reduced total token usage by 47% whilst maintaining the same accuracy on 250 tasks across 36 different tool servers.
That means agents built on GPT 5.4 are meaningfully cheaper and faster. Now let me talk about the cybersecurity findings because this is buried in the technical report and I think it's really important. GPT 5.4 thinking is the first general purpose model that required OpenAI to implement specific safety mitigations for higher capability and cybersecurity. When they tested it against professional level capture, the flag challenges, it achieved an 88% success rate.
In a simulated network event, the model executed complex multi-step attacks, including exploiting vulnerable web applications to steal credentials and moving laterally through a network. Under OpenAI's preparedness framework, high cybersecurity capability means they believe the model could potentially automate end-to-end cyber attacks against hardened targets. Every model generation has seen these scores go up. GPT 5.2 scored around 47%, GPT 5.3 scored around 80% and now GPT 5.4 is rated high with 88% on capture the flag.
The question I have and I think we all should be asking is what happens when GPT 6 or GPT 7 hits a critical level? Because under OpenAI's own framework, critical means it could cause catastrophic large scale damage autonomously, attacks on power grids, water systems, financial networks. And right now anyone with an API key can access GPT 5.4. The capabilities are advancing faster than the access controls.
I'm not advocating for any specific policy here, I'm just saying the trajectory is worth watching. Now let me also cover pricing because this matters if you're deciding which model to use. A GPT 5.4 standard costs $2.5 per million input tokens and $15 per million output tokens. That's actually slightly more expensive per token than GPT 5.2 but because it uses fewer tokens to solve the same problems and tools.
cuts usage by 47%. The effective cost of getting work done can be lower. GPT's 5.4 Pro, the most powerful version, costs $30 per million input tokens and $180 per million output tokens. That's really expensive.
For comparison, Quad Opus 4.6 is $5 input and $25 output. Gemini 3.1 Pro is $2 input and $12 output. So the pricing creates interesting trade-offs. Gemini gives you similar or better reasoning at a fraction of the price.
Quad gives you better coding at a moderate price. GPT 5.4 gives you the best professional work and competing user at a premium. And GPT 5.4 Pro gives you maximum performance on everything at a price that's really only practical for enterprises. For regular ChatGPT users, GPT 5.4 Thinking is available on the Plus plan at $20 a month.
Pro is $200 a month. And I do want to point out something kind of ironic. Remember all those people saying intelligence would be too cheap to meter. That's the cost of AI.
That the cost of AI would just keep falling towards zero. Well, $30 and $180 per million tokens for the Pro version is not exactly cheap. The base level stuff is getting cheaper, but frontier reasoning, the really powerful stuff, is actually getting more expensive. And I think that creates a really interesting dynamic where the best AI becomes a premium product that mostly enterprises can afford to use at scale.
Okay, let me also cover the competitive pace because the context matters. OpenAI released GPT 5.3 on March 2nd. Then two days later, March 5th, they released GPT 5.4. Two models in three days.
And before that, Anthropic released Claude Opus on February 4th. Google released Gemini 3.1 Pro on February 19th. Three frontier models from three companies in one month. The competition is driving improvement at a pace we've never seen.
And the creative writing side is worth mentioning because OpenAI took a lot of heat for GPT 5.2 being, for lack of a better word, robotic. Sam Altman, gone on stage and admitted they messed up. They focused so much on maths and coding that the model lost its personality. GPT 5.4 has improved significantly.
On the LM arena, creative writing leaderboard, it's already ranked second with only 390 votes so far. I tested it myself and it genuinely feels more natural, more warm and more human. That matters because nobody wants to use a model that's annoying to talk to no matter how smart it is. And the coding demos have been impressive.
Someone built a full theme park simulation game from a single prompt using the new playwright interactive feature. The model built the code, opened a browser, played the game, found bugs, went back, fixed them and repeated. That full loop of build, test, see, fix is genuinely new and it's going to change how software gets made. Samuel Albany created five one-shot demos.
Every ML paper on arc sieve from 2012 to now as a particle simulation. Every naval battle, Lord Nelson IV. Isometric king's cross station, a self-assembling conference poster. All single prompt, all first try.
These aren't cherry pick, this is a pattern. Now here's where I need to talk to you directly about what all of this means. The world economic forum's future of jobs report says 40% of employers are planning to reduce staff this year. In the first six months of 2025, almost 78,000 tech jobs were attributed to AI.
A survey of US companies using chat GPT found that 49% have already replaced workers and 37% of business leaders say they expect to replace workers with AI by the end of 2026. And now you layer computer use on top of that. Chat GPT inside Excel on top of that. Million token context windows for financial data integrations with Moody's and SAP global on top of that.
But look, I don't just want to scare you because the other side of this is genuinely exciting. The barriers to building things have never been lower. Sam Altman says he believes a one person company can now have and become worth more than a billion dollars. And whether you agree with him or not, the direction is right.
The tools that required teams of dozens of people are now available to individuals. I hear people every single week inside the AI profit boardroom who are building real businesses with their tools. Real revenue, real customers, not hypothetical stuff. A person with a laptop and the right skills producing more than a small team did two years ago.
And this is exactly why I created the AI profit boardroom because I've been making these videos every single day watching this unfold in real time and the pattern is undeniable. The people who learn these tools first are the ones who benefit the most. Every week the evidence gets stronger. The AI profit boardroom is a community of people who are actually using AI to build automations, launch businesses and generate real income.
Not watching from the sidelines. So if you want to join feel free to check it out, link in the comments description or just go to theaiprofitboardroom.com. Now let me give you some concrete next steps. If you're a coder test GPT 5.4 but also keep testing Claude Opus 4.6.
Opus still leads on pure software engineering. GPT 5.4 is better for agent workflows and computer use. Use both root tasks to the model that's best at each specific thing. If you work in knowledge work, spreadsheets, presentations, analysis, reporting, get on chat GPT for excel the moment it's available in your region.
Learn to prompt it effectively. The people who master this are going to produce in one hour what used to take a day. If you're a manager start thinking about what your team does versus what AI should do. Gartner says multi-agent systems are moving from pilot projects to enterprise standards in 2026.
If you're not piloting this you'll be hired. If you're a founder or creator this is the best time in history to start building. The tools have never been more powerful, more accessible but you have to actually learn them. And if you're someone who doesn't want to pick between models honestly that's a smart play.
The best approach in March 2026 is to use GPT 5.4 for professional tasks and computer use. Use Gemini 3.1 for cost sensitive tasks and Claude Opus for heavy coding. No single model wins everything. The people who understand that and root accordingly are the ones who will get the most out of AI.
So to answer the question in the title did OpenAI just build the world's smartest AI on professional work? Yes. 83% GDP value, 91% big law, 87% investment banking, 75% computer speed in humans. Those numbers are real and the best in the world.
On raw reasoning no. Claude wins that. But here's the thing most people the professional work scores are the ones that actually matter because most people aren't doing abstract reasoning puzzles or debugging complex code bases. Most people are producing deliverables at desks and at that GPT 5.4 is genuinely the best AI model in the world.
Make of that what you will. If you got value from this hit subscribe, hit the bell and share with someone who needs to hear it. I guarantee you there's someone in your life who's still pretending this isn't happening. I'll see you in the next one and feel free to check out the AI Profit Bulletin if you haven't already.
More episodes