AI News Today
← All episodes
Episode 1 · February 18, 2026 · 15:04

Claude Sonnet 4.6, Computer Use Benchmarks, The AI Repricing Event - with julian goldie

Julian Goldie breaks down why the release of Claude Sonnet 4.6 is more than a simple update—it is a structural repricing event for the entire AI industry. We explore the massive leap in computer-use capabilities, the 1 million token context window, and how flagship-level performance is now available at a fraction of the cost.


TIMESTAMPS

00:00 Claude Sonnet 4.6 & The Repricing Event

01:15 OS World: AI Using Computers Like Humans

02:45 Breaking the 72.5% Benchmark Barrier

04:30 The Cost Revolution: Flagship Power, Mid-Tier Price

06:00 Enterprise Results: Box, Replit, and 1M Context

08:15 Claude Code & The Future of AI Development

10:30 The Economic Collapse of Legacy Software

12:00 Advice for Developers & Knowledge Workers

14:15 The Compounding Advantage of AI Native Orgs

16:00 The Human Element: Staying Ahead of the Curve

Full transcript

Something just happened that I don't think most people fully understand yet. So Anthropic dropped CloudSonic 4.6 and on the surface this seems like a routine model update. Another version number, another benchmark chart, another press release. But when you actually dig into what's changed and more importantly what it means for the pricing equation that governs how AI gets deployed at scale, you start to realize this is not a routine update.

This is a repricing event and repricing events have consequences that ripple far beyond the people who read AI news. Let me start with the number that stopped me in my tracks, 72.5%. That is CloudSonic's 4.6's score on OS World. Now OS World is the benchmark that tests whether an AI can actually use a computer like a human uses a computer.

Not answers questions about computers, not write code, actually sit down at a screen, move a cursor, click buttons, fill out forms, navigate tabs, open spreadsheets and get work done the way you and I would do it. 72.5%. Now here's why that number is wild. In October 2024, 16 months ago, the very first cloud model with computer use capabilities scored 14.9% on that same benchmark.

Anthropic themselves called it still experimental, at times cumbersome and error prone. And they were being honest, it was proof of concept, to demo that something was theoretically possible. Sonic 3.7 released in February 2025 and that got it to 28%. Sonic 4 hit 42.2% by June, Sonic 4.5% climbed to 61.4% in October and now Sonic 4.6 is at 72.5%.

That is not a trendline, that is a phase change. In 16 months, the score nearly quintupled. The model went from kind of works to, you know, sometimes in controlled conditions to scoring 72.5% on tasks involving real software, Chrome, LibreOffice, VS Code in a simulated environment. And early users aren't just seeing benchmark numbers, they are reporting human level capability on tasks like navigating complex spreadsheets and filling out multi-step web forms across multiple browser tabs.

And I keep saying this and I will keep saying this until people really hear it. The thing that matters most about AI right now is not raw intelligence, it is the ability to actually do things. Not know things, do things. And Sonic 4.6 just crossed a threshold where it is doing things at a level that was until very recently only possible with models that cost 5 times as much to run.

Let's talk about that cost thing because it is the engine driving everything. Claude Sonic 4.6 is priced at $3 per million input tokens and $15 per million output tokens. That's the same price as Sonic 4.5. There's no increase, it's the same pricing.

Anthropic's flagship Opus models cost $15 per million input tokens and $75 per output tokens. Five times the price, five times. And what Anthropic is now saying and what their own testing backs up is that Sonic 4.6 performs at Opus level on a wide range of real world tasks. Not all tasks.

There are domains where Opus still wins, but for coding, for office work, for computer use, for the kind of tasks that businesses are actually trying to automate right now, Sonic 4.6 is right there. And you can feel what this means for enterprise deployment. If you're a company and you were just waiting because you needed Opus level performance but couldn't justify Opus level pricing at scale, that trade-off just disappeared. The argument for holding back on deploying AI agents into your workflows at scale just got significantly harder to make.

Box, the enterprise content management company, got early access to Sonic 4.6 and ran it through their actual testing suite, the kind of heavy reasoning tasks that enterprise customers use every single day. Sonic 4.6 hit 77% accuracy on those heavy reasoning benchmarks. That's up 15% points from 4.5's 62%. In public sector work, it hit 88% accuracy.

In healthcare context, 78%. In retail, 94%. These are not toy benchmarks. These are domain-specific tests that Box uses to evaluate whether a model can actually serve enterprise customers.

And Sonic 4.6 is clearing those bars at a price point that makes broad deployment viable. Michel Katasta, president of Replit, said it plainly, the performance to cost ratio of Claude's Sonic 4.6 is extraordinary. It's hard to overstate how fast Claude models have been evolving in recent months. I want to pause on that sentence.

Hard to overstate how fast they're evolving. That's not a content creator saying that. That's the president of a company that builds developer tools and has been watching Claude models closely because their business depends on understanding which models to build on. When someone in that position says it's hard to overstate the pace, you should probably take it seriously.

Now, let me explain the technical thing that actually matters here because there's one capability improvement in Sonic 4.6 that I think gets undersold in most of the coverage I've seen. The 1 million token context window. And I know, I know, context windows have been expanded for a while. Everyone keeps talking about them.

But stay with me because the way this plays out in practice has changed. Imagine you hired a consultant to review your company's entire code base. Millions of lines across hundreds of files and give you a diagnosis. Old AI models were like hiring that consultant but telling them they could only reach 50 pages at a time and they had to pull out all the other pages away and they could never look at two sections simultaneously.

They had to build everything from fragments. They were essentially working blind. A 1 million token context window means you can hand Claude Sonic 4.6 an entire code base, an entire legal contract, dozens of research papers, and it doesn't just hold them in memory. Anthropic explicitly says 4.6 reasons effectively across all of that context.

Not just holds it, but reasons across it. For developers using Claude Code, Anthropic's terminal-based coding tool that has become, and I do not use this word lightly, a cultural phenomenon in Silicon Valley, the real world difference is already showing up. Early testing found that users preferred Sonic 4.6 over Sonic 4.5 roughly 70% of the time in Claude Code. But here's the part that really got my attention.

Users even preferred Sonic 4.6 to Opus 4.5, Anthropic's previous flagship, more powerful, expensive model, 59% of the time. A mid-tier model beating the previous flagship at one-fifth of the price. They rated Sonic 4.6 as significantly less prone to over-engineering and laziness. They reported fewer false claims of success, fewer hallucinations, more consistent follow-through on multi-step tasks.

So if you're a developer and you've been frustrated with AI coding tools that confidently tell you they've completed a task and then you look at the code and it's half-finished or broken, that specific complaint is what Anthropic have been addressing here. Fewer false claims of success. That is such a specific and important thing to fix because the failure mode of an AI coding assistant that lies to you about what it did is way more dangerous than one that just says, I don't know, or I couldn't do that. False confidence costs engineering hours.

Now I want to zoom out because I think there's a bigger story happening here that the model release itself is actually just a data point inside of. We are watching the gap between frontier and near frontier AI models collapse in real time. Six months ago if a company needed flagship level AI performance for a serious enterprise application they had one option, pay for the flagship, accept the cost, build around it, price it into the product. Now they don't.

Now the mid-tier models does most of what the flagship did and in a few months the new Sonic will probably be even better and Haiku, the smallest, fastest, cheapest model in the Claude family, will eventually catch up to where Sonic is today. This is Moore's law behavior applied to AI capability relative to price. The same capability keeps getting cheaper. What costs you five dollars today will cost you one dollar in 12 months and what costs you one dollar today will eventually cost fractions of a cent.

The implications of that compression for software business models are significant. We are watching software stocks sell off. The iShares expanded tech software sector ETF, a basket of major software companies, has dropped more than 20% year to date. Investors aren't just being nervous, they're doing maths.

If AI can do knowledge work at human competitive quality for dollars per million tokens and if that price keeps falling, what does that do to the value proposition of software that has been charging premium prices for tools that help humans do that knowledge work more efficiently? I'm not telling you that software is dead. I'm telling you that the economic assumptions baked into a lot of software pricing are quietly being stress tested right now. And Claude Sonic 4.6 is one of the things doing the stress tests here.

Now let's talk about what this release actually means in practice for different people watching this. If you're a developer, you should be running Claude Sonic 4.6 and Claude Code today. The preference data is clear. 70% of users prefer it over Sonic 4.5.

59% prefer it over the previous flagship. The context window means you can throw entire projects at it. The improvements in instruction following and follow through mean you lose less time cleaning up hallucinated input and output. This is your default tool now.

If it isn't already, make it happen this week. If you are a non-technical knowledge worker, a marketer, an analyst, a consultant, an ops person, a financial person, the computer use improvements in Sonic 4.6 are actually more immediately relevant to you than the coding improvements because computer use is AI that can navigate the actual software you use every day. That legacy CRM system that doesn't have an API. The government portal where you have to manually fill in like 15 fields.

The multi-step approval workflow that takes you an hour because you have to touch seven different tools. Sonic 4.6 is getting meaningfully closer to being able to do those things for you. Box tested this. Their enterprise customers are real world knowledge workers.

80% accuracy on public sector tasks. 78% in healthcare. 94% in retail. These are the sectors with the most bureaucratic, multi-step, legacy software dependent workflows.

And this model is clearing those benchmarks. Now, I get the skepticism. I genuinely do. I also get that people are exhausted by the pace of AI announcements.

Every day there's something new. You've heard this changes everything enough times that it starts to sound like a boy crying wolf. But here's what I'd ask you to actually pay attention to. And it's not the benchmark numbers.

It's the pricing curve. Sonic 4.6 performs like Opus 4.5 used to. It costs one fifth of what Opus costs. This is not a marginal improvement.

This is a structural shift in the economics of deploying the technology at scale. And every time that shift happens, every time capable AI gets dramatically cheaper, the calculus for weighting changes, the ROI case for deployment gets stronger. The number of workflows where AI is a sensible choice expands. If you are a business leader, I want to be blunt with you.

The organizations that are figuring out where AI can be deployed in the operations right now, not because it's a trend, not because it looks good in a board presentation, but because they're genuinely doing the workflow analysis and piloting the tools, those organizations are going to have a compounding advantage over the ones that are still watching. Because this isn't just about efficiency, it's about what happens to the capability gap between an AI native organization and a non-AI native organization over the next 24 months as these models keep getting better and cheaper. That gap is going to widen faster than most executives are currently modeling for. Anthropic also just closed a $30 billion funding round at a $380 billion post-money valuation.

That's more than double their September valuation. To put that into context, that is a company that six months ago was worth roughly $175 billion that's now worth $380 billion. The capital is coming in because the people who understand this technology and the trajectory at the deepest level are willing to pay for a piece of it at those prices. That is Signal.

Now, let me be clear about what I'm not saying. I'm not saying 4.6 is AGI. I'm not saying it replaces human judgment on complex decisions. I'm not saying every job is going away tomorrow.

The model scores 72.5% on computer use and that means it fails more than a quarter of the time on those tasks. The best humans score way higher. There are still real limitations, but I keep coming back to that trajectory. 14.9% to 72.5% in 16 months.

What does that curve look like 16 months from now? What does it look like at 90%, 95%? And at what point does the fail rate on computer use tasks become low enough that businesses treat it like hiring a very competent, very cheap, always available employee who just occasionally needs to be checked? That point is not as far away as it feels.

And here's the thing underneath all of this that I think matters most, and it's not actually about the technology. The technology is accelerating. That part is clear. The question that doesn't get asked enough is what do we do with the people?

Because every efficiency gain from AI is a workflow that used to require human time and now doesn't. And that math compounds. The people whose jobs were those workflows, the junior analysts, the data entry workers, the form fillers, and the first level code reviewers, they are not abstract statistics. They are people with mortgages and kids and skills they've built over years and are being quietly repriced by a benchmark chart.

I don't have a neat answer to that. I think anyone who tells you they have a neat answer is selling you something. What I do know is that the worst thing we can do as individuals is pretend the curve isn't there. Because the curve is there, the numbers are real, the pricing is real, the adaption is real.

The only question is whether you are positioning yourself to be someone who uses these tools or someone who's going to be replaced by someone who does. And if you're watching this, I'm going to assume you are in the first category. So go run Sunit 4.6, play with computer use, throw a big document and that 1 million token context window and see what it does with it. Build something this week that you couldn't have built two months ago.

Not because I'm telling you to. Because the only way to actually understand what's happening here is to have your hands on it. The benchmark numbers are useful, but the moment it actually clicks is when you're sitting there watching it navigate something on your screen that you used to have to do yourself and think, oh, oh, this is real. That moment is available to you today.

More episodes

Browse all episodes →