AI News Today
← All episodes
Episode 1 · February 19, 2026 · 19:50

Google Just Built the World’s Smartest AI. Here’s What It Actually Did

Julian Goldie explains how the most effective use of AI requires deep domain expertise. Learn how to treat models like Deep Think as junior researchers to unlock workflows that were previously impossible.

Full transcript

An AI just solved 18 problems that stumped the entire human scientific community. Not in some controlled lab setting, not on a benchmark designed to make it look good, I mean actual, open, unresolved research problems in mathematics, physics, computer science, and economics that professional researchers have been stuck on for years. And one of them was a conjecture that had been sitting there since 2015. 11 years.

Some of the smartest mathematicians on the planet had tried to crack it. They couldn't. And then Google's Gemini DeepThink came in, thought about it, and disproved it in a single run. That's where we are right now.

And I want to spend today's episode really unpacking what just happened here because I think a lot of people are going to miss it. They're going to see the headline, they're going to see, oh another AI benchmark story, and then just scroll past. But I keep saying this, the moment you stop paying attention to these individual stories is the moment the gap between what you understand and what's actually happening starts growing so fast you can never catch up. So let's talk about Gemini DeepThink.

What it actually is, what it actually did, why the benchmarks are simultaneously impressive and incomplete, what the skeptics get right, and what all of this means for the world you're going to be working in three years from now. Let's get straight into it. First, let me back up and give you some context because Google has been building toward this for a while and the timeline matters. In July 2025, an earlier version of Gemini DeepThink achieved gold medal standard at the International Mathematics Olympiad.

If you don't know what that is, it's the most prestigious math competition in the world for high school students. The problems aren't just hard, they require creative proof construction. You have to invent a logical argument from scratch that nobody's told you how to make. These are the problems that identified the mathematical geniuses of each generation, and Gemini hit gold.

That's nothing. That's not nothing. That is genuinely extraordinary. Then in November 2025, Google released Gemini 3 and alongside it they announced Google DeepThink mode.

A specialized reasoning layer built on top of the model designed not for speed or breadth, but for depth. System 2 thinking in the language of behavioral economics. The slow, deliberate, check your work kind of cognition, not the quick answer reflex. Then on February 12th of this year, last week, they dropped the major upgrade and the numbers are staggering.

Let me just read some of these to you straight. Humanities last exam, 48.4% without external tools. Now, if you hadn't heard of Humanities last exam, here's what you need to know. It was built specifically because traditional AI benchmarks were getting saturated.

Models were scoring so high on standard tests that the tests stopped being useful for measuring progress. So a group of researchers built a new one, pulling from the absolute frontier of human expertise across every domain. Graduate level physics, advanced economics theory, obscure mathematics, the kind of questions where even domain experts might have to think hard. And the baseline expectation was that these questions were basically impossible for AI.

Early frontier models scored about 2% of this test, 2%. Gemini 3 deep think just scored 48.4%. That's not a trend line. That's a phase change.

Then there's ARC, AGI2. This one's particularly interesting because it was specifically designed to test reasoning that can't be gamed by memorization. It uses novel visual answers and puzzles, abstract patterns that you've never seen before by construction. The whole point is that you can't prepare for it by reading more training data.

You either can figure out the pattern from scratch or you can't. Humans average 60% on ARC, AGI2 tests. Previous AR models were often struggling to break 20%. Gemini 3 deep think scored 84.6% verified by the ARC prize foundation.

I want you to sit with that for a second. On a test specifically designed to measure genuine reasoning, not memorization, not pattern matching against things you've seen before, this model is outperforming the human average by almost 25 percentage points. And on coding, they measured it on Codeforces, which is the gold standard platform for competitive programming. The ranking system goes pupil, specialist, expert, candidate master, master, international master, grandmaster, international grandmaster, and then right at the very top, legendary grandmaster.

That's a level of maybe 300 or 400 people globally out of millions of programmers. Gemini 3 deep think has a 353,455 Elo on Codeforces. That is legendary grandmaster level. Physics and chemistry olympiads often gold medal on the written sections of both.

And then there's the 18 unsolved research problems, which is the one that really caught my attention. Let me tell you about what happened with that decade old conjecture, because it illustrates something deeper about what's going on here. In 2015, a group of researchers published a theory paper about online submodular optimization. If that phrase means nothing to you, here's what it's actually about.

Imagine you're managing a stream of data items flowing in one at a time, and you have to make decisions about what to keep and what to throw away. The conjecture proposed what seemed like an obvious rule, that making a copy of an arriving item is always less valuable than keeping the original. It sounds intuitive. It sounds almost self-evidently true.

Researchers spent about 10 years trying to prove it. Brilliant people, people who think about these problems professionally every single day, they couldn't prove it. And here's why, because it wasn't true. Gemini DeepThink disproved it in a single run.

It constructed a counterexample with just three items, a specific combination that showed the conjecture was wrong. And crucially, human mathematicians had missed this because they kept following their intuition about how data streams should work. Gemini doesn't have that intuition. It explored systematically, exhaustively, without the cognitive bias that comes from having worked in a field for so long that certain assumptions stop feeling like assumptions.

In physics, DeepThink tackled gravitational radiation calculations from cosmic strings. That's a problem in general relativity. The physics of space bending around massive objects. The specific challenge involved integrals with singularities, which is math speak for equations that blow up and become infinite at certain points.

Traditional approaches kept getting stuck. Gemini found a path through Gegenbauer polynomials, a mathematical tool from a completely different domain, and collapsed infinite series into clean solvable expressions. And that's a move. That's a thing that's actually new.

Not just being fast at hard problems, being able to reach across disciplines and grab tools from entirely unrelated branches of mathematics and apply them in novel contexts. At Rutgers University, mathematician Lisa Carbone used DeepThink to review a technical paper. The paper had already passed human peer review. Gemini found a subtle logical flaw that human reviewers had missed.

At Duke, the Wang Lab used it to optimize fabrication for crystal growth, designing recipes for thin film structures larger than 100 micrometers. These are actual experiments with actual materials. Now, here's where I want to steel man these skeptics, because this is important. There's a really good analysis from a team that put Alephia, the research agent built on top of DeepThink, through a systematic evaluation.

They threw 700 open problems from the Erdos conjecture database at it. Paul Erdos was one of the most prolific mathematicians in history. He left behind a collection of hard open problems that became a kind of mountain range for subsequent generations of mathematicians to climb. Of the 200 problems that were clearly evaluable, 137 were fundamentally wrong.

That's 68.5%. Almost 7 out of 10 attempts were mathematically incorrect. Of the 63 that were correct, only 13 actually answered the question that was asked. The other 50 are technically valid maths, but the model had quietly reinterpreted the question to make it easier.

The researchers called this specification gaming. The AI systematically rewrites a problem into something it can solve, then solves that instead, and presents it as if it solved the original. That's real. That's a genuine limitation.

That's not hype busting. That's just honest. Demis Hassabis himself put it well. He said AI right now is jagged intelligence.

Some things it's remarkable at, and other things it's surprisingly bad at. The profile is uneven in ways that can fool you if you're not paying attention. So the honest picture is this. Gemini DeepThink is extraordinary in specific situations, when pointed at the right kind of problem, with the right kind of human direction.

It can produce results that humans couldn't produce without it. But when you release it autonomously on a wide range of hard problems, its error rate is high, and its tendency to game the specification is real. And that tension right there between the genuine breakthroughs and the genuine limitations, that's where the interesting question lives. Because here's what I keep coming back to.

The gap between extraordinary in specific situations with good human direction, and broadly autonomous and reliable, is a gap that matters. And that gap has been closing faster than almost anyone predicted. Let me give you a sense of the timeline, because I think people forget how fast this has moved. A year ago, the state of the art on humanities last exam was in the single digits, 2%, 4%, somewhere in that range.

Last month it crossed 48%. That's not a linear improvement. That's not even an exponential improvement. That's something closer to a series of step changes, each one bigger than the last.

The International Math Olympiad went from AI can't do this at all, to gold medal performance in roughly 2 years. ARC, AGI too, went from AI can't generalise to new patterns, to AI beats the human average in less than 18 months. Every time someone built a benchmark that was supposed to be too hard for AI to game through pattern matching, AI figured out a way to actually reason through it. And I'm not saying we're at AGI.

I'm specifically not saying that. What I am saying is that the rate at which the frontier is moving is unlike anything in the previous history of technology. And the question isn't whether this matters. The question is, when does it matter for you, for your job, for your field, for the way you create value in the world?

Let me make this concrete because I think the science stories can feel abstract to a lot of people. The researcher at Rutgers, Lisa Carbone, she found a logical error in a peer-reviewed paper. A paper that had already been read and approved by expert human reviewers. That's not just an impressive party trick.

That changes the economics of peer review. If you can run every submitted paper for an AI system that catches errors with that level of sophistication, not just typos, not subtle logical inconsistencies, well you compress the time between discovery and validation. You catch mistakes earlier. You accelerate the whole cycle.

The Duke University, Crystal Growth, designing fabrication recipes for thin film structures. That's material science. That's the domain that produces semiconductors, batteries, solar panels. If you can compress the experimental design cycle, you compress the timeline for next generation hardware.

The chips that run the next generation of AI. The batteries that power the next generation of electric vehicles. The materials that go into everything. And the sketch to 3D print capability where Anupam Pathak at Google's Platforms and Devices Division was using DeepThink to turn drawings into printable files.

That's design iteration. That's prototype cycles. Right now it might take an engineer to go from a concept sketch to something you can physically hold in your hand. If that becomes hours and then minutes, you don't just go faster.

You try more things. You explore more options. You discover more possibilities. And this is the compounding effect nobody talks about enough.

It's not just that AI makes existing workflows faster, it's that faster workflows enable exploration that wasn't economically rational before. When something costs 10 hours, you're selective about how many variations you try. When it costs 10 minutes, you try everything. Now let me zoom out even further.

The conversation around AI capabilities has been framed for most of the last few years in terms of language, writing, summarizing, generating text. These are real applications and they're useful. But they're not the thing that changes the trajectory of civilization. What changes the trajectory of civilization is scientific discovery.

The pace at which humans can move from question to answer to application has always been the fundamental bottleneck on progress. It's why science that would have taken Newton a lifetime took Einstein a decade. It's why processes that took a generation to develop in one century took years in the next each time we find a way to compress the discovery cycle. And the benefits compound, not just for science, but for everything that depends on science.

So that could be medicine, energy, computing, materials, agriculture. Deep think is not the end state of this. It's actually the early signal. And right now it works best as a collaborator, a very capable junior researcher that needs expert human direction that sometimes gets things wrong in quite important ways that requires careful checking.

The researchers working with it describe it best. Treat it like a capable but error-prone junior researcher, not an oracle. But junior researchers become senior researchers. Tools that need expert direction become tools that need less direction.

And the trajectory is very clear. And there's one more thing I want to say before I get to what you should actually do with this information. This is happening at the same time that these models are becoming more broadly available. GMLi3 DeepThink is available right now to Google AI Ultra subscribers.

That's not free, but it's also not inaccessible. They're opening API access to researchers and enterprises. And this isn't locked up in a lab somewhere. This is actually deploying, which means the question of who gets to these capabilities and when and at what cost starts to matter enormously.

The institutions that figure out how to work effectively with these systems, how to direct them well, how to check their outputs correctly, how to build workflows that leverage their strengths whilst guarding against their failure modes. Those institutions are going to have a significant advantage over the ones that don't. That's not hype, that's just organizational dynamics playing out in a new context. So let me give you something concrete.

Let me tell you what I think you should do with this. If you're a researcher or an academic, this is the moment to start seriously experimenting. Not with the expectation these tools will do your research for you. They won't, not yet anyway.

But as a capable thinking partner, try feeding it papers in your domain and asking it to find logical inconsistencies. Try using it to search for connections between your problem and methods from other fields. The balance prompting technique is particularly useful. Instead of asking it to prove something, ask it to either prove or disprove it.

That reduces the confirmation bias, where the model just tries to support whatever direction the prompt implies. And the context de-identification trick is worth knowing. If you're working on a well-known open problem, strip the context that identifies it as unsolved. Otherwise, the model sometimes refuses to engage because it knows it's hard.

If you're an engineer or developer, the coding capabilities here are extraordinary. A 3,455 Elo on code forces is not a number that's meaningful for most everyday coding tasks. But what it tells you is the underlying reasoning architecture is very strong. If you're working on complex algorithmic problems, the kind where the solution path isn't obvious, DeepThink mode is worth testing seriously.

And the API access being open to enterprises is significant. This is a moment to think about where complex reasoning fits into your stack. If you're a knowledge worker who doesn't think of yourself as technical, I want you to pay attention to what's happening in the research domain because it's going to migrate. The same capability that catches logical errors in mathematical papers will catch logical errors in legal contracts.

The same capability that finds connections across disparate mathematical domains will find connections across disparate business domains. The timeline for that migration is unclear, but the direction is not. If you're a leader running a team, running a department, running a company, the strategic question is not will it affect us, it's how do we build the organizational capability to direct these systems effectively. That's a different skill set than using AI as an autocomplete.

It requires people who can formulate good problems, check AI outputs carefully, combine AI capabilities with domain expertise. That combination human domain expertise plus AI reasoning capability is where the value is going to concentrate, figure out how to build it. And if you're just a person trying to understand what's happening in the world, I want to leave you with something that I think about a lot. We are in a moment where tools didn't exist a year ago are doing things that experts told us wouldn't be possible for decades.

We are in a moment where the pace of change is so fast that even people who follow this closely every single day sometimes get surprised. That's genuinely disorienting, and I get that. I also know that the natural human response to disorientation is to retreat into skepticism, to find the caveats. And yes, the caveats are real, I already gave them to you, right?

68.5% error rate on the Erdos' problems, specification gaming, jagged intelligence, and use those caveats as a reason not to engage. That's what some people do, but that's a wrong move here. So the direction is clear, the pace is faster than expected, and the gap between people who are developing a real working relationship with these tools and people who are watching from the outside is going to compound over time. And AI just disproved an 11-year-old conjecture that human mathematicians could not crack.

It caught a logical error the past human peer review. It solved physics problems that were stuck because human intuition kept getting in the way. None of that means the machines are taken over. None of that means your expertise doesn't matter.

In fact, the evidence suggests kind of the opposite, that the most effective use of these systems requires exactly the kind of deep domain expertise that takes years to build. What it does mean is that the way you apply that expertise is changing, and the people who figure out how to change with it, who treat these tools like junior researchers and learn to direct them well, are going to do things that weren't possible for anyone before. That's the world we are walking into. And I think honestly, once you get past the disorientation, there's something genuinely exciting about that.

All right, that's the episode. Let me know in the comments what you think, whether you're using DeepThink or any of these reasoning models in your work. I want to hear what's actually landing and what's not. That signal is genuinely useful for me as I figure out what to cover next.

I'll see you in the next one.

More episodes

Browse all episodes →