Julian Goldie breaks down the launch of Google's Gemini 3.1 Flashlight, a model that is 45% faster and significantly cheaper than previous versions. This update represents a major shift in the AI industry, offering doctoral-level reasoning at a fraction of the cost. Learn how to leverage adjustable thinking levels and 1-million-token context windows to scale your business operations.
Full transcript
Google just dropped an AI model that's faster than it's more expensive ones, costs less than a cup of coffee to run a million times, and just launched today. And most people have no idea what that means for them. My name's Julian, and whether you're a business owner, a student, a developer, or someone who just heard the word AI for the first time, I promise you, by the end of this video, you're going to understand exactly what happened today, why it matters, and most importantly, what you need to do about it. Because what Google released this morning on March the 3rd, 2026, is not just another tech update.
This is the AI price war going nuclear, and the people who understand what's happening right now are going to have a massive edge over the people who don't. So let me break it down for you. The model is called Gemini 3.1 Flashlight. It dropped this morning in preview, available to developers right now through Google AI Studio and through Google's enterprise platform, Vertex AI.
Now, before your eyes glaze over at those words, AI Studio, Vertex AI, developer tools, I want you to stay with me because the words do not matter. What matters is what this model does, what it costs, and how that changes everything. Here's the headline. What does that mean in plain English?
Think of it like this. Imagine you hired a worker who could type 249 words per minute. Pretty fast, right? Now, imagine you found a different worker who could type 363 words per minute.
Same work, same quality, but almost half again as fast. Now, imagine that second worker also charged you less money. That's what Google just did. The new model costs 0.25 per million input tokens and $1.5 per million output tokens compared to the old Gemini 2.1 Flashes, 0.30 cents and $2.50.
Let me say that again in plain terms. The new model, the faster one, is also the cheaper one. The old, slower model costs more money per use. The new, faster model costs less.
That is not supposed to happen. When you buy a faster car, you usually pay more. When you hire a more advanced and experienced doctor, you pay more. When you get a bigger apartment, usually you would pay more.
In every part of our lives, better costs more. Google just completely flipped that on its head. Now, you might be sitting there thinking, OK, cool, tech people care about this. Why should I care?
I hear you. Let me give you a simple picture of what this actually means in the real world. Imagine you run a customer service team. You've got 20 people answering emails every day.
Every email takes about five minutes. Your team handles maybe 2,000 emails a month. Now, imagine you could run an AI system that answered those same 2,000 emails faster, more consistently in 100 different languages around the clock. No sick days, no lunch breaks, no vacation requests.
And imagine that the cost to do all of that, 2,000 responses, was now less than a single cup of coffee. That's not science fiction anymore. That's what's on the table right now today with what Google just launched. Google's 3.1 Flashlight is built specifically for high-volume workloads at scale, like content moderation, translation, and processing large numbers of requests, where cost is the main thing holding them back.
When the cost drops this far, the question stops being, can we afford to use AI, and starts being, can we afford not to? Let me give you a little history here, because this is important. A year ago, if you wanted a really smart AI, one that could understand context, follow complex instructions, handle images and audio, reason through problems, you were paying a lot of money for it. The big, powerful models were expensive.
The cheaper models were, frankly, kind of dumb. They'd misunderstand instructions. They'd give you basic, low-quality answers. It was like the difference between hiring a top-shelf consultant and paying for an intern who maybe read some books.
So businesses had to make a choice. Do they pay premium prices for smart AI, or do we use cheap AI and accept lower quality? That trade-off was real. It was a central tension in AI deployment for years.
What Gemini 3.1 Flashlight changes is this. It matches the quality performance of Gemini 2.5 Flash, the bigger, more expensive model, whilst costing less and running faster. Read that sentence slowly. Matches the quality of the bigger model, costs less, and runs faster.
The trade-off is gone. The cheap model is now as good as the expensive one, and it runs faster. This is what Google CEO Sundar Pichai has been saying for years. He kept telling people, AI is going to get faster and cheaper with every generation.
And people kept nodding politely and not quite believing him. He was right. And if this trend continues and there is reason to believe it will, then in one year from now, we'll look at today's models the same way we looked back at the first iPhone. Impressive at the time, comically limited by current standards.
Let me talk about something called thinking levels. And I know that sounds weird, like we're talking about a robot having feet. We're not. But here's what it means and why it's a big deal.
Gemini 3.1 Flashlight comes with adjustable thinking levels, minimal, low, medium, or high, which lets developers choose how much reasoning the model does for each task. Think about it like this. When you're grocery shopping, you don't sit down and make a spreadsheet comparing every brand of pasta. You just grab the one you usually get and move on.
But when you're buying a house, you do sit down. You think carefully. You weigh pros and cons. You maybe even hire someone to help you think through it.
You scale your thinking to the top. Most AI models couldn't do that before. They ever always thought carefully, which made them slow and expensive. Or they never did, which made them fast but dumb for complex tasks.
Gemini 3.1 Flash now lets you dial that up or down depending on what you need. If you need a quick translation job, set it to low thinking, fast, cheap, done. If you need a complex reasoning task, maybe building a user interface or running a simulation or following a complicated set of instructions, dial the thinking up. Let it work through the problem carefully with the same model price structure you controlled the dial.
This is genuinely new. And for businesses, the implications are enormous. Now, I want to be honest with you here. I always try to steel man these skeptics.
I don't think it serves you to just be a cheerleader and tell you everything is perfect and you should be excited and buy something. So let me give you the honest version. Gemini 3.1 Flashlight has a time to first token of 5.18 seconds. On the higher end compared to other models in a similar price tier, where the median is about 1.82 seconds.
In plain English, it's answering a little slower than some competitors. That matters in certain situations. If you're building a real-time chat app where someone is sitting there watching a blinking cursor, that five-second wait before anything happens on screen might feel sluggish. So it's not perfect.
No model is. But here's the thing, speed to first response is one metric. Output speed, how fast the model produces all the words of its answer once it starts, that's a different metric. And on output speed, Gemini 3.1 Flashlight is among the fastest in its class.
So if you're running batch job, if you're processing thousands of documents overnight, if you're building translation pipelines or content moderation systems that work asynchronously, this model is exceptional. It's a question of what you're building, whether the model's specific strengths fit your specific use case. That's a nuance most AI coverage doesn't bother to make. I'm making it now because you deserve the full picture.
Let me tell you about some of the companies already using this kind of technology, because this isn't hypothetical anymore. It's not, here's what AI could theoretically do someday. It's what companies are doing right now. Heygen uses Gemini Flashlight to translate videos into over 180 languages, letting them provide personalized global experiences for their users.
Think about what that means. A company that creates AI avatar videos can now take any video and translate it, not just words, but the whole thing into 180 different languages automatically at scale. The person on screen can appear to be speaking Portuguese or Swahili or Japanese without ever recording in those languages. That's not a party trick.
That's a business transformation. Every creator, every educator, every business with global customers can now reach the entire world of their content in the native language of their audience. Doxhounds turns product demo videos into full documentation by using Flashlight to process long videos and extract thousands of screenshots, transforming footage into comprehensive training materials for AI agents much faster than traditional methods. Let me translate that for the non-developers in the room.
You know how every software company has to write documentation, like the instructions that tell you how to use a product? That used to require technical writers, hours of work, lots of back and forth. Now you record yourself using the software once, you feed it into an AI. The AI watches a video, it takes thousands of screenshots, and it writes the documentation automatically.
What used to take a team of writers two weeks now takes an afternoon. And that's not AI replacing human creativity, that's AI doing the parts of human work that were never creative in the first place. The tedious, systematic, repetitive parts. The someone has to do this work.
And that frees the humans up to do parts that actually require judgment and creativity and relationships and insight. At least that's the optimistic version. And I want to be honest, not everyone who gets freed up from the old task gets a new, better task. That's a real conversation we need to be having, and I'll come back to that later.
Satellite is building a platform that processes satellite data in real time using Gemini flashlight. They actually reported a 45% reduction in latency for critical onboard diagnostics and a 30% decrease in power consumption compared to their previous models. I want to stop here for a second. Satellites, we're using cheap, fast AI models to process data from satellites in orbit.
Think about that for a moment. A few years ago, if you wanted to process satellite data in real time, you needed massive, expensive computing infrastructure. You needed specialized engineers. You needed serious investment.
Today, a relatively small company is doing it with a general purpose AI model that costs fractions of a cent to use. The access gap is closing at a rate that most people don't fully appreciate yet. The tools that used to be available only to big governments and billion dollar corporations are now available to a founder working out of a coffee shop in Bangkok. I'm not exaggerating.
I know people building real businesses right now with AI tools that would have been impossible, literally, technically impossible three years ago. And the question is, are you one of these people or are you watching from the sidelines? Let me give you the competitive picture because this launch doesn't happen in isolation. In Google's own benchmark documentation, they compared Gemini 3.1 Flashlight against GPT-5 Mini, Quad 4.5 Haiku and Grok 4.1 Fast.
These are the other budget AI models from OpenAI, Anthropic and XAI. This is a war between the biggest tech companies on earth and guess who they're fighting over? You. Google wants to be the AI you use.
OpenAI wants to be the AI you use. Anthropic, the company behind Claude, guess what? They want you to be using their AI. Elon Musk's AI wants to be the AI you use.
They are all competing for the same thing, to become the AI structure that the entire world runs on. And right now that competition is making AI better, faster and cheaper for everyone. The winners of this war, honestly, might be the regular people who just get to use progressively more powerful tools for less and less money. But there's a catch.
The people who learn how to use these tools really use them, not just dabble, are going to have a massive advantage over everyone who doesn't. The tools exist. The question is whether you know how to pick them up and build something with them. Let me give you some numbers that should wake you up.
The Gemini 3.1 Flash Lite model scored 34 on the Artificial Analysis Intelligence Index, well above the median of 19 for other models in its price stick. In plain English, it's almost twice as smart as the average model that costs the same amount. When you go shopping for shoes and you find a pair that's twice as good at half the price, you would buy them. The same logic applies here.
Flash Lite scored 86.9% on GPQA Diamond, a benchmark-to-test doctoral-level science knowledge, and 76.8% on MMU Pro, which tests multimodal understanding across different types of information. This is a budget model that can pass doctoral-level science questions at nearly 87%. Now, let me put that in perspective. A few years ago, the world's most powerful AR models, the ones that cost hundreds of thousands of dollars to run, couldn't come close to 87% on a test like this.
Today, the cheapest tier of models are doing it. And the price to run this model is $0.25 per million input tokens and $1.4 per million output tokens. One million tokens is roughly 750,000 words. For $1.5, the model will read and respond to 750,000 words worth of content with near-doctoral-level reasoning.
A human expert who could do that would charge you $1,000 an hour, probably more. I want to take a step back and talk about what's happening at a bigger level now, because the Gemini 3.1 Flash Lite launch is not the story. The story is a trend it represents. Think about where we were two years ago.
Two years ago, if you wanted a capable AI model, you were looking at GPT-4. That was the gold standard, and it was powerful, but it was expensive and slow by today's standards. A year ago, models were dramatically better. And today, the cheapest models available to any developer with a credit card are outperforming the premium models of two years ago on every single metric that matters.
Speed, quality, cost, context window, which is how much information the model can hold in its memory at once. Gemini 3.1 Flash Lite has a one million token context window, one million tokens. Let me put that in perspective. The entire Lord of the Rings trilogy, all three books, is about 500,000 words, or roughly 650,000 tokens.
This model can hold more than the entire Lord of the Rings trilogy in its working memory, and then reason about it, and then answer your questions about it for fractions of a cent. A year ago, that kind of context window was only available on the most expensive, most powerful models. Today, it's table stakes on the budget tier. The acceleration is real, and it is not slowing down.
Let me tell you about something that happened this week that nobody's covering enough. The timing of this launch had a competitive element. Claude Anthropic's AI was still dealing with issues after an outage on March 2nd, whilst Google was rolling out Flash Lite in preview. This is not a coincidence of timing.
These companies watch each other constantly. When one AI goes down, another one's product team sees an opportunity, and this is a war. A polite tech company war, fought with blog posts and benchmark charts and press releases, but a war all the same, and what that means for you is this. Every week, sometimes every day, there is new ammunition being added to the arsenal of tools you have access to.
The battlefield is moving so fast that people who learned AI six months ago are already partially out of date, which is why being part of a community that tracks this stuff, that filters the noise and tells you what actually matters, is not a luxury anymore. It's a competitive necessity. And let me talk to you about what I see happening over the next 12 months. Right now, we're in what I'd call the infrastructure phase of AI.
Companies like Google are racing to build the cheapest, fastest, most capable foundations. They're fighting each other for market share, and every time they fight, prices drop and capabilities go up for everyone. What comes next is the application layer. Regular businesses, not tech giants, not Silicon Valley startups, but the small business down your street, the law firm, the accountancy practice, the marketing agency, the restaurant chain, they're all starting to plug these cheap, fast models into their workflows.
Not because they want to, not because they're excited about AI, because the ones who don't are going to be left behind by the ones who do. And when your competitor can produce 10 times the content in half the time at a 10th of the cost, your only choice is either to match them or lose. That's not a hypothesis, that's already happening in certain industries, and the wave is coming for everyone else. The question is whether you're going to be standing on a surfboard when it arrives or getting knocked over by it.
Now, I want to talk to some specific people, because different people in this audience need to think about this differently. If you're a developer, you need to be testing Gemini 3.1 Flash Lite right now. Today, it's in preview, it's available via the Gemini API in Google AI Studio, and the pricing makes it viable for high volume production workloads in a way that previous models weren't. The configurable thinking levels means you can optimize the speed on simple tasks and dial up reasoning when you need it.
The model is specifically built for translation, content moderation, generating user interfaces, and creating simulations. If any of those are in your stack, test this today. If you're a business owner who doesn't code, the specific model name doesn't matter to you. What matters is this, the cost of running AI in your business just dropped again.
The quality went up, the speed went up. This means the tools that your developers build for you or the access through consumer products that run on models like this in the background just got better. If you haven't started automating parts of your business yet, the excuse of it's too expensive or it's not good enough yet gets harder to make with every launch like this one. If you're a knowledge worker, so for example, a writer, a marketer, a researcher, a consultant, your job isn't to run AI models.
Your job is to produce insight, strategy, words, ideas. What AI models like this one do is compress the time you spend on the mechanical parts of that work, the research, the first drafts, the translation, the summarization. So you can spend more time on the parts that require your actual judgment. The people who learn to use these tools as a multiplier for their own thinking are going to be dramatically more productive than the people who don't.
And in any competitive field, dramatically more productive people usually means dramatically more valuable. If you're a student, you're entering a job market where AI fluency is going to be as basic a requirement as computer literacy was for your parents' generation. The time to learn this stuff is not after you graduate, it's now. The tools are cheap, the learning curve is not that steep.
And the advantage of starting early compounds over time in a way that's very hard to catch up later. If you're a manager or an executive, your competition is looking at this launch today and thinking about what it means for their cost structure and their team size. Your job is to have that same conversation, but faster and more thoughtfully. The companies that are going to win over the next decade are the ones whose leadership actually understands what AI can and can't do.
Not at a service level, not I've heard of chat GPT, but at a real working level. That means getting your hands dirty. It means experimenting. It means being curious.
Now, let me talk about the thing that I know some of you are worried about, but haven't said out loud, which is jobs. And I know, I know, you've heard this before. AI is coming for jobs. Every year, someone says it, and then, you know, people still have their jobs.
But I want to be honest with you. I think the scale of change coming in the next few years is different from anything we've ever seen before. Not because AI is going to wake up and want to fire everyone, but because economics is economics. Google built this model specifically for high volume tasks where cost is a priority.
So translation, content, moderation, classification. Those are jobs, real jobs that real people have. So translators, content moderators, and data labelers. Yes, the first wave of AI automation is already going through those fields.
But here's what I keep saying, and I believe it. The response is not to pretend it isn't happening. It's not to fight the tool. It's to be the person who can use the tool rather than the person who the tool replaces.
The textile workers who learned to operate looms didn't all lose their jobs. The ones who adapted, who became machine operators, who moved up the value chain, they did fine. The ones who refused to engage, who hoped the change was just passing by, they didn't go so well. And we're at that same moment again, just a lot faster this time.
And this is the part where I want to talk to you about what it actually takes to be the person who uses this tool. Because I talk to a lot of people who say, I know I need to learn AI, I just don't know where to start. And I get that really, I do. The space moves fast.
There's a hundred tools. Every week there's a new model. The jargon is confusing. And most of the content out there either assumes you're already a software engineer or talks to you like you're a five-year-old who just learned what a computer is.
There's no middle ground for the business person who's smart, capable, and just trying to understand what they need to know to stay competitive. And that's what we built the AI Profit Boarding for. The community for people who are serious about using AI in their businesses and their careers without needing to become a developer, without spending months on YouTube watching tutorials that are already out of date. We track every major AI development, launches like Gemini 3.1 Flashlight today, and we filter it down to what actually matters for people who are building things, running teams, and trying to stay ahead.
We cover practical AI automation, real workflows, real case studies from real businesses, not theory, not hype, not speculation, the actual stuff you can use. If today's video made you feel like there's a lot happening that you might be missing out on, that feeling is correct. And the AI Profit Boarding is where people go to make sure they're not left behind. The link is in the comments and description, or you can go to the AIProfitBoarding.com.
Come check it out. Let me come back to the Gemini 3.1 Flashlight story for a minute. Because there's one more technical detail I want you to understand, and it's the one I think that matters most. It's called reasoning.
Gemini 3.1 Flashlight is built on top of Gemini 3 Pro, the most powerful model in Google's lineup, and then compressed and optimized for speed and cost. Think about what that means. The brain that powers Gemini 3.1 Flashlight is fundamentally the same architecture as Google's most powerful AI. It's like the difference between a sports car and a commuter car built by the same engineering team.
The commuter car isn't a luxury vehicle, but it shares the same fundamental engineering innovations, the same safety technology, the same fuel efficiency breakthroughs, just optimized for a different purpose. When Google gets better at building smart AI, that intelligence flows down to the cheaper models too, which means every time DeepMind, Google's AI research lab, makes a fundamental breakthrough in how AI reasons, how it understands images and audio, how it follows complex instructions, that gets built into the next generation of models at every tier. The cheap models of tomorrow will be smarter than the expensive models of today. We know this.
We can already measure it. The curve is going up and to the right on every metric. Quality up, cost down, speed up. That trajectory doesn't stop unless something fundamentally breaks.
And right now, nothing fundamental is breaking. Gemini 3.1 Flashlight ships with configurable thinking levels in AI Studio and Vertex AI, giving developers direct control over how much reasoning the model applies to each task. Lean for bulk translation or content moderation and higher reasoning for UI generation, simulations or complex instructions. And I want to come back to this one more time because I think it represents something that doesn't get talked about.
We've been trained to think about AI as one thing. You use it, you get an answer that's simple. But what's actually happening right now is that AI is becoming a spectrum. You have a choice about how hard it thinks, how much time it takes, how much it costs, and you are in control of that dial.
That is a fundamentally different relationship with AI than what most people have experienced right up to now. It's not like ask ChatGPT a question and get an answer. It's basically design a system where AI's thinking is a resource that you allocate wisely, kind of like electricity or time. And the businesses that get good at managing that allocation are going to build things that are simultaneously faster, smarter, and cheaper than businesses that treat AI as a monolithic yes or no tool.
Understanding that is genuinely valuable. And it's the kind of understanding that takes a little bit of time to build but compounds enormously once you have it. So let's just stack up what happened, right? So it's not just about Gemini 3.1 Flashlight.
This week, the week before this, Claude had a major service outage that affected businesses around the world. That's a reminder that relying on a single AI provider is a risk. The businesses that know how to use multiple models, how to switch between them when one goes down, how to evaluate which model for each task, those businesses are more resilient. A few weeks before that, Gemini 3.1 Pro dropped, the big model, and before that, Google had already released Gemini 3 Flash.
The cadence of releases is accelerating. It used to be a major model releases like every few months. Now it's closer to every few weeks, and this is not slowing down. The people who are trying to wait until things settle down before I learn about AI, I have bad news for you.
Things are not going to settle down. This is the new normal, fast, constant improvement, new models every few weeks or every week, and the price is dropping constantly. The only sensible strategy is to build a piece of staying current, not spending 10 hours a day on it, not becoming an AI researcher, just having a reliable system for knowing what matters and what doesn't for your specific situation, your job, and your business. So let me give you the practical takeaways.
Three things you can do right now today based on this launch. Number one, if you have any kind of AI workflow in your business, any tool, any automation, any system that uses AI in the background, find out which model is running them. Ask your developer, look at the settings. You might be paying more for an expensive model when a cheaper, faster one would actually serve your needs just as well or better.
This launch is a good excuse to audit what you're using and why. Number two, if you've never actually used AI to automate something in your business, pick one repetitive task this week and find an AI tool to help with it, not a hypothetical task, a real one, something that actually takes you to translation, summarization, drafting responses to a common type of email, writing product descriptions, processing receipts, one task this week, just to start building the muscle. Number three, follow what's happening, not obsessively, but regularly, because the gap between people who are informed and people who aren't is widening every month. You don't need to be an AI researcher.
You don't need to read every technical paper, but you do need someone in your corner who can translate the noise into signal. That's what this channel is for. Hit subscribe if you haven't. And if you want something deeper, a real community, real frameworks, real case studies from people using AI to run more efficient businesses right now, come check out the AI Profit Boardroom.
Link is below or go to the AIProfitBoardroom.com. Here's what I want you to leave you with. So every generation has a moment when the rules of the game change. Grandparents had the internet.
Your parents had smartphones. And somewhere in there, people understand what was coming and got to ride the wave. And people who didn't got bowled over by it. We are in one of those moments right now, maybe the biggest one.
The model that just dropped today, Gemini 3.1 flashlight, is faster and cheaper and smarter than models that were considered cutting edge just one year ago. And 12 months from now, today's cutting edge will look like yesterday's news. The acceleration is real. And the thing that separates the people who benefit from this moment from the people who get left behind from it is not intelligence.
It's not technical skill. It's not even money. It's awareness and willingness to learn. So are you aware of what's happening?
Well, you are today. I have to listen to this. Are you willing to learn? That's up to you.
But if you've watched this far, I have a feeling you're the kind of person who wants to know, who wants to stay ahead, who takes his stuff seriously and good because you should. The world is changing faster than most people realize. The people who are paying attention, who are building understanding month by month, who are showing up consistently, they're going to look very, very smart in five years. Be one of those people and I'll see you in the next one.
More episodes