AI News Today
← All episodes
Episode 1 · May 19, 2026 · 13:41

NEW Google Gemini AI Agent is INSANE!

Google Gemini AI Leaks Before I/O 2026: Desktop Agent, Spark Mode, Stream-to-Cursor, Veo Video + Gemini 3.5/3.2Ahead of Google I/O 2026, leaked reports suggest major upgrades to Gemini, including a new desktop app with standard chat plus “Spark Mode,” a local agentic workspace that connects to folders, reads files, runs scripts, organizes documents, and syncs with Google Drive. The script highlights “Stream to Cursor,” a floating overlay that lets Gemini read the context of whatever window your cursor is on, alongside a persistent voice interface (“Gemini Live”) and reported local “skills” support for attaching custom scripts to automate repeatable tasks. Leaks also mention Veo video generation integrated into the Gemini desktop experience, plus new models reportedly Gemini 3.5 Pro and Gemini 3.2 Flash, with testers emphasizing speed and workflow impact. The video ties these leaks to building a connected system via an “Agent OS” dashboard integrating Gemini, Claude, Hermes, OpenClaude, and NotebookLM for business automation.00:00 Massive Gemini Leaks01:24 Spark Mode Desktop Agent02:36 Stream to Cursor Overlay04:00 Build a Connected System05:46 Local Skills Support06:28 Veo Video Inside Gemini07:27 Gemini Live Voice Agent08:02 Gemini 3.5 vs 3.2 Flash10:03 Why This Matters for Business12:13 Agent Layer Arms Race13:04 Boardroom Walkthrough and Wrap

Full transcript

Google's biggest AI leaks just dropped ahead of Google I.O. 2026 and if what's being reported is accurate, Gemini is about to become the most powerful AI agent most people have ever used. We're talking about a desktop agent that connects directly to your files, a feature that watches your cursor and understands what you're already working on, a voice overlay that sits on your screen all day, video generation baked right in, and a brand new version reportedly Gemini 3.5 Pro and Gemini 3.2 Flash. The early testers are saying this makes Claude Code an OpenAI codex look, and this is a direct quote from a software architect, nerfed.

Now none of this is confirmed by Google yet, these are leaks ahead of Tuesday's official announcements, but the direction is clear and if you understand what's coming before it lands, you'll be able to plug it straight into your workflows the moment it goes live, whilst everyone else is still figuring out what it actually does. So in this video I'm going to break down every major Gemini leak, what it actually means for your business, and show you the one setup that lets you plug Gemini, Claude, Hermes, OpenClaude, Notebook, LM, all into a single operating system that runs your lead gen content and research without you doing it manually. Stick with me to the end because the last part is the most important thing that I'll show you today. Let's get into it.

So let's go through the leaks one by one. The biggest one is the Gemini desktop app. According to what's being reported ahead of Google I.O., the app is being split into two separate modes. The first is a regular chat mode, so nothing new there.

The second is what's being called Spark mode, and Spark, if the leak is accurate, is a full local agentic workspace that runs tasks on your actual computer. What's reportedly inside Spark is a direct connection to folders on your machine, the ability to read files, run scripts, organize documents, and sync automatically with Google Drive. The agent isn't sitting in a browser tab waiting for you to copy things to it. It's already connected to what you're already working on.

It knows what's there and it can act on it. Here's a simple example of what that looks like in practice. You've got a folder of client proposals. You ask Spark to find the ones with no follow-up in the last 30 days and write a re-engagement message for each one.

It reads a folder, checks the files, and goes off and does it. What used to take an afternoon takes minutes. Again, this is a leak, not a confirmed feature, but the detail in what's being reported is specific enough to take seriously. Now here's the feature that actually stopped me.

It's being called stream-to-cursor, and the concept is this. Gemini watches where your cursor is on screen. Whatever app or window your cursor is hovering over, Gemini can read the context of that window. There's a floating overlay running quietly in the background.

You can share that screen, that window, or even a camera feed without switching apps or copying anything. So you're in your email client, cursor on a thread. You want a quick summary. You don't copy it.

You don't switch tabs according to leak. Gemini already sees what you're looking at. You just ask it. You're inside your CRM, for example, a project management tool.

You want to know which tasks are overdue. You ask. It reads the window. It answers.

If stream-to-cursor ships close to how it's being described, the friction in your daily workflow basically disappears. You stop shuttling information between tools and just get to work. The leak also mentions you'd be able to switch between Gemini 3.2 Flash and Gemini 3.5 Pro directly from that floating overlay. Flash for speed, Pro for depth, and right there mid-task without closing anything.

That matters more than it sounds. Different tasks need different models, and having that switch in your workflow rather than in a settings menu is genuinely useful. And also the fact that Gemini 3.5 and also Gemini 3.2 Flash are coming out is exciting in itself. Now I want to stop here for a second because I know a lot of you are watching and already trying to run AI in your business.

Maybe you've been using Claude. Maybe you've played with Hermes, OpenCLAW, or Gemini, and you've probably had the experience of hearing about a new tool every week, not knowing what to actually do with it. Here's what I keep saying. The people who get real results aren't the ones using the best individual tool.

They're the ones with a connected system, something they can slot into any new tool. For example, if Gemini drops, they can slot it in. If Hermes updates, they slot it in. Claude adds a new feature, just add it in, right?

And it's immediately useful because the system is already set up and running. Inside the AI Profit Boarding, we've built out exactly that. It's called the AgentOS, a local AI operating system that runs on your computer with a dashboard that connects Claude, Hermes, Gemini, OpenCLAW, NotebookLM, and more all in one place. Right now today, for example, you can plug NotebookLM into it and have it automatically generate podcasts, videos, infographics, and slide decks from one single click.

No manual clicking between tools. You can run Hermes Agent Swarms from the same dashboard, assigning tasks that go off and run in the background whilst you do other things. You can plug in your SEO workflows, your content studio, your memory system, all from one control panel. At the moment, the official Agent from Gemini features drop on Tuesday.

We'll have a step-by-step tutorial live inside the boardroom showing you exactly how to add Spark Mode and the new Gemini models to that same AgentOS. You're not starting from scratch. You're adding a powerful new tool to a system that's already working for you. We have four weekly live coaching calls every week where we go deep on exactly that, how to connect Gemini, Claude, Hermes, and OpenCLAW together so they're generating leads and sending you hours, not just sitting in separate tabs, right?

We also have 3,000 business owners inside. They're already running this stuff. Link in the comments description or go to theaiprofitboarding.com. Back to the leaks.

There's something being reported called local skills support. According to what's leaked, skills are custom scripts or capability folders you attach directly to the agent. Now, you can already do that, right? You can build or download a script that handles a specific type of task, say formatting a weekly client report or pulling data from a tool and turning it into a summary.

But now, you can attach it to Gemini as a skill, and from then on, you just ask. The skill runs, the output appears, right? This is where you can really customize your workflow. So the difference between using Gemini as a sort of general assistant and using it as a purpose-built operator in your specific business is the skills that you build, right?

You could have three or four well-built skills and you're saving hours every week on things that currently eat up your time. Now, let's talk about the Gemini Omni angle. The leaks refer to it internally as Veo for Omni, which, if accurate, means Google is baking Veo for video generation directly into the Gemini desktop experience. What that would look like in practice, for example, is if you're working on a script or a brief inside Spark mode.

Let's say, for example, you want a short video, you can ask Gemini to generate it in the same window. There's no jump into a separate tool, no export, no import. The content stays in one place. For anyone producing content, for clients, for social media, etc., it's a meaningful change to how the production pipeline works.

You're not just stitching together five different tools to go from idea to output. It's all in one place, all with an agent. And also, I can imagine that would link to your skills that you set up, and also, it would link to the context on your screen if you're trying to create video about something that you're currently working on. Now, the direction this points to is clear, and it's consistent with where every major AI company is actually heading right now.

There's also Gemini Live in the leaks, which is a persistent voice overlay. So, according to what's been reported, it's described as still a work in progress internally. It's expected to ship in early form, if it ships at all on Tuesday. But the concept is a voice interface to the agent that stays on your screen all day, right?

It's kind of like having Jarvis on your computer. So, you don't open it, you don't close it, you just talk to it. For example, like, summarize a client email thread from yesterday, or what tasks are due today? Draft a follow-up for the proposal I sent on Monday, right?

You say it, the agent handles it, and now we're going to talk about Gemini 3.5 and 3.2 specifically. Now, Eden Kalakanu, probably totally mispronounced that, a software architect, tested what he described as Gemini 3.5 Flash inside a platform called Antigravity. His words are, it makes Claude Code and OpenAI Codex look nerfed. He said it built a physics-enabled voxel castle in under a minute.

Now, I want to be honest about what that actually means. One person's test on one platform is not a benchmark. There are people in the same conversation, Fred, who also said they've never met anyone who uses Gemini for coding, which shows how wide the gap is between where Gemini is heading and where people's current experience of it currently sits. What's being more consistently reported across multiple testers is the speed.

Gemini 3.2 Flash is apparently very fast, and for agents running multi-step workflows, speed matters in a way that's hard to overstate. When you're running a 10-step automated workflow, latency compounds. A fast model that's good enough beats a more capable model that's slow. That's why Flash matters, even when pro tends to get more attention.

Reports also mention Gemini 3.2 Flash already available on Antigravity. I tested it out this morning. I checked it and logged in. I couldn't see it at all, but it's interesting to see, and I think some people have it released.

I actually saw that some people haven't got it in the pro plan, whereas on the free plan, more and more people are seeing 3.2 Flash inside Antigravity. Now, according to what's also being discussed, next week we're reportedly getting Sonic 5 from Anthropic, according to some rumors, and GPT 5.6 from OpenAI. Sam Altman actually just tweeted about that today, and Gemini 3.5 from Google all in the same window. These companies are watching each other closely, right?

The releases are clustering. The competition is forcing the pace up, and the people who benefit most from that are the ones who actually plug these tools into their business. So what does all of this actually mean for someone running a business or doing client work like you? Well, if the leaks are accurate, the Gemini desktop agent is going to be one of the most capable local agents available to anyone with a Google account.

Stream to cursor alone would save hours per week for anyone using multiple tools daily. Spark mode would be immediately useful for anyone managing files, client documents, or content pipelines. Skills is where the power users will pull ahead, but the pattern I keep seeing is this. A new tool drops, people try it for a few days, it doesn't connect to anything they're already doing, doesn't sit into their current workflows, it just sits in a tab and so they stop using it.

But the people who actually get results have a system, right? Not a collection of tools, a system where everything is inside one nice dashboard. So something they can slot new tools into. And the moment each new tool drops, it's already working because they already have a system and a mission control running.

I also want to be straight about the Gemini reputation question. That post I mentioned about nobody using Gemini for coding, that's real, right? That came from the same discussion thread around these very leaks. So Gemini has had a rough reputation, especially for coding work, and I think that's a fair read of where it's been.

I don't really use it to build out websites. I would rather use something like Claude or Hermes Agent to help me instead. But three things are shifting based on what's been leaked. First, the speed on 3.2 Flash is being independently reported across multiple testers, not just one.

Second, the desktop app architecture is a completely different play, right? They already have a desktop app, but it doesn't really have an agent inside it. So this could totally change everything for them. And third, the Veya 4 integration, if it ships as leaked, would be a really powerful way to generate videos and a whole new step up from what we previously had.

So the honest read is don't write this off based on past experience with Gemini. Tuesday's announcements will tell us how much of what's leaked actually made it to launch, but the direction is meaningful. I do remember Gemini and Google I.O. last year.

During the event, so many new updates came out, and it was almost like Gemini had suddenly gone from being down here to all the way up here, right? It suddenly switched, and these things can change so quickly in AI. So let's see what happens. Every major AI company now though, is racing for the same thing.

Not the chat layer, not the model layer, the agent layer, right? The part that connects to your files, your apps, your tools, and actually implements stuff. OpenCore is building a piece of it. Hermes is building the open source version as well.

Google is building the consumer version, one that ships pre-installed on hundreds of millions of devices, already wired into Drive, Gmail, and Search. And for business owners, that's good news, because the barrier to running real AI-powered workflows keeps dropping. And the question isn't whether these tools are capable. They are, right?

They are very, very powerful. The question is whether you have a good system for using them, or whether you're watching the announcements and feeling behind every week. The people who get set up now, who know how to connect Gemini to their files, build skills, and use these sort of AI agents to automate everything inside their business, those people are going to have a compounding advantage over the next 12 months. It's very hard to close later.

And when the official features drop, Tuesday, the AI Profit Boarding will have a full walkthrough live within 24 hours. Here's what that looks like specifically. We'll show you how to add all of these features inside our AgentOS. We're going to show you how to integrate this stuff with Claude and Hermes Agent and OpenClaude.

And we'll show you how to automate pretty much everything in your business using this stuff to get the most valuable possible, right? Also, inside the AI Profit Boarding, we have four weekly coaching calls, tutorials, we have a 30-day roadmap, and 3,000 business owners already running AI automations for themselves as well. So, hope to see you inside there. Cheers for watching.

I'll see you in the next one. Cheers, bye-bye.

More episodes

Browse all episodes →