Julian Goldie breaks down Alibaba's massive Qwen 3.5 release. Learn how a tiny 9B model is outperforming giants like GPT-4 and running entirely offline on basic consumer hardware. We explore the death of API costs, the future of private AI agents, and how to set up Ollama in minutes.
Full transcript
Alibaba just dropped something nobody was ready for and I'm going to be completely real with you right now because this is one of those moments where if you're not paying attention, you're going to wake up six months from now and feel like the world moved on without you because it did. On March the 2nd, 2026, a company called Alibaba, one of the biggest tech companies in the world, based in China, just released four new AI models and the thing that makes this so wild isn't how smart they are, it's how small they are and I mean in the best possible way. The smallest one is called Quen 3.5 0.8 B, that's billion, that's C 0.8 B and it runs on your phone right now today for free without the internet. Let me say that again, a real AI, the kind that can look at pictures, watch videos, understand 201 different languages, help you write, help you think, help you code, now runs on your phone offline with zero monthly fees, zero API costs, zero rate limits.
Alibaba just completed the Quen 3.5 lineup with four small models, 0.8 B, 2B, 4B and 9B, all natively multimodal, all with a 262,000 token context window, all open source under Apache 2.0 license. That's the thing I need you to understand before we go any further. Apache 2.0, that means it's free, not free like a trial, free like you own it, free like you can download it, use it, build a business with it, run it on your server, use it in your product and nobody can ever charge for it or take it away from you. This is not a chat GPT situation where you pay $20 a month and open AI can change the rules whenever they want.
This is yours. So let me walk you through exactly what just happened, why it matters and most importantly what you need to do about it right now. Because if you're a business owner, a freelancer, a marketer, a developer, a manager or honestly just a regular person who uses the internet, this changes things for you. Specifically, personally, let's start with the basics.
What is Quen? Because I know most people watching this have never heard of it and that's completely fine. Quen is Alibaba's AI project. Alibaba is sort of like the Amazon of China.
They do e-commerce, cloud computing, logistics and their AI research lab has been quietly building something that keeps shocking the AI world every few months. Alibaba's Quen team just unveiled the Quen 3.5 small series and the company has been continuing to push the frontier on what open source AI models are capable of and this week they did it again. Now here's the thing that blew my mind. I keep saying small models.
What does that actually mean? Because if you're not a tech person, the word parameters and billion gets thrown around and it sounds impressive but it also sounds like it means nothing. So let me give you the simple version. Think about your brain.
Your brain has roughly 86 billion neurons. More neurons in theory means more thinking power and models work sort of the same way. The more parameters, think of parameters like tiny little connection points inside the AI's brain, the smarter and more capable it tends to be. OpenAI's big flagship models, we're talking hundreds of billions of parameters, trillions even, they're massive, they require data centers, they require millions of dollars of electricity every single day to run.
Now here's what Alibaba just did. The Quen 3.59b model, just 9 billion parameters, now outperforms previous generation Quen 330b on most benchmarks and beats GPT-5 nano on vision tests. 70.1 versus 57.2 on MMU Pro and 78.9 versus 62.2 on mass vision. Read that back solely, right?
So a 9 billion parameter model just beat a 30 billion parameter model. That's like a person who weighs 100 pounds just out arm wrestle someone who weighs 300 pounds and then also beat a top tier competitor from a different company that was supposed to be in its league. This is not a benchmark improvement, this is a different era of AI and the reason this just happened comes down to one thing, architecture. The way the AI is built on the inside, not just how many connections it has, but how smart those connections are.
Traditional small AI models are usually just shrunken down versions of larger ones, so Quen 3.5 small takes a different route. It uses a hybrid architecture combining gated delta networks with sparse mixture of experts. Let me translate that into plain English. Old small AI was like taking a big engine from a truck and cramming it into a car.
Kind of works but it's clunky and inefficient. What Alibaba did is build a tiny engine that was designed to be small from the very beginning with smarter parts and better efficiency so it runs at a smaller size than most people thought was possible. The mixture of experts part is especially important. Here's what that means.
Instead of turning on the whole AI brain every time you ask a question, the AI only switches on the parts it needs. It asks a maths question and it switches on the math expert. Ask it to look at an image and it switches on the vision experts. Everything else stays off.
This saves a massive amount of energy. Alibaba's 35B A3B model for example activates only 3 billion parameters at a time despite having 35 billion total. Proving that better architecture and data quality outweighs raw skill. So you're getting the intelligence of a giant model but running at the efficiency of a tiny one.
And that is how a 9 billion parameter model beats a 120 billion parameter model from OpenAI. Benchmark data shows Quen 3.59B matching or surpassing GPT-OSS 120B across multiple valuations including GPQA Diamond which tests graduate level reasoning at 81.7 versus 71.5 and advanced math competitions at 83.2 versus 76.7. Graduate level reasoning on a model that runs on your laptop, on a model that fits in six gigabytes of memory. Think about that.
Six gigabytes. Your iPhone has more storage than that. Your MacBook has way more than that. This AI fits on a thumb drive.
Now let me talk about what this thing actually does because I don't want you thinking this is just a text box you type it into. This is a full stack AI agent and every single one of these four models, even the smallest, could do all of this. The goal for this series is for autonomy. An agent that can think through reasoning, see through multimodal vision, and act through tool use.
And these small models are designed to do all three. Can see images, can watch videos, can read documents, can write code, can use tools, can browse the web with the right setup, can answer questions in 201 languages, can hold a conversation with a context window of 262,000 tokens which is roughly 200,000 words. That's like having a conversation where the AI remembers everything you've ever said to it across the equivalent of two full novels. Quent 3.5 supports up to 265, uh 256,000 tokens of context.
That's roughly 200,000 words for about two full novels. And think about what that means in practice. You could give this entire AI your business plan, your entire financial history, every email you've ever sent, your entire code base, and it'll hold all of that in its head whilst it talks to you. No AI assistant I've paid for does that so reliable without tracking things.
And this thing actually does it offline for free. Now let me tell you about the four models and who each one is for. The first one is the 0.8b. That's a baby.
0.8 billion parameters. This one is designed to run on the phone, not just any phone, your phone right now. The 0.8b and 2b models are designed for speed and minimal footprint. They're perfect on device AI where latency and privacy matters most, where you need instant responses without needing to be connected to the internet.
So let me paint a picture. You're on a plane, no wi-fi, someone texts you a question in Japanese. You don't speak Japanese. You pull out your phone, the AI translates it instantly, locally, in real time.
No signal needed, no API called to a server in California, no 35 cents per thousand tokens being charged to your credit card. Just happens on your device. Well think about it this way. You're a small business owner, you want an AI assistant that keeps your business conversations completely private, you don't want to feed your customer data into open AI servers, and you don't trust anyone with your data.
With the 0.8b model, your data never leaves your device, period. That is a privacy guarantee that no cloud-based AI can match. The second model is the 2b, two billion parameters. Still phone size, still free, but noticeably smarter.
The 2b model scores 84.5 on OCRBench for reading text from images, and 75.6 on VideoMME for understanding video. Numbers have put it ahead of many 7b class models from just a year ago. A 2b model that reads documents from better images, then models four times its size from last year. This is the pace that we're moving at right now, and this is why I keep telling people the ground is shifting, not slowly, right now in real time.
The third model is the 4b. This is where it starts to get seriously impressive for everyday use. The 4b model acts as a surprisingly strong multimodal base, striking a balance between capability and footprint that makes it ideal for lightweight autonomous agents. Either can take actions, not just talk.
An autonomous agent, that's a key phrase. Not just a chatbot, an agent that can see, think, and do things. At four billion parameters, running on a laptop with eight gigabytes of RAM. Now, I want to stop and acknowledge the skeptics here for a second, because I know some people watching this are thinking, yeah, yeah, benchmarks.
Benchmarks don't always translate to the real world, and they're not wrong. Benchmarks can be gamed. Anthropic CEO Dario Amodi recently pointed out that some Chinese models appear tailored to perform well on benchmarks without being equally impressive in real world use, and that's a fair concern. But here's what I'd say to that.
First, there are multiple benchmarks, different types of benchmarks, third-party benchmarks, and they're all pointing in the same direction. Second, and this is a more important point, developers are already running this thing in the real world and sharing their results. People who built stuff with this model this week and their experiences matter. An AI assistant answers in a fraction of a second, never hits a rate limit, keeps your conversations completely private, and costs nothing beyond your electricity bill.
That's what running a local model gets you, and the Quen 3.5 release has made local AI genuinely excellent. One developer said, and I love this quote, I'm now running a SONNET 4.6 level LLM, local, private, no rate limits, 100% free. That was way too easy. That's a real person who ran this thing and came back to tell everyone about it.
That's not marketing copy, that's lived experience. Now, the fourth model, the 9B, and this is the one that genuinely shocked the AI community. The 9B model strikes a rare sweet spot. Small enough to run on consumer hardware, but capable enough to compete with models 3-9 times its size on serious benchmarks.
At 4-bit quantization, it drops to just 5GB, viable on an RTX 360 or a MacBook M1 with room to spare. 5GB on a MacBook M1. An AI that beats OpenAI's open source 120 billion parameter model on a MacBook M1 with room to spare. I don't use the phrase game-changer lightly.
I've been covering AI for long enough to know that most things called game-changers are actually just incrementally better things. This is different. The 9B model outperforms the previous QN3-80B on long-context benchmarks despite being nearly 9 times smaller, scoring 55.2 on LongBench v2 vs the 80B score of 48. 9 times smaller, better performance on long documents.
This is efficiency at a level that almost doesn't make logical sense until you understand the architecture. Now let's talk about how to actually get this running because this is where I want to separate us from everyone else who just reads the headlines. The easiest way to run any of these models is through something called OLAMA. OLAMA is a free tool that makes running local AI as simple as a single command.
If you've ever installed an app before, you can do this. Step 1. Go to OLAMA.com. Download OLAMA.
It's free. It's one file. You install it like any other app. Step 2.
Open your terminal. That's just a text window on your computer. Mac has it built in. Windows has it too.
And you just type a single command. On OLAMA, all 4 QN3.5 small models support native tool calling, reasoning, and multimodal capabilities. You run them with commands like OLAMA run QN3.5 colon 9B or OLAMA run QN3.5 0.8B. That's it.
Run command. The mode stops. The model downloads. It starts up.
You type to it. It talks back. No API key. No credit card.
No monthly subscription. No monthly bills. For a 16GB VRAM GPU like for example the RTX 4090, the 9B model is the right choice. It fits entirely into VRAM with room to spare for a long context window, generating 80 to 120 tokens per second, making conversations feel instant.
That's faster than most people can read. It's not like some sluggish local experience. That's actually responsive, usable, and daily driver speed. Now, I want to make sure we cover the people who are thinking, OK, but I'm not technical.
I've never used a terminal. This sounds like it's for developers. And I hear you. The honest answer is, this is getting easier every single week.
A year ago, running a local AR model required a PhD and a spare weekend. Today, it's 3 steps and about 15 minutes. But here's what I really want to say to the non-technical people in the room. You don't necessarily need to run this yourself.
What matters is that you understand what this means. What matters is that you understand the world you're stepping into. Because right now, the cost of AI is collapsing. The hosted version of the 35B model, called QN3.5 Flash, costs just 10 cents per million input tokens.
Roughly 1 thirteenth the cost of Claude SONNET 4.6 for compatible tasks. 1 thirteenth. And that's the hosted, pay-per-use version. The local version costs literally zero.
The only thing it costs is your electricity. We're talking pennies. Compare that to what businesses were paying a year ago. $100 a month, $200, $500, just to give their team access to decent AI.
Now, the same quality of AI, or better, is available for close to nothing. So what does that mean for business? Well, it means a competitive advantage is no longer who can afford AI. The competitive advantage is who knows how to use AI.
And that is a very different game. Let me give you a real example of what I mean. Let's say you're running a small marketing agency. And six months ago, you were spending $400 a month on AI tools.
Your big competitor was spending $4,000 a month and getting way better results. They had resources you don't have. Today, you can run the same quality AI they're running, locally, privately, for free. The playing field just got leveled.
And the only question is, do you know what to do with it? Do you know how to build AI workflows? Do you know how to automate repetitive stuff in your business? Do you know how to use AI agents, not just a chat box, but actual automated systems that do the work whilst you sleep?
Because that's the new edge, not the tool itself, the knowledge of how to use it. And I see this pattern everywhere I look now. The businesses that are pulling ahead aren't the ones with the biggest AI budget. So the ones where somebody, maybe one person, maybe two, sat down and figured out how do these tools actually work, how to connect them, how to automate them, how to make them work for specific business goals.
And then they built systems. And those systems ran 24-7 and everything got easier. This is exactly why I put together the AI Profit Boarding, because there's so much noise out there every week about AI, right? Every week there's a new model, every week there's a new tool.
Most people are stuck in the exhausting cycle of trying to keep up with it all without ever actually doing anything about it. Inside the Boarding, we cut through all of that. We focus on what actually moves the needle for real businesses, not theoretical stuff, not research papers, but actual workflows, actual automations, and actual systems you can copy and use from real people sharing what's working right now. If you've been watching from the sidelines, if you've been meaning to get started, if you feel like AI is too fast to catch up, this is exactly who this is for.
Because the gap between people who understand the stuff and people who don't is growing every single week. And I'd rather you be on the right side of that gap. So check the link inside the comments or the description for that, or go to the AI Profit Boarding if it sounds like something you want. Okay, back to Gwen.
Back to what's actually happening. Let me zoom out for a second because I want to give you some context on why this release matters so much beyond just the technical stuff. There's a very real race happening right now between American AI companies and Chinese AI companies. And for most of 2023, 24, 25, most people in the world assumed the American companies were far ahead.
OpenAI, Anthropic, Google, they were setting the pace and China was catching up. And that story really started changing with DeepSeek in January 2025. So DeepSeek released a model that matched GPT-4 level performance at a fraction of the cost. The AI world lost its mind.
American tech stocks dropped. People started asking, wait, how did they do that? China's AI labs are playing a central role in driving the convergence between open source and closed proprietary models. And the gap between what you can get for free and what you pay for keeps narrowing.
And then Alibaba kept going. And now we're here. A model with 9 billion parameters beating a model with 120 billion parameters from one of the world's biggest AI companies. Think about the efficiency improvement here.
A year ago, getting this kind of performance required a server rack. It required, you know, having a big setup, an expensive setup. Today, it requires a gaming laptop. In two years, it'll probably just require a phone.
And maybe it already does for the smaller models. I you can run Quen 3.5, 0.8B on a 17 iPhone Pro. Now, Alibaba has endowed these small models with human aligned judgment through reinforcement learning across million agent environments, allowing them to handle complex multi-step objectives like organizing files or turning gameplay footage into code. Million agent environments.
That's how they train these things. Not just showing them text, actually running millions of simulations of AI agents, completing real tasks and using those results to make the model smarter at doing actual things in the real world. This is why the shift from chatbot to agent is such a big deal now. A chatbot answers questions.
An agent does things. A chatbot can tell you how to write a marketing email. An agent can write the marketing email, put it into your CRM, schedule it to send and follow up with anyone who clicks it. A chatbot, it can explain how to organize your files.
An agent can actually organize your files. And now you can run that agent locally, privately for free on hardware you already own. Now, let me tell you about something that I think is really important and that most coverage of this story is actually missing. And that's the privacy angle.
So we talk a lot about AI capability. We talk about benchmarks and costs and who's ahead. We don't talk enough about the fact that when you use Cloud AI, any Cloud AI, you are sending your data to someone else's server. Every prompt you type, every document you upload, every conversation you have.
Now, most of the big companies have policies about not using your data to train their models. And I believe most of them honor these policies most of the time. But most, and most of the time, it's not the same as definitely always with zero risk. For personal use, that's probably fine.
But for business use, that's a real consideration. With models designed for on-device interference, where privacy is paramount, your data never leaves your device. That's a guarantee no Cloud AI can match. Your data never leaves your device.
So if you're a lawyer, a doctor, a financial therapist, an advisor, a business owner with sensitive client information, a government employee, anyone who handles confidential information, this is not a small thing. This is a massive thing. And the ability to run a genuinely powerful AI locally, privately, with no external server involved, changes the risk calculus for a lot of organizations that were previously saying we can't use AI tools because of data policies. And now they can.
Let me also talk about what this means for developing countries, for people who don't have reliable internet, for people who are in parts of the world where Cloud AI services are expensive or restricted. Quen 3.5's 201 language support with genuinely inclusive worldwide deployment means AI now works for people across regional dialects and languages that most AI systems completely ignore. 201 languages, not just like English or Spanish or French and Mandarin, 201 languages including regional dialects. This thing speaks languages that most major AI companies have never even considered including.
That's access to AI for communities that have been locked out of this conversation entirely. And I think that's worth spending a moment on because a lot of AI coverage is very Silicon Valley centric, very focused on what this means for tech companies in San Francisco. in New York, developers in London, but AI that runs on a phone offline in 201 languages, that's not a Silicon Valley story, that's a global story, that's a farmer in a rural area getting access to expert agricultural advice, that's a student in a country with unreliable internet getting access to a tutor, that's a small business owner with limited tech resources getting access to a tool that gives them capabilities that used to require an entire team, and that is what democratization of AI looks like. Not a press release about democratization, actual AI democratization.
So let's come back to earth a little bit, let's talk about limitations because I don't want to oversell this. The 0.8 model is small, and small means it has limits. It's not going to write you a 10,000 word business strategy document, it's not going to do complex multi-step reasoning at the level of the big frontier models. For truly complex high-stake works, the big cloud models still have an edge in multi-step agentic workflows.
A small error in an early step can lead to a cascade of failures where the agent pursues an incorrect or nonsensical plan. This is one operational challenge teams must monitor for small models, and that's real. Small models can go off the rails in complex workflows, they can make small mistakes early that compound into bigger mistakes later. You need to be aware of that, you need to build in checkpoints, you need to verify outputs.
While these models excel at writing new code from scratch, they can struggle with debugging or modifying complex legacy systems, and that's a real limitation in production environments. So if you're a developer thinking, oh I can replace my entire dev workflow with a local 9-bit model, pump the brakes a little bit. It's excellent for many things, it's not perfect for everything, but for everyday use, for most tasks, most people actually use AI for. So for example like writing, summarizing, translating, answering questions, analyzing documents, helping with decisions, generating ideas, writing code for simple to medium complexity projects, the 9-bit model handles all of that offline for free.
Now let me talk about where this is heading because I want to give you the trajectory, not just a snapshot. Think about what happened with smartphones in 2007. You know, the iPhone came out, it was impressive, but it was slow, the camera was bad, the apps barely worked, and most people thought it was a toy for tech enthusiasts. Five years later, it was running people's entire business.
Fifteen years later, the phone in your pocket has more computing power than the computers NASA used to send people to the moon. Entire industries were restructured around it, and we are at the 2007 moment for local AI. Now, a 9 billion parameter model running on a laptop is impressive. In two years, maybe less.
The phone in your pocket will be able to run a model that would probably blow this away. As agentic AI moves past simple chatbots toward genuine autonomy, models that can think, see, and act, small models that are local enable agentic loops to run for a fraction of the cost of cloud inference, effectively democratizing the agentic era. The agentic era. That phrase keeps coming up in everything I read about AI now, and it's important because right now most people are using AI as a fancy autocomplete.
Ask it a question, get an answer, copy-paste. That's not the destination, that is really just a starting point. The destination is AI agents that manage your calendar, process your emails, update your CRM, run your social media. It's why OpenClaw is so impressive, because it can do all of this automatically, or running locally, or private, and all free.
And that world is not 20 years away. That world is two to three years away, and pieces of it are available right now. The people who are building with this now, figuring out how AI agents work, how to chain them together, how to automate business processes with them, those people are going to have an enormous advantage over everyone else. Not because they're smarter, because they started earlier.
This is the thing I try to hammer home every single time. In a world where the tool is free and the knowledge is the edge, the people who invested and understand in the tools early are the ones who are going to win. Now let me give you a practical framework for how to think about what to do next. If you're someone who uses AI occasionally, but hasn't gone deep yet, start with Olama, download it this week, run the 9b or the 0.8b model, ask it everything you'd normally ask Chachibity.
Feel the difference between paying for something versus owning something. That's a mindset that matters. If you're a business owner or freelancer, start asking what are the repetitive tasks in my business that I do over and over. Those are the first automation candidates, writing first drafts, summarizing documents, answering standard customer questions, generating social media captions, translating content.
Every one of those is something a local AI model can do right now. If you're a developer you already know what to do, but if you haven't run the 9b locally yet, do it today. The 262,000 token context window, a local inference speeds with zero cost is genuinely useful for production work. If you're a manager or team leader, your job right now is to figure out which people on your team are going to be your AI champions.
Who is naturally curious about this? Who's already playing with these tools? Invest in those people, give them time to explore. Because the team that figure this stuff out first are going to run circles around teams that are waiting for formal training programs.
And if you're someone who feels behind, who feels like AI has been moving too fast and you haven't caught up, I want to say something directly to you. You're not as far behind as you think. The tools that are available right now are better and easier to use than anything that's existed even six months ago. The community of people learning this together is enormous.
The resources are everywhere and you can start today and be genuinely capable in 90 days. The only wrong move here is deciding not to start, because the honest truth is we are in the middle of the biggest technological shift in generations, maybe in our lifetimes, right? And the difference between people who thrive in that shift and the people who get left behind is almost entirely a function of whether they decided to learn. The tools are free, the information is out there.
The question is whether you're going to invest the time. A year ago running a multi-modal locally meant needing a 13b plus parameter model and a serious GPU. Now a 4b model with 262 K context, handles text, images and video from 8 gigabytes of VRAM. The calculus for on-device AI has just changed.
Think about that rate of change. A year ago 13 billion parameters minimum, serious hardware required. Today 4 billion, 8 gigabytes of RAM, your laptop handles it. So where will we be in a year from now?
And I don't know exactly, nobody knows, but I know the direction, I know the trend line. And the trend line is pointing toward a world where AI capability is essentially free and unlimited for anyone who has a device and internet access. And even for people who don't have internet access with models like these running offline. That world rewards knowledge over capital.
It rewards people who understand these tools over people who can merely afford them. And that is a genuinely exciting shift if you're willing to show up for it. So let me leave you with this. Alibaba just gave away an AI brain that fits on your phone, that runs without the internet, that speaks 201 languages, that beats models 13 times its size, that you can download today and never pay a dollar for.
The question isn't whether AI is going to change your work, it already is. The question is whether you're going to be the person who figured out early, who built the skills, the systems, who positioned yourself ahead of this wave, or whether you're going to be the person who kept waiting for the perfect moment to start. And the perfect moment is now. This stuff is hard to learn.
Alone it's confusing, it can be very noisy and overwhelming, and there's a million different tools and models and frameworks, and everyone says something different. That's real. I'm not going to pretend like the learning curve doesn't exist. But there are communities of people working for it together.
The AI at Profit Boardroom is one of those places. Real frameworks for real businesses, real automation systems you can implement this week, not theory, actual stuff that works. Because the thing that I've seen over and over again, is that the difference between someone who gets stuck and someone who actually builds momentum, isn't intelligence. It's not even time, it's having the right map, the right framework, the right people around you who have already been through it.
So if you're serious about not getting left behind, I'd start with the AI at Profit Boardroom. Check it out, link in the comments description, or go to the AI at Profit Boardroom.com. But regardless of what you do next, go download Alarmer, run Quen 3.5, ask it something you'd normally pay $20 a month to ask, feel how fast it is, feel how private it is, feel how it's just yours, and then start imagining what you're going to build with it. Because the tools are here, they're free, they work, the only question left is what are you going to do with them?
And I genuinely believe that question, that choice, is one of the most important ones you'll make this year.
More episodes