AI News Today
← All episodes
Episode 63 · July 5, 2026 · 07:36

A1 Agent: New FREE Chinese AI is INSANE!

Agent A1: New Free Local Chinese Agentic Model (35B MoE) + Benchmarks, Speed & Setup

The video reviews Agent A1, a newly released free Chinese local model designed for agentic tasks, showing it running privately and offline inside an Agent Factory workflow alongside free Claude code. The host demos outputs like a landing page, a neon snake game, a “Dragon Realm” open-world game, a keyboard mini app, and a 3D solar system visualization, noting fast replies and that builds save to a workspace with optional voice control. Agent A1 is a 35B mixture-of-experts model (3B active) with a 256k context window, created by Intern Science in Shanghai and optimized for long-horizon search, engineering, scientific research, instruction following, and tool calling. It ranks #2 on the creator’s local leaderboard, outperforming Gemma 4 but trailing Qwable, and is reported around 95 tokens/sec locally on a Mac Studio M4. The video explains strengths/weaknesses, shows running via Ollama and configuring Hermes profiles, and promotes an Agent OS bundle plus OmniRoute routing across 90 free models.

00:00 Meet Agent A1
00:31 Live Local Build Demo
00:53 Leaderboard And Examples
01:36 What A1 Actually Is
02:29 Benchmarks Breakdown
03:08 More Builds In Action
03:45 Specs Speed And Context
04:40 How To Run Locally
04:43 Strengths And Weaknesses
05:41 Hermes Setup Steps
06:13 Agent OS And OmniRoute
06:51 Get The Full Setup
07:34 Wrap Up

Full transcript

Today we're going to be looking at a brand new model called Agent A1 and this is a new free Chinese model, you can run it locally. Let me show you what we've built with this so far. So we actually tested it out inside our agent factory, plugged into free cloud code with a local model and it's pretty good. Like I mean, check out this local landing page we created.

Again, this is free, this is local, it can run privately, it can run offline, we can be on a plane, we can code with this bad boy. And it's pretty easy and simple to use and set up. So you can see some stuff that we've built here. If we want to use it, basically we can just go inside our agent factory here and we could say, okay, build a keyboard, for example.

And what I'll actually do is start using the live build and then start creating here. The other thing I noticed with A1 is that it's pretty fast to apply. If you're wondering how it performs on the benchmarks, we'll come on to that second. If you're wondering how it performs against all the other local models that I've tested recently, on the local leaderboard, it is now ranking number two.

It's not as good as Quable, but it has outperformed GMF4 on our test so far. And if you want to see like some examples of what it can create, I actually created this open world game called Dragon Realm, which is probably one of the best outputs so far I've seen from a local model when it comes to building something like this. It created that landing page I showed you a second ago. It's actually pretty good for coding out local mini apps as well.

And you can see it coding right here. So it's quite fast when we use it directly. We can also control it with our voice using the agent factory and everything that we build gets saved to our workspace too. Now, if you want to learn more about what Agents A1 is, so it's a 35 billion parameter mixture of agents, agentic model.

It came from InternScience, which is a lab in Shanghai. There's a few different versions of the model. This just dropped 24 hours ago, and it's designed for agentic tasks. That means, for example, you could plug it locally into something like Hermes as well and use it locally as a coding agent.

If you're wondering who are InternScience, so they describe themselves as the open source hub of AI for Science Center at Shanghai AI Laboratory. They've created quite a few models. So intern agent as well as something else they've built. And this is designed for long horizon trajectories as well.

So it can work on long horizon tasks according to this. Now, how does it perform on the benchmarks? I always think like test yourself, particularly when it comes to local models. But according to these benchmarks, for example, if we look at HLE, it's outperforming Quen 3.6 and Step 3.5 Flash.

Kimi, DeepSeq and Chachipiti are not too far off on that. And it's particularly good at science benchmarks too. So actually, if you look at the A1 projects page, they say it's optimized for long horizon search, engineering, scientific research, instruction following and tool cooling. Now, if we go back to that task we just gave it, we've got the keyboard here.

We can switch the volume. We can change the oscillator. If you want to see some other stuff that we built with it, the language page was pretty nice. It created like this, you know, basic sort of neon snake game.

Bear in mind, like it's not a frontier model. So it's not like going to compete with Fable 5 or something crazy like that. But it can build basic mini stuff. So for example, like mini apps, landing pages, small games, etc.

Also created this quite nice solar system task here. So this is an example of a 3D visualization from a solar system. It built that out, no problem, which is pretty amazing itself. It's also 256K context window, which is enough for most tasks.

You're not going to hit the token context window for most tasks. And it's designed for agentic reasoning, tool use, long context and instruction following. So it's a 35 billion parameter model. It only activates 3 billion at a time because it's a mixture of experts.

And also in terms of speed, it actually ran faster than Jemma 4 on MOX when we tested out. So it ran at 95 tokens per second, fully local on a Mac Studio M4 compared to Quable, which runs way slower and Jemma 4 Coda way, way slower. In terms of how it performs on the benchmarks versus everything else that we tested, scored 4.8 out of 10 versus Quable at 7.14. So this is still the strongest model I've tested, but it is a lot slower.

And then we've got Jemma 4, which scored lower than Agents A1. But you can also check out GoldieBench.com if you're not sure, if you want to check it out yourself and have a look at the builds we created. Now, if you want to run it locally, you can just run it with a couple of commands here. Now you also might be wondering, okay, what are the strengths and what the weaknesses of it?

So it's tuning is really fine-tuned at search tools and science. So it's not really going to be good for like visual stuff, like UI and that sort of thing. This is more for research and agentic tasks. In terms of the strengths, it's agent-tuned.

So it claims state-of-the-art on SEAL, zero long-horizon search, IFBench instruction following, and BrowseComp. It runs fully free locally as well via Olama. And it's best for like local agentic loops, tool calling, long-horizon, research tasks, and free offline agent work. So for example, you could actually create like a new profile for Agents A1.

We've done that locally here. If we test it out, it actually replies. It's not that fast when it replies, but I think that depends on your setup and what you've got running in the background. However, if you do want a free local model and you want to run Hermes offline, for example, or privately, you can do that no problem with this model right here.

Pretty amazing. Now, if you're wondering how do you actually set up in terms of a new local profile, just make sure that you go into Hermes inside your terminal, switch the Hermes model to Olama, make sure you have Agents A1 set up, and then you would go to profiles over here and just create a new profile with Hermes agent. So that's basically it. That is Agents A1, pretty decent model.

It's free, it's local, actually works, can build some nice stuff. Not super visual, but good for agentic code. And that's what it's designed for. If you want to get our full setup with the agent operating system, we have local systems.

We've got Hermes agent ready to go here. We can create multiple different profiles for local models. We also have the local coding system. We have free cloud code ready to go.

And we have a memory system in here so all your local models can run locally for free. One thing we've also added, if you like free models, is OmniRoot. So we've built in OmniRoot so that you can build and code for free. It's an open source project that just dropped, and it automatically routes across 90 different free models so that you don't get interrupted whilst you're coding with free models, you don't get rate limited, et cetera.

So if you want to get all of that, feel free to get it inside our agent operating system inside the AI Profit Boarding. Link in the comments description or go to the AIProfitBoarding.com. Inside the classroom here, you can get access to our agent OS. We've got a video tutorial.

You can see when it was last updated. You can get the zip file to install it. Every time something new and useful drops, we add a new video tutorial and a step-by-step guide on how to use it. And then also inside the community here, I answer these questions personally every single day with a video tutorial.

Plus the whole community helps you because there's always someone online 24-7. Inside the calendar, you can ask questions and jump on four weekly coaching calls, and then inside the map, you can meet people in your local area who are building with local AI agents and stuff like Hermes. So feel free to get that. Link in the comments description or just go to the AIProfitBoarding.com.

Thanks for watching.

More episodes

Browse all episodes →