Set Up Ollama + DeepSeek V4 Flash (Cloud) and Use It with Claude Code, Codex, OpenCode, OpenClaw & HermesThe video shows how to set up Ollama with DeepSeek V4 Flash (a cloud model) and use it for free with multiple coding and agent tools, including Claude Code, Codex, OpenCode, OpenClaw, and Hermes. The presenter updates Ollama via a terminal command, selects the DeepSeek V4 Flash model on ollama.com, and demonstrates running DeepSeek directly in the terminal, managing multiple terminal tabs, and noting cloud usage limits. They compare capabilities across tools: OpenClaw can drive browser automation while DeepSeek in Ollama struggles with web search; Hermes is described as a smoother agent for getting tasks done; OpenCode and Codex are used for coding outputs like a blog page and a ping pong game. The video emphasizes that results depend on the “harness” (agent tool) controlling the DeepSeek API and promotes additional courses and community support.00:00 Setup Overview00:30 Install and Update Alama00:52 Pick DeepSeek V4 Flash01:08 Run DeepSeek in Terminal01:39 Terminal Tabs Workflow02:09 Connect to Claude Code02:40 Cloud Model Limits03:12 Spin Up Multiple Agents04:00 OpenClaw Browser Test04:49 Hermes and Web Tasks05:32 Comparing Agent Strengths06:27 Coding Demos OpenCode Codex08:53 Harness vs API Explained09:30 Training and Community Outro
Full transcript
Today I'm going to show you exactly how to set up Olama with DeepSeek v4 and how you can use it directly with the latest version DeepSeek v4 on Claude Code, Codex, OpenCode, OpenClaw, Hermes AI Agents 2 and it is so simple and easy to set up. So let me guide you through it. The other thing about this is you can use it for free and just try it out for free and then see if you like it. So let me show you exactly how it works.
You can see for example I've got OpenCode running and building projects right here. So what we can do is you want to first make sure that you have Olama set up. To do that all you do is copy this command then you're going to go to your terminal and you just paste that in. Make sure you update to the latest version.
The reason that you want to do that even if you have Olama set up is that you need to have the latest version running to get this working right. Now once you've done that, it literally takes two minutes as you've seen, then you're going to go over to models on olama.com and from here you would select DeepSeek v4 flash. This is the new model that just dropped, DeepSeek v4 flash. Now once you've done that, you can actually use this locally.
So if you just want to run this inside your terminal, you know, you can just take this like so, plug it in and you can actually talk to DeepSeek directly inside your terminal. As you can see, if you don't know how to open up your terminal, all you do is you press command and space if you're on a mac, type in terminal and then boom you've opened it up right. It's as simple as that. From here, for example, we can just check are you working.
So if we go inside the chat here and just double check is this working, this should be working inside a terminal. As you can see it's thinking in Chinese but it's replying in English. That's what we like. So from here we're going to open up a new tab.
Now if you're wondering how do you open up a new tab, like you can see we've got multiple tabs inside a terminal to run different AI agents. What you do is you press command and t if you're on a mac and that will help you easily open up these tabs and that makes it much easier to organize everything and keep everything together. So you can switch between these coding agents and also DeepSeek itself and then just use them in parallel right and that's going to make you like five times faster, five times more productive by doing this. Now from here what we can do is we can plug this in and start using it with Cloud Code and we can use it with Cloud Code for free right if we're trying this out.
Bear in mind if you're using DeepSeek for V4 Flash you don't need any hardware to run this because it's a cloud model. You're running a cloud model through Olama. What that essentially means is it's not running locally right. You've not downloaded the model, it's running through the cloud servers at Olama and the great thing about that is you can get access instantly.
You don't need to wait for it to download, you don't need any fancy hardware to run it which is perfect. Now also if you're running on cloud models one thing to note here is that there will be usage limits. You can see examples of usage limits right here and you can get a free plan of this and then you can just try it out see if you like it and then from there decide okay you know do I need do I even need like a you know an upgraded plan. So from here we're going to type in for example Cloud Code.
We've already opened that up and now we've got OpenCode. We already have DeepSeek running inside a terminal and we have Cloud Code running with DeepSeek V4 Flash cloud as well right. So easy to do all this. So from here we're going to run ChatGPT codecs right.
Now at this point we've got so many terminal tabs open right we might as well start building stuff out. So I'm going to show you exactly how this works step by step. So we've got for example we've got OpenAI Codecs running with DeepSeek V4 Flash here. We have for example OpenCode over here.
We have DeepSeek inside a terminal as well. I'm going to just make sure that we have OpenCore and Hermes set up and then from there once that's done we can start building stuff out right. I'll show you some of the cool stuff you can do all this. So you can see here how it's beginning to work with all of them which is great.
So we have for example the terminal we have OpenCore. We have Hermes over here. We've also got OpenCode and then we have Codecs as well right. Beautiful.
Now if we want to access OpenCore what we need to do is actually go over to the local gateway host over here as you can see. So let's just test this out now. Let's make sure it actually works. So for example we've got DeepSeek V4 Flash.
We've run all the commands. We've got this all set up. We can see the benchmarks here by the way. Pretty good model.
So from here what we can do just to test this out is we are going to plug this in. So I'm going to say for example go to go and speak to chat GPT right. Just to see okay can it use browser use etc. So it's going to we're going to say go and speak to chat GPT now and let's just check that works.
Boom look at that. It's just opened up chat GPT and it's just opened it up inside our browser. So OpenCore actually works with DeepSeek and it's pretty good for using browser automation. I mean that was super fast and easy to use.
Perfect. Alright let's move on to the next one. So I'm just going to say okay you know can you search the web for the latest AI automation news and we'll do that inside our terminal. Then over here we're going to run a example with Hermes.
So if I say okay schedule in researching the latest AI automation news daily for me. It's going to initialize the agent and then we've got that going as well. Alright and so we've used OpenCore that's working. This is struggling to actually search the web right here.
So I'm just going to go inside the chat and say okay can you create an SEO calculator for me and it should just go off and build that. So just something to note there is like if you're using this inside OpenCore it can access your browser it can search the web etc. If you use it inside Olama it does seem to struggle with searching the web which I was surprised by. But that's basically how you know the difference between them.
So depending on what you need and what you want to do these all have different capabilities right. OpenCore is a great agent and it's also easier and less technical to set up OpenCore with Olama than anything else. And then for example if you're using the terminal well it's pretty simple and easy to like just build stuff out locally. If you're using Hermes, Hermes is great for like just smooth agents right.
And what I mean by that is like if I tell it to go off and do something it's actually going to get it done. Whereas with OpenCore I don't know I might be like 50 50 on whether it's going to do the job or not. So I actually really like Hermes as a smooth AI agent that just works directly for me. So this is running here and you can see it's already coding.
Hermes works perfectly, OpenCore works perfectly. Let's just check open code so this is what now and it's actually coded out something for me earlier right. So if we have a look here and we scroll up you can see that we actually said create this page for us right and this was a blog post that it wrote directly. So if we copy that url it's actually coded locally for us.
So we're just going to copy that go to our window and I'm just going to say open this up and one thing you might wonder is like why would you use these different models like what is each one for. So OpenCore is pretty good for coding you see it's just built out a nice little blog post for us here. OpenCore is pretty good for just running scheduled tasks. Hermes is kind of like a smoother easier version than OpenCore and it you know it's a bit easier.
It's also easier to use inside the terminal which I like too. It's got a better terminal user interface I would say and then for example if you just want to run something like this as you can see you know if you just want to use a chat basically then I would run DeepSeq inside the terminal right. So it's not really that good for web search, not good for agentic tasks inside the terminal directly but DeepSeq v4 is pretty good just as a little chat bar right like inside the terminal here. And then you got Hermes, you schedule tasks, connect it to your phone etc.
You've got OpenCode over here pretty good for coding tasks like for example websites and stuff like that and then you have Codex right and Codex is really cool. I'll show you some examples here so if we say okay build out a ping pong game right it's going to start building this out as you can see here it's creating the game planning out the HTML file etc. Sometimes it does give an error but it actually works right like it actually does work and it does do the job. It's going to create a plan for building out and then it will actually code that directly for us.
So you know OpenCode and Codex are really for coding tasks like building out websites, games, tools etc. Even like mini apps. Then you've got stuff like for example the terminal directly with DeepSeq and that's more for like chat. Then you've got OpenCore which is a good agent and then you've got Hermes which is a good agent and that's the difference between all of them.
How to set them up one click, how to try them for free, how to set up Olama and how to get access to all this directly. So if you want to get more training by the way it's just create the page as well which is pretty cool. So if we click here we can see you know that we've got the game ready to go and that was pretty easy and simple to set up. So that's basically how you can use all these tools together to build out whatever you want.
The one thing that I will say is like whatever you do with DeepSeq Flash and DeepSeq V4 Flash what really matters is the harness that you put it inside right. So for example if we go directly to chat.deepseq.com that is okay but it's not going to give you agentic outputs whereas you put it inside a harness like OpenCore and all of a sudden it's much more powerful right. Why is that? It's because the harness tells the API what to do and it basically controls the API itself.
So the API is DeepSeq V4 Flash and the harness is like OpenCore or Hermes and these are different ways that you can control the API and mold it in the way so you want right. So thanks so much for watching if you want to get more training on DeepSeq on Hermes etc feel free to get that inside the AR Puffer Boardroom. We have a two-hour course on how to use Hermes and a full six-hour course on OpenCore. We have a full guide on how to build and automate anything with DeepSeq V4 and we've also got loads of good trainings.
We update this daily with like all the new stuff that's actually useful as you can see right here. If you want to learn more about Hermes how to build and automate anything we've got a full video tutorial right here with a step-by-step guide on exactly how to use it. We also have a community where you can ask questions get help and support you can jump on weekly coaching calls and you can meet people in your local city using this stuff as well. So that's all inside the AR Puffer Boardroom link in the comments description or go to the ARPufferBoardroom.com you can also connect with me personally inside there you can ask questions and you can also DM me whenever you need help too.
So feel free to get that and thanks so much for watching that is how to use OLAMA with DeepSeq V4 Flash.
More episodes