The episode covers the release of DeepSeek V4 (Pro and Flash) alongside other model updates, showing that DeepSeek V4 is available free to try in instant and expert modes with a 1M-token context window, open-source licensing, and API access, while older DeepSeek API versions will be retired after July 24. The host reviews reported benchmarks versus Claude Opus, GPT 5.4, and Gemini, explains the Mixture-of-Experts architecture (1.6T/284B total parameters with far fewer active), and summarizes efficiency, attention innovations, training approach, and three reasoning modes (non-think, think high, think max). In live tests, DeepSeek’s non-think coding outputs for a landing page look dated, while deep-think improves results but is slower; the host still prefers Claude and found GPT 5.5 more impressive. The video ends by promoting an AI Profit Boardroom community and guides.
Full transcript
China just dropped DeepSeek version 4. On the same day the Chiefty 5.5 came out, OpenCore had a new update, Hermes v0.11 came out, DeepSeek version 4 is here. There's three different models like you can see and it is available to try out for free on expert and instant mode today at chap.deepseek.com. This was literally announced and you can see the details right here.
You can get access to expert and instant mode. There are two different models here so pro is 1.6 trillion parameters and DeepSeek flash is 284 billion parameters. It's a mixture of experts model which means that it has less active parameters depending on what you're using it for and these both have a million token context length but here's the biggest thing these are both open source models and available on the API so this is an open source model there's a million token context window we can start testing out over here and we could try so you can switch between instant and expert on the chat directly and then you can start using this and get access to it so this is one of the most anticipated models ever released and it's just come out which is crazy stuff you know what's interesting as well is like if you actually ask it what version are you using it will say DeepSeek v3 right now but actually it's using DeepSeek v4 as we've seen inside the announcement directly from DeepSeek over here right now if you're wondering okay what are the benchmarks how does this perform so if we look at the current benchmarks this is being tested against older models because it's just been so many updates recently right so if we have a look at this and the updates you can see for example DeepSeek v4 pro max is being compared against Claude Opus 4.6 max and GPT 5.4 high and Gemini 3.1 pro high and as we can see right here DeepSeek is crushing on benchmarks versus many of these models right so you can see here 57.9 and Simple QA verified versus 46.24 Claude Opus 4.6 max for 45.3 versus GPT 5.4 high bear in mind the GPT 5.5 has dropped today so that is another update directly from the AI as well right so basically if we're looking at okay what has changed what number one enhanced agentic capabilities so open source state of the art in agentic coding benchmarks if you're wondering okay what is open source basically it means that you can run this locally depending on how big the model is and how powerful your setup is and also there will be quantized versions which means that you can run smaller models even on a local setup which is pretty cool too rich world knowledge so it leads all current open models trailing only Gemini 3.1 pro on world class reasoning beats all current open models in math stm stem sorry coding rivaling top closed source models and you can see the benchmarks over here too in terms of how it's performing so we've got v4 pro v4 flash k 2.6 glm 5.1 4.6 etc right so these are all being measured against each other and how they perform and what's going on over here all right so we'll test it out in a second in fact you know what i'm gonna go over to deep seek right now and we'll select expert mode and you know what we can do we can run Opus side by side versus this so we'll have Opus open over here and then we'll use deep seek over here right now you have got this option for deep think as well i'm not going to use that because i think that takes quite a long time to use and then we can switch between expert and instant we can also like just test multiple different options here right so we can have for example instant on one tab we can have expert on this tab and then we can have claude opus i still prefer 4.6 anyway to be honest and we'll compare and see how they perform so what we're going to do is you know people always say do something useful so what i'm going to do is create a new landing page for my website so this is my website over here i'm going to take this landing page content i'm going to go over to deep seek and say okay create a new fun interesting alternative to this all right and then i'm also going to say make sure you embed the video from the page as well all right so we'll grab the embed code for this copy that embed that inside then we'll go back to deep seek here and we'll see what it comes back with and how quickly are to perform as well so this is instant it is super fast look at that it's just replying straight away and then also you can see here you'll have a run option once it's completed right so you can run this inside the canvas of the chat and actually see what you've created in a second if you go to this tab now we can do the same thing with expert mode and then we can try this also inside claude opus 4.6 if you want to see 4.7 as well i will do that but just let me know inside the comments if you want me to do that and basically from here we've got all of them coding out and we'll test out what's performing and how how they perform overall now if you go back actually so this will just trigger reasoning so it's actually using expert but it's using without the thinking mode i think it's just going straight into coding which is pretty quick super fast lightning fast to be honest and then we've got claude opus over here i will say like the ui still feels a bit dated if you're using chat deep seek you know compared i mean this is literally the same interface it was like a year ago which is crazy when you look at the way that chat gpt is improved way of analysis improves um there's a big difference and you can feel it when you're using it to be honest but that's basically what we're looking at right here let's see what else we've got whilst we're waiting for those to code out so there's deep seek v4 flash right this is uh something that has reasoning capabilities closely approaching v4 pro performs on par with v4 pro on simple agent tasks so you could use this inside your agents uh v4 flash bear in mind it just still has a million token context window which is great so it's smaller parameter but faster response times and also pretty efficient now let's have a look at that performs so this is actually being compared against you know kimi k 2.6 which literally just came out this week really so it's good that they are comparing it against some newer models but i think things are moving so fast that it's hard for these models to compare themselves against newer models that you know literally just get released day to day it's like every day there's something new um so if we have a look down here we've got loads of cool stuff so long context agentic etc let's see how they perform all right so 87.5 on mmlu pro flash is 86.2 so there's not a huge difference on on this kimi k 2.6 thinking is actually outperforming flash but you've got deep seek v4 pro which is you know crushing all of them interesting stuff all right then we have structural innovation and ultra high context efficiency right so it is a million token context window that's a standard across all deep seek options peak efficiency so world leading long context with drastically compute and memory and then novel attention not sure about that but i would need to to research more about that novel attention and understand what it means so you can see the differences here in terms of like how they perform bear in mind like i think deep cv3 what was that might have been like 250 or 100k context window now you know like v3 particularly just v3 itself was released like early last year so it's been quite a while since the last update and the api is available too so you can get that and start coding with it by the api as well and bear in mind that the old versions will be retired after july the 24th right so after july the 24th deep seek hyphen chat deep seek hyphen reasoner will be fully retired and inaccessible but you can get access to this via the api and directly as well all right so if you want to start using this you can go to platform.deepseek.com and then from here you can start using this you can just grab an api and then you know plug that into ai agents so you could for example because it's like you you could run this for example with something like hermes with open claw etc and i think that'll be pretty easy to to set up and get working so let's have a look at the website outputs now so this is this is oh this interesting right so this is the the new website that's created and this is on instant right so this is the lowest power version now we didn't use deep think we didn't use reasoner mode um but it still created a nice little website as you can see set up um i honestly don't like the design on that i'll be 100% honest with you let's try this version this is a lot nicer so it looks better but it still looks and feels dated you know when you look at the outputs here i'm not going to tell you it's amazing if it's not amazing when i look at this it's it's it feels old it still feels super old whereas you look at for example claude opus 4.6 and bear in mind this is not newest coding model it's nowhere near the same level is it let's be honest um claude opus 4.6 is still still crashing on the the outputs here and then if you want to see something even bigger and better let me show you an example right here so this is one that we actually created with gpt 5.5 earlier today and that is the the website that gpt 5.5 coded out right and it's a lot more complex it's a lot more interesting but it looks a lot more modern right um compared to for example the outputs from deep seek which still feel honestly like version 3 when i look at the outputs i'm like that that feels like version 3 um maybe it's because we didn't have reasoning switched on but look at that it's super basic it's not so i wouldn't recommend it for coding tasks maybe it's good for agentic stuff uh bear in mind you can get access to deep seek for free so this is on the free plan if i could choose between all of these i would go with something like um i would go with claude you know claude is the best lmb i've tried so far honestly and then gd says do you think deep seek v4 is insane model honestly i'm looking at this and i'm going not sure about that to be honest with you i mean look at the outputs here like it's so bad it's so bad uh maybe it's because we didn't have reasoning on so we should give it another shot to be honest but at the same time am i impressed so far if that's your question no no not so far at all um all right so let's try out some more prompts here i'm going to grab some more prompts and then we'll try something else out so i'm going to take this one and we'll try and build out a game with deep seek we're going to use expert mode and we're going to use deep think this time so maybe that was a mistake not using deep think right so we're going to type that in and then over here we're going to say okay make sure that it is computer versus human all right this is something else i tested earlier today with codex and gpt 5.5 all right so if you want to see an example of what we did with gpt 5.5 and this is something that i tested earlier today and here's an example i think it's going to be a little bit laggy inside the tab here but this is an example what we built earlier today right with gpt 5.5 so that's the quality that we're looking at now we're going to test this out with deep seek and see if it could do something better right um yeah claude claude is it you know it can get pricey but at the same time it's the best right and it's always that sort of trade-off between you know do you want the best or do you want like you know something amazing um let's see what else we've got hopefully deep seek is better at that stuff so i think it's it's going to be more efficient but at the same time like the outputs are not not impressive from what i've tested so far at all so we'll see what it comes out with on this one so you can see this is still in the thinking mode it's not created anything yet so it will actually show you the thinking and reasoning before it gives you the output and this is much slower to use right so if you use deep think it's going to take a little while to start coding out start writing anything as you can see here probably think and reason for a few minutes before it actually creates anything gpt 5.5 the few images i've seen look fantastic yeah it's really good for the coding it's great i've done a full review on it already but yeah for coding it's pretty good inside the codex is pretty good as well i don't have access inside chat gpt directly i have to use inside codex the only problem was like the usage limits on it were very very short right you could use it for like 5-10 minutes and then um it seemed to stop working with codex seem to hit the limits quickly even when you upgrade claude is for developers who can't design i think claude is just a really powerful tool you know um i think it's hard to match anything else but if it helps you achieve more if it helps you do more then why wouldn't you use it you know chat gpt is gold standard i haven't heard that one in a while that is a new one for me do you really think that like i think gpt 5.5 yeah for sure like today it's overtaking everything but for the last you know the last month few months or so it's just no it's been nowhere near the same and then also they've said like just rely on our main x account for news so there's gonna just be you know just be careful you get your new sources from um i would just rely on that directly now you can see if we go to the api here it's now available on web app and api right so we can use this directly inside the api here which is pretty cool i'm also interested to see what other people are saying about some people are saying it's gpt5 and gpt 5.5 and claude opus 4.7 level but i just don't see that from what i've tested so far let's have a look at the tech report as well so it's available on hug and face as you can see right here and this is the full tech report i think it's all based on your prompt because i heard the same thing about 4.7 honestly i've tested that same prompt with you know all sorts of things gpt 5.5 claude opus 4.7 and i've never had that sort of output from those models recently right again this is taking quite a long time so bear in mind like you might get better outputs but it's going to be a lot slower bear in mind like for example claude opus just delivered the page like you know within a few minutes so opus 4.6 is a lot faster if you're going to use reasoning mode and if you're not going to use reason mode i just don't think the outputs are very good from deep seek from what i've seen so far all right so it's running here it's beginning to code out we'll test out in a second see how it performs i hope it's good like i really liked deep seek you know when it first came out last year so i hope this is good um let's talk about this so i've actually been through the technical guide which is available on hugging face so this is the technical guide as you can see right here basically it'll show you like you know it's a long document like 55 pages very deep on all the research and how it works etc and i'll talk you through exactly what it means how it works and everything else as well in a second so if we come back we'll wait for deep seek to finish this report and in the meantime let me walk you through this so there's two versions of deep seek you got v4 pro which is the biggest powerful one you got flash which is the fastest one right and the cheapest one so why should you care right now well basically this is an open source model that beats or matches on benchmarks the best closed source models in the world of many tasks so for example we're talking about gpt 5.4 called opus 4.6 in gemini 3.1 pro and it's open source meaning anyone can download and use it right so you can go over to lm studio and if we go to lm studio i'm pretty sure you'll see this on olama soon if it's not already been released if you go to the model section and we type in deep seek before you can see that deep seek v4 will be released eventually i actually can't see it inside let's have a look inside hugging face see what it's called ah look that's what it's called right yes it's actually i can't find it inside lm studio yet but i'm pretty sure later today you'll find all the models for deep seek before you can get it directly from hugging face if you prefer to do that as well right so it's open source anyone can download and use it and the numbers matter as well right so it's super fast right deep seek v4 pro uses only 27% of the computing power that the previous version needed right so it uses 10% of the memory called kv cache compared to v3.2 the flash version is even more efficient so that just uses 10% of the compute and 7% of memory that means it's way cheaper to run and way faster as well right in terms of benchmarks on simple qa v4 pro max scored 57.9% beating claude opus 4.6 max at 46.2% and gpc 5.4 at 45.3% on coding completions it scored a rating of 3206 the highest of any model tested and on the hardest benchmark apex shortlist it scored 90.2% beating every other model including gemini 3.1 pro so we've got the update here let's run this now it did take a long time to build it but it's definitely better than the normal version you can see though it's still like a little bit laggy right so you see how the computer is lagging the pink one on the right hand side it seems like a little bit buggy there i would definitely fix that it took a long time to respond as well i don't know if that's everyone's using it or whether it's just like the thinking mode is is a lot slower um but this is not it's not bad i don't know if it's as good as something like claude opus 4.6 um or something like for example gpc 5.5 i was a lot more impressed with gpc 5.5 today to be honest let's see what we got inside the comments here it's just claude is not really good at agentic when there's a big code everyone's talking about claude compare token oh okay yeah sure one second let's do that so if we go over here we can actually compare them side by side to see how they perform in the meantime let's have a look as well so benchmark performance live coding tasks yeah hit 93.5 the best score ever recorded right so deep cv4 pro max currently ranks 23rd among human candidates on the code force is leadable how does it actually work and so the architecture let's talk about that so it's a mixture of experts kind of like you know you want to think of it like a company with 384 specialist employees when a question comes in only six of those specialists work on it at a time right which means that you get the intelligence of a massive model um but you get it for cheaper right so v4 pro has 1.6 trillion total parameters but activates only 49 billion at a time and v4 flash has 284 billion total parameters but activates only 13 billion at a time now let's talk about the attention system as well this was something i didn't really understand when i first came across it so the old way ai reads long text um is like reading every single word on every single page every single time right that goes very slow so deep cv4 uses two clever fixes to fix this right compressed sparse attention so it squashes every four tokens hits one then only looks at the most important ones i think we're like reading chapter summaries instead of entire chapters and then hca heavily compressed attention um this basically squashes 128 tokens into one so it's kind of like reading just a book title and table of contents so by combining both of these the model can handle a million tokens without using so much compute and that's basically how they've got around the compute problem also manifold constrained hyper connections so this is a fancy way of saying better wiring between the layers of the ai brain right normal ai models pass information from one layer to the next deep cv4 expands that connection to e four times wider this lets more information flow between layers without losing anything important and then you've got the muon optimizer so this is the engine that trains the model so most ai models use an optimizer called adam w deep cv4 switch to something called muon for most of its training this made it learn faster and more stably as well the diego says first force better than chat gpt 5.5 i wouldn't say so at all yeah i don't think it's on the same level and what i've tested and the outputs i've seen not on the same level or something like gpt 5.5 but i will say it's probably more affordable for most people so you know it's one of those and it's like what do you want what do you prefer training data process dc was trained on over 32 trillions of data 32 trillion tokens of data web pages code math problems scientific papers long documents etc they used a special technique of gradually increasing the text length during training and they started at 4k tokens then 16k then for 64k all the way up to 1 million right it's called progressive training and that's why it handles long context so well and then it's free reasoning modes right so you've got non-think mode which is fast instant answers for simple questions you got think high mode which takes more time think step by step great for medium difficulty problems you got think max code which goes absolutely all out on reasoning slow but incredibly powerful right use up to 384k tokens of context right the max mode is what produces most of these benchmark scores so that's how it's performing so well most of these benchmarks which explains the outputs so when we tested out earlier we were using the non-think mode and that was not that good but the outputs from think mode was much better it's just really really slow compared to you know something like for example claude 4.6 so we can use it for you know content gradient writing coding research analysis ai agent workflows and i'm going to have i'm going to put a full guide inside the ai profit bottom on there you can see how it performs um in terms of i mean the costs are way cheaper aren't they look at that right if you look at this table huge difference right there massive difference um and you can see here as well wow big big difference look at that 4.7 versus deep cv4 okay that makes sense that makes sense i mean that's a that's a game changer in itself particularly for like developers and that sort of thing and you can see the full breakdown here right so i mean the output the input is the difference in input of deep seek versus cord opus is huge look at that all right thanks very much for watching if you want to get the full guide from today i will put that inside the ai profit boardroom and also if you haven't checked out the air profit boardroom yet already basically this is my a automation community that helps you learn and grow with a automation to save time to get more customers to generate more leads etc now if you type in deep seek over here at the top you'll get all of my best trainings on how to use deep seek and basically build anything with it right so we show you all the best ways we've got full courses and breakdowns on how to use deep seek how to get the most out of it etc um all of my best trainings on this loads of different ways to use it we have full courses on this and it's all inside the air profit boardroom inside the community you can also get access to 2900 members who are building with this stuff too you can get access to all of my best trainings with new advanced daily tutorials inside the classroom here inside the calendar you can jump on four weekly coaching calls where we go deep on this sort of stuff and then also inside the map you can meet people in your local city who are also using ai agents like deep seek like chat gpt 5.5 like open core and hermes and we drop new daily tutorials inside here as well right so if you want to for example get the full um 30-day roadmap on using deep cv4 i've added that in right here as you can see right so it's a full guide uh with a 30-day roadmap on exactly how to use it a step-by-step operating procedure and you can get that all inside the air profit boardroom along with a full six-hour course on how to use hermes another uh two-hour course right here on open core as well you got all our trainings for example you know all of our new trainings on on everything and we drop new video tutorials with guides every single day on all the new cool stuff that comes out just like this so feel free to get it link in the comments description or go to the aiprofitboard.com to get access thanks for watching i'll see you on the next one cheers
More episodes