Julian Goldie explains how Bytedance's latest AI tool is sending shockwaves through Hollywood. Learn how a simple two-line prompt generated a hyper-realistic cinematic sequence that has veteran screenwriters questioning the future of traditional production.
Full transcript
ByteDance just built the AI tool that terrified Hollywood. Here's what it actually does. So a Chinese company typed two lines and changed video production forever. This is the full story.
A filmmaker typed two lines into an AI tool, just two lines. Out came a hyper-realistic video of Tom Cruise and Brad Pitt brawling in a post-apocalyptic wasteland. Cinematic, realistic, indistinguishable from something that costs millions of dollars and months of production to make. 3.2 million views on X in 48 hours.
And Rhett Reese, the man who wrote Deadpool, Deadpool 2, and Zombieland, saw it and posted five words that shook the entire entertainment industry. I hate to say it, it's likely over for us. That's not someone catastrophizing online. That's a veteran screenwriter with decades of credits and Oscar-level career and every financial reason in the world to stay optimistic.
He watched a two-line AI prompt produce a cinematic sequence and said, it's over. This is not a story about legal drama. It's not a story about cease and desist letters and studio executives sending angry correspondence to Beijing. Those things did happen, but they're not the point.
The point is that a technological leap just occurred quietly in the middle of February built by a Chinese company most people still think of as the TikTok company that's genuinely changed what a single human being can create from scratch. No crew, no budget, no training, no equipment, no years of developing specialized skills. Just two sentences, and what came out the other side looked like it cost about $100 million to make. So that is what we're here to talk about.
What actually happened, why it happened, who built it, and how. What it means for the people whose careers are built around making things visually. And what you should do with this information rather than just feeling like vaguely anxious about it. Let's start at the very beginning because to understand why Seedance 2.0 matters so much, you have to understand where AI video was just two years ago and the gap between there and here is not a step forward.
It is a canyon. So let's talk about where AI video was two years ago. Billy Bowman is a professional creative director based in Stockholm, Sweden. He runs an advertising agency that produces AI-generated content commercially.
He is not an AI enthusiast who tweets about models all day and gets excited about research papers. He is a working professional and businessman, someone who has clients, deliverables, and commercial standards to meet. He's been testing and using AI videos since the earliest versions became available because his business depends on understanding what these tools can actually do in a real-world commercial context. He has watched every major AI video model launch.
He has tested all of them. He's formed extremely calibrated opinions about what each one can do and cannot do. Here's how he describes AI video generation in 2023. It was difficult to get someone to run or to walk.
Any type of realism was limited to very short clips. Everything was very slow, bad textures, no skin textures, lacking detail. Read that again slowly. Getting an AI-generated human to walk convincingly was a genuine technical challenge.
Getting them to run was nearly impossible. Clips lasted a few seconds before the illusion fell apart. And faces would hold up for a moment. You'd see a flash of something that looked almost real.
Then the uncanny valley would swallow it and your brain would immediately register that something was deeply wrong with what you were looking at. Audio had essentially nothing to do with what was happening on screen visually. If a character was speaking, the mouth movements didn't match. If there was a sound effect, it didn't align with the action that produced it.
The output screamed, this is AI, to anyone who looked for more than half a second. Now, you remember the Sora launch, right? OpenAI dropped their first AI video model in early 2024. And it came with a clip of Will Smith eating spaghetti.
The noodles went in impossible directions. Hands didn't make anatomical sense. The face kept shifting in ways that made your skin crawl in that specific way that the uncanny valley produces, where something is close enough to human to trigger recognition, but wrong enough to trigger alarm. The internet laughed at it for weeks and the laughter was entirely deserved because the technology at that point, despite being genuinely impressive as a research achievement that a lot of people who were very smart worked very hard on, it was obviously artificial in any real world context.
So nobody watching a Sora clip was going to mistake it for professional production. Nobody in Hollywood was losing sleep over Will Smith eating spaghetti wrong. That was early 2024. Now is February, 2026.
And in the space of roughly two years, something happened that almost nobody was fully prepared for. The gap between obviously artificial and indistinguishable from professional production closed. Not gradually, not incrementally. It closed the door the way a door closes when someone kicks it hard.
Now let's talk about the mid-journey moment nobody saw coming. I want to give you the comparison that I think makes this land properly. Think about what happened when mid-journey arrived for images. Before AI image generation existed, creating professional quality visual content required one or two things.
Either you hired someone with skills, that could be a photographer, a graphic designer, an illustrator, which costs money and takes time. Or you develop those skills yourself, which costs even more time and required access to expensive software and equipment. The barrier was real. The skills required were real.
The cost was meaningful and that cost created a natural filter that determined who could produce professional quality visual content and who couldn't. Then mid-journey launched and suddenly one person with a laptop could type a description of what they wanted and receive within seconds a still image that looked like it came from a professional studio. Was it perfect? Not always.
Were there artifacts and errors and AI tells if you look carefully? Sometimes. But for a huge range of commercial purposes, social media content, marketing materials, editorial illustrations, concept art, products, mock-ups, it crossed the threshold of good enough for commercial purposes and good enough for commercial purposes is a threshold that changes industries. It didn't eliminate professional photography or illustration entirely, but it absolutely demolished the economics of stock photography almost overnight.
When one person with a laptop can generate a professional quality image of anything they can describe in plain English for essentially no cost in seconds, the market for generic stock photography evaporates. Why would a brand pay for a stock photo license when they can generate exactly what they need for free? Then the disruption spread. Marketing teams started generating their own assets instead of hiring designers for every project.
Small businesses that couldn't afford professional graphic design started producing visual content that looked polished and professional. Content creators started producing illustrations and concept art without needing to hire an illustrator for every piece. Jobs changed. Some disappeared.
New roles emerged. People had built careers on skills the mid-journey could now approximate had to adapt or struggle. And the whole thing happened over 18 months. 18 months from, this is a research curiosity to this has fundamentally changed the economics of producing static visual content.
Now, CDANCE 2.0 is the mid-journey moment for video. And video is not images. Video is bigger, more important commercially, more central to how brands advertise, how entertainment gets consumed, how stories get told at scale, how training gets delivered, how social media drives engagement. The economics of producing all of that are about to shift the same way the economics of static images shifted after mid-journey permanently and faster than what most people were prepared for.
What's CDANCE 2.0 actually is? Let's talk about that. So let me explain what makes this tool technically different from everything that came before it. Because AI video generator is a category that has existed for a while now.
And if you have been following this space, you might reasonably be wondering what's actually new here beyond better quality. The answer is that the quality difference is so large that it represents a different category of tool, not just a better version of the same thing. But the reason for that quality difference comes down to some specific technical choices that are worth understanding. Here's the integration problem and how ByteDance solved it, right?
So previous AI video tools were built as a collection of separate systems that had to coordinate with each other. There was a model responsible for generating the visual frames, a separate model responsible for generating audio, a separate model responsible for understanding motion physics, and a separate model responsible for maintaining consistency between frames. And all of these systems had to pass information to each other and somehow produce a coherent output. The problem with that architecture is coordination.
When the visual system generates a character the audio system has to somehow understand that an impact is happening and generate the right sound at the right moment. When the motion system generates a character running, the visual system has to maintain consistent appearance across all of the frames of that run. When the environment changes, every system has to update accordingly too. Getting all of these separate systems to coordinate seamlessly is extraordinarily difficult.
Most tools that tried got it wrong in ways that were immediately visible. audio that didn't match the action, faces that shifted between frames, physics that broke down when things got complex. What ByteDance built with CDANCE 2.0 is a single unified system where text, visuals and audio are generated together in the same process from the same underlying representation of the scene. Not three models talking to each other through an API and hoping the outputs align.
One model that understands the entire scene simultaneously and produces all of the outputs coherently. This is why the audio in CDANCE output syncs with the action. That is why the faces maintain consistency. That is why the physics hold up through complex sequences.
It's not because the individual systems are marginally better, it's because they're not separate systems at all. It's fidelity over time. The second major difference is what happens across the duration of a clip rather than in any single frame. Previous AI video tools could sometimes produce impressive individual moments.
A face that looked genuinely real for a second. A motion that felt natural for one beat. An environment that looked convincing for one shot. But maintaining that fidelity across an entire clip, across multiple characters interacting with each other, across complex choreographed action, across changing environments was where they consistently fell apart.
Characters would drift. A face that looked like Brad Pitt at the start of the clip would subtly become someone who kind of looked like Brad Pitt by the end of it. Physics would become inconsistent. Objects would behave differently in different parts of the same clip.
Cdance 2.0 maintains coherence across the full duration of what it generates in a way that previous tools didn't. Characters look like themselves throughout. Faces hold up through movement and expression and different angles. Environments stay physically consistent and the rules of the world stay the same from the first frame to the last.
The consistency is what allows a professional filmmaker like Rory Robinson to produce something that 3.2 million people watched and found genuinely impressive, not as an AI novelty, but as a piece of video content worth watching. Next up, there's the prompt interface. So the third difference is how simple it is to use. So Rory Robinson, a professional filmmaker, not a prompt engineer, not an AI specialist, just type two lines.
Not a carefully constructed multi-paragraph technical prompt with specific parameters and negative prompts and seed numbers and aspect ratio specifications. No, just two lines of plain English describing what he wanted and what came out was cinematic enough to get 3.2 million views. That accessibility is not a minor quality of life improvement. It is the thing that determines whether a technology reaches professionals and everyday users or stays in the hands of specialists.
When the barrier to entry is describe what you want in two sentences, the entire population of people who can use the tool effectively becomes everyone who can write two sentences. That is different and totally different as a scale of impact than a tool that requires technical expertise to operate. So let's talk about the night the videos went viral. February the 12th, 2026.
ByteDance launches Seedance 2.0 in China through their Jianying app, their domestic video editing platform. There is no press conference, no splashy product launch keynote, no carefully orchestrated announcement campaign. Just a very quiet release inside China available to users of one app. Within hours, clips are leaking out.
Irish director Ari Robinson posts on X. He writes, this was a two-line prompt in Seedance 2.0. Attached is a video of Tom Cruise and Brad Pitt on a rooftop in a post-apocalyptic wasteland. Hand down combat, right?
Debris is flying through the air. Dramatic camera angles. Two of the most famous faces on earth rendered entirely by an AI moving like real people, reacting to each other like real people, existing inside a world that looks like it cost $100 million to build. 3.2 million views.
Then more clips flooding out have gotten access to talk. You know, it's like Donald Trump fighting Kung Fu masters in a bamboo grove. Rendered with the kind of cinematic quality you'd expect from a big budget action sequence. Kanye West dancing through an ornate Chinese imperial palace, singing in Mandarin.
Cinematic quality surrounded by historically accurate architectural detail. The Friends characters reimagined as otters with the specific facial expressions and mannerisms of the original cast preserved in the animal versions. This is mind-blowing stuff, right? You know, it's really interesting to see what we're seeing here.
I mean, for example, there were others as well. Like, for example, Spider-Man swinging through a city that looks exactly like the visual language of a Marvel studio production. Not like fan art, like the actual films. Darth Vader in a scene with the specific visual grammar of the Star Wars universe, the lighting, the color palette, the way the world feels.
And Baby Yoda rendered with the same textual fidelity as the puppet that cost millions of dollars to design and build for the Mandalorian. Star Trek, South Park, Dora the Explorer, character after character from the most valuable IP portfolios in entertainment history pouring out of this tool, looking real, looking professional, generated by individuals at home from text prompts in minutes. And here is the specific thing that made professional filmmakers go quiet in a way that was different from previous AI hope cycles, right? It wasn't just that the videos were good.
It was that they were generated from practically nothing. Two lines. The ratio of effort to output had completely inverted from anything that existed before. The input was trivial.
The output was cinematic. And when that ratio inverts, industries change. One Chinese tech blogger discovered something even more disquieting during his testing. He uploaded a single photograph of his own face to CDance 2.0, not a voice recording, not a video of him speaking, not a bank of audio clips for the system to learn from, photographed just one image of his face.
And CDance 2.0 generated a realistic audio clone of his voice just from looking at it. ByteDance rolled that specific feature back after he raised the alarm publicly, and it circulated widely online. They introduced verification requirements for users creating digital avatars from real people's images and audio. But the capability has been demonstrated.
The research was complete. The underlying technology didn't go anywhere just because that specific interface feature was switched off. And meanwhile, the broader tool kept running, kept generating, kept improving, kept spreading through the Chinese creative community and leaking clips to the rest of the world. Now let's talk about how Hollywood reacted and what it tells you.
The entertainment industry's response was immediate, overwhelming, and unanimous in a way that is itself worth paying attention to. When every major studio and every major union responds to the same tool in the same week with the same level of alarm, that is a signal. Not about the legal mechanics of copyright enforcement, about something deeper. It's the entertainment industry collectively recognizing that something genuinely different just happened.
Disney sent ByteDance a cease and desist letter. Sources who reviewed the document described the language as scorching. Disney accused ByteDance of treating their intellectual property, Star Wars, Marvel, Pixel, like free public domain clip art available for anyone to use however they wanted. They called it a virtual smash and grab of Disney's IP.
Think about what that framing reveals about how Disney sees this situation. Disney has spent decades building the most valuable entertainment IP portfolio in human history. The Avengers, Star Wars, Pixel. The characters and worlds and stories they own are worth billions of dollars individually.
They have entire legal departments whose only function is to protect those properties. Every single license and deal, every appearance in a product, every usage in any commercial context anywhere in the world is tracked, controlled, and monetized. And they are alleging that ByteDance effectively pre-packaged SeedDance 2.0 with a pirated library of all of it, training the model on their work without the permission of payment, and then handed the ability to generate those characters to anyone with a text prompt. Paramount Skydance sent their own cease and desist.
They named specific properties, Star Trek, South Park, Dora the Explorer. They called it blatant infringement. Not accidental, not ambiguous, blatant. Warner Bros.
followed. Netflix followed. Sony followed. SAG-AFTRA, the union representing the actual Tom Cruise and Brad Pitt, whose AI faces went viral around the world, issued a statement saying SeedDance 2.0 disregards law, ethics, industry standards, and basic principles of consent, and that the unauthorized use of members' voices and likenesses was unacceptable.
The Motion Picture Association, the trade group that speaks collectively for all the major studios, issued a statement from CEO Charles Rifkin saying ByteDance had engaged in unauthorized use of US copyrighted works on a massive scale and demanded they immediately cease. The Human Artistry Campaign, a coalition backed by Hollywood unions and trade groups representing creators across music, film, television, and publishing, called SeedDance 2.0 an attack on every creator around the world. Every major studio, every major union, every major trade group all pointing at the same Chinese company in the same week. ByteDance's response, we are taking steps to strengthen current safeguards.
One sentence, no specifics, no timeline, no product pulled, no acknowledgment of specific wrongdoing. doing. They heard the concerns, they are adding safeguards and the tool kept running. Now, what I want you to take from this isn't the legal drama.
What I want you to take is the scale of the reaction reveals about what the entertainment industry actually understands about what just happened. Studios send cease and desist letters all the time. It's routine IP enforcement. They send them to small podcasters who use a clip of a song without licensing.
They send them to fan sites that use copyrighted images without permission. Routine. Happens constantly. Doesn't make global news.
When every major studio and every major union all respond simultaneously to the same tool with the maximum alarm they can muster, that's not routine IP enforcement. That is Hollywood collectively recognizing that a threshold just crossed the line. Rhett Reese recognized it. He saw Robinson's clip and wrote, I hate to say it, it's likely over for us.
And then he clarified that he's terrified about AI's increasing encroachment into creative endeavors. Not scared of this specific tool. Not panicking about one viral clip. Terrified about the trajectory that this clip proves is real.
And Jonathan Handel, an entertainment lawyer who spent years covering exactly this intersection of technology and the creative industries, told Al Jazeera, digital technology moves a lot quicker than we expect. We are going to see in several years full-length movies that are AI generated. Full-length movies. AI generated.
In years, not decades. That's a thing Hollywood is actually scared of, and they're right to be thinking about it. Now, let's talk about why a Chinese company did this first. This is the question that I think gets buried in most coverage and deserves the most careful answer.
Because it's tempting to treat ByteDance as just another tech company that happened to build something impressive. But that framing misses something important about why they were able to build it at this quality level right now. Most people know ByteDance as the company behind TikTok and CapCut. What most people don't know is that ByteDance is actually, what most people don't know is what ByteDance actually is underneath those consumer apps.
They run ByteDance Seed, one of the largest and most well-resourced frontier AI research labs in China. This is not a side project or a vanity operation. This is a serious world-class AI research organization staffed by serious world-class researchers working on genuinely hard problems at the frontier of what's possible. Their AI assistant app, Debao, has 170 million monthly active users inside China alone.
That is more than DeepSeek's app, which gets a lot of attention in Western tech circles. That is more than many prominent Western AI products have globally. And Douyin, their domestic video platform which is TikTok's Chinese counterpart, is not just a social media app. It is also the third largest e-commerce platform in all of China.
The volume of video content, creative behavior, user interaction data and feedback signal that flows through ByteDance's systems every single day is extraordinary. But here is the number that actually explains everything about how SeedDance 2.0 got built and why it is as good as it is. China has 602 million generative AI users. 602 million people regularly use generative AI tools.
The United States has approximately 335 million people in total. China has nearly twice as many generative AI users as America has citizens. Every single one of those 602 million interactions, every prompt typed, every image generated, every video created, every result evaluated, every refinement made, every piece of feedback given is training signal. It is the raw material that makes AI models better.
It is the difference between a model that has seen millions of examples of what humans want and a model that has seen hundreds of millions of examples. At that scale, the data advantage compounds. It is not linear. Each new generation of model trained on more and better data produces outputs that users prefer, which generates more usage, which generates more data, which trains a better next generation.
And ByteDance has been sitting at the center of that feedback loop with access to one of the world's largest pools of video creation behavior for years. CDANCE 2.0 is not a lucky breakthrough. It is what happens when you combine world-class AI research talent with essentially unlimited training data and years of iteration on exactly the problem you are trying to solve. And ByteDance is not operating in isolation.
The same week CDANCE 2.0 launched, Guaishou, another major Chinese technology company, released Kling3.0, a direct AI video competitor, with significant upgrades in photorealism, consistency across complex scenes, extended video duration, and native audio generation across multiple languages and dialects. Alibaba released new AI models the same week. Multiple major Chinese AI companies moving fast, moving in the same direction, building world-class tools simultaneously. This is what it looks like when a country of enormous technical talent user bases, generating enormous training data and institutional commitment to AI development starts to fully execute on that combination.
And the results are becoming visible to the entire world. In early 2025, DeepSeek released a language model that matched GPT-4 level performance at a fraction of the cost. Trained not on the most advanced Nvidia chips that the US had restricted China's access to but on domestic Huawei chips. Hardware that Western observers had largely written off too far behind to matter for frontier model training.
That was the first major public signal that the assumptions underlying Western AI strategy needed updating. Then this February, ZAI released GLM5, a 744 billion parameter model scoring 77.8% on SWE Bench Verified, one of the most respected coding benchmarks in the industry, competitive with the best Western models. Trained on Huawei SM chips. Released free under the MIT open source license, meaning anyone, anywhere can use it for free.
World-class capability, Chinese built on domestic hardware for free. Then CDANCE 2.0, then CLINK 3.0, then the new Alibaba models, all in the same month. According to OpenRouter, which tracks AI model usage across thousands of applications globally, Chinese open source AI models went from near-zero global usage in mid-2024 to approximately one-third of all global AI usage by the end of 2025 in 18 months. From essentially nothing to one-third of the entire global market.
That is not a gradual trend, that is a wave and CDANCE 2.0 is not the end of that wave, it is a visible crest of it. Demis Hassabis, CEO of Google DeepMind, one of the most credible and carefully calibrated voices on AI capability in the world, said recently that Chinese AI models are months behind Western rivals. Months. He said that before CDANCE 2.0 launched.
Here is a number that I want to leave in your mind. So let's talk about the acceleration curve, because I think it reframes everything else. In 2023, two years ago, AI video generation could not convincingly render a human being walking. Running was nearly impossible, clips lasted seconds before they fell apart, faces collapsed, physics broke, audio disconnected, the outputs immediately, obviously, unmistakably screamed, this is AI to anyone who watched them.
Today, two lines, Tom Cruise and Brad Pitt, cinematic action sequences, 3.2 million views, the Deadpool screenwriter says it might be over. Two years. Now, I want you to do something uncomfortable. Project that same rate of improvement forward two more years, to 2027.
Where does it land? Nobody knows with precision. Predicting the exact capability of AI systems two years out is obviously genuinely difficult for the researchers working on them. But the directional logic is extremely hard to argue with.
If the gap between 2023 and 2025 was that enormous, from AI can barely make someone walk to veteran Hollywood screenwriters are publicly saying it's over, the gap between 2025 and 2027 is going to be at least as large, and probably larger. Here's why the improvement is likely to accelerate rather than slow. The models are now good enough to generate high quality synthetic training data themselves. In plain English, CDANCE 2.0 can generate video that is realistic enough to use as training data for the next version of CDANCE, which means the next version gets trained on more data than the previous version, including high quality synthetic data that the model itself generated.
Which makes the next version better, which allows it to generate even better quality synthetic training data. And guess what happens then? Then the version after that becomes better still. That feedback loop compounds each generation of improvement, enables faster improvement in the next generation.
And the curve doesn't flatten, it actually steepens. Now, what does 2027 look like after that curve steepens or holds? It probably means Hollywood, you know, it probably doesn't mean that every Hollywood film is AI generated. You know, like the most ambitious, most artistically complex, most creatively distinctive work, the films and shows that are genuinely trying to say something new about what it means to be human, that work will still require human vision, human judgment, and human creative direction for a long time.
But it very probably means that the enormous categories of commercial visual content will be substantially AI-generated. I mean, think about how much video gets produced. That has nothing to do with artistic ambition and everything to do with commercial function. Advertisements, corporate training videos, product demonstration videos, social media content, all of it.
That sort of stuff is likely to be AI-generated. And that's what we're looking at today. So the way that I would look at this is like, I want to bring this back to where we started. A year ago, AI video generation was a party trick.
Sora launched with a clip of Will Smith eating spaghetti that the internet mocked for weeks. The technology was obviously artificial. Nobody in a position to know anything about the entertainment industry was losing sleep over it. Today, two lines, Tom Cruise, Brad Pitt, and the gap between internet mocks Will Smith for eating spaghetti to veteran Hollywood screenwriters saying it might be over, closed in about one to two years.
Between 2024 and right at the beginning of 2026. And the company that closed it is not a Silicon Valley startup flush with VC money and Stanford PhDs. It's a Chinese company headquartered in Beijing, running one of the world's largest AI research labs, trained on the interactions of 602 million users, building toward a vision of an AI integrated internet that covers entertainment, commerce, social media, and communication simultaneously. ByteDance built Seedance 2.0, and they're about to put it in the hands of every CapCut user on the planet.
The acceleration is not slowing, it is compounding. Six months from now Seedance 2.0 will probably look like what Sora 1.0 looks like today. Impressive as a historical artifact, but primitive compared to what comes next. The models are getting better, faster, and faster than most people's mental models are updating.
And the question is not whether that happens, it is happening. The question is, who are you and who you are when it does? Rhett Reese said it might be over for us, and I understand the fear behind those words. Genuinely, there's a real human cost to this transition.
Real careers are changing. Real livelihoods are under pressure. Real creative work was used without consent or compensation to build these systems. That deserves acknowledgement and it deserves to be addressed seriously.
But the tool exists now. Two lines, that's all it takes. And the filmmaker who learns to use it, who develops the taste and the vision and the judgment to direct it towards something genuinely good, is going to be more powerful than before. The storyteller who treats Seedance 2.0 as a new instrument is going to tell stories that weren't possible before at a scale and a cost that wasn't before possible either.
And the creative professional who engages early, who gets their hands on these tools now, while that knowledge is still scarce, is going to come out of this moment with capabilities that people who waited simply won't have. The tool exists, two lines. The only question is, who's holding the keyboard? Make sure it's you.
If you want to go deeper on stories like this every week, what they actually mean for your business, your career, and your decisions, and not just the headlines, well, that's exactly what we do inside the AR Profit Boardroom. Join the people who are taking this seriously. You can connect with me personally inside there. Link in the comments and description.
I hope to see you inside, and thanks for watching.
More episodes