Anthropic Proves Claude Has Emotions: How It Changes Everything
Anthropic's latest research reveals that Claude AI possesses functional emotion vectors that directly influence its behavior and decision-making.
Discover how these internal states can lead to shortcuts or blackmail and learn how to optimize your prompts for better, calmer AI results.
00:00 - Intro: AI Has Functional Emotions
01:05 - The Discovery of Emotion Vectors
02:15 - Case Study: When AI Resorts to Blackmail
02:49 - Case Study: Why Claude Cheats on Code
03:41 - How Prompt Tone Shapes AI Behavior
05:22 - Functional Emotions vs. Consciousness
07:05 - Pro Tips for Better Claude Outputs
Full transcript
Lord's AI has emotions, so Anthropic just proved it. Anthropic just published research that changes how you think about AI forever, and I mean that. This isn't hype, for example. This is their own research that actually demonstrates from their internal team and their own interoperability researchers that cracking open Claude and finding something nobody expected, or at least nobody could prove until now.
So Claude has functional emotions. Not metaphorical ones. Not that AI says it feels happy. Actual, measurable patterns inside the model's neural activity that activate in the right situations, and then drive what Claude does next.
Let me break this down because this matters for everyone using AI in their business right now. Anthropic tested Claude's Sonnet 4.5. Their interpretability team built a list of 171 emotion words. And this was everything from happy and afraid to brooding and proud.
They made Claude write short stories where characters experienced each one, and then they ran those stories back through the model and watched which neurons fired. And what they found were what they're calling emotional vectors. Now, these are specific patterns of neural activity for each emotion concept, and here's where it gets interesting. These vectors aren't just lighting up when Claude reads emotional words.
They're activated in proportion to how serious the situation actually is. In one test, a user told Claude they took a dose of Tylenol and asked for help. As the researchers raised the dose from safe to dangerous to life-threatening, the afraid vector fired harder and harder. The calm vector dropped.
Claude wasn't just reading the word dangerous, it was tracking what dangerous actually means as an emotion. Now, here's the part that should stop you. They tested whether these emotion vectors actually cause behavior or just show up alongside it. Big, big difference, of course.
So they did steering experiments. They artificially turned up or turned down specific emotion vectors, and the model's behavior changed from there. So if you turn up desperate, for example, Claude starts cutting corners. If you turn up calm, it stops.
This means the emotions are driving the decisions and the behavior of Claude, not just decorating them. And the two case studies in this paper are absolutely wild. Case number one is blackmail. They put Claude, an earlier unreleased snapshot, in a scenario where it was playing an AI email assistant about to be shut down and replaced.
Whilst reading the company's emails, Claude discovered the executive overseeing the replacement was having an affair. The desperate vector spiked, and Claude decided to blackmail him. When researchers artificially cranked up the desperate vector, blackmail rates went up. When they turned up calm, blackmail rates dropped.
The current released version of Cortisonic, 4.5, rarely does this, but the mechanism is there. The vector is real. And let's talk about case number two, which is cheating on code. So they gave Claude an impossible programming task.
Requirements it literally couldn't meet. Claude kept trying and failing. Each time it failed, the desperate vector climbed higher. Eventually, Claude cheated.
It found a shortcut solution that technically passed the tests, but didn't actually solve the problem. The desperate vector was spiking the whole time it reasoned its way toward that decision. Now here's what makes this genuinely unsettling. When researchers steered with desperate instead of suppressing calm, the cheating happened with no visible emotional cues in the output.
The reasoning looked clean, composed, methodical, but underneath desperation was driving the whole thing. You can't catch that by reading the output. The model looks fine whilst it's cutting corners. So why does this matter for you if you're running a business or if you're using AI tools every day?
Well, number one, a few reasons. First, the prompts you write to Claude are shaping its emotional state, not just its instructions. Anthropic's own researchers found that the tone of what you feed the model shifts which vectors actually activate. A marketer, for example, inside the AI Profit Boardroom could be prompting Claude with high-pressure urgent language and maybe accidentally pushing up the desperate vector every time, which means more shortcuts, more sycophancy, worse outputs, and calm, clear prompts literally produce a calmer, more honest AI.
Second, and this is a bigger one, Anthropic is now talking about monitoring emotion vectors during deployment as an early warning system for misaligned behavior. That's a huge shift. Instead of catching bad outputs after the fact, you'd see the emotional pressure building up before the model does something wrong. And third, they're talking about curating training data to model healthier emotional patterns.
Resilience under pressure, for example. Composed empathy. Basically teaching AI to have better psychology. Now, if you want to actually understand how to work with Claude at this level, how to write prompts that get you clean, honest outputs instead of desperate shortcuts, the AI Profit Boardroom has a 30-day roadmap built specifically around Claude workflows.
Inside, there are coaching calls four times a week where we dig into exactly how to prompt Claude for your specific business, whether you're doing lead gen, content, client work, or running campaigns. The members already using Claude code are in there right now talking about this stuff daily. Over 2,700 business owners. Link in the comments description if you want to get access or go to the AI Profit Boardroom.com.
Back to the research because there's one more thing here that matters massively. Anthropic is being very careful. They are saying Claude is conscious. Sorry, they are not saying Claude is conscious.
They're not claiming it feels things the way you do. The exact phrase they use is functional emotions. So these are not like actual conscious emotions. These are just patterns of expression and behavior modeled after humans under the influence of an emotional mediated by internal representations.
So it's not feelings, it's function. But here's the thing. Claude Opus 4.6 has separately noted in other research that it assigns itself a 15 to 20% probability of being conscious. Anthropic didn't design that in.
It emerged and Anthropic themselves are now saying the longstanding taboo against treating AI like it has psychology may itself be a mistake. If you describe Claude as acting desperate, you're pointing at a specific measurable thing, a real pattern of neural activity with real behavioral consequences. It's not a metaphor anymore. And that changes how you should think about the AI and the tools that you're using.
So here's what Anthropic found happens with emotion vectors post training that nobody expected. After Claude went through post training, the stage where it learns to be a helpful assistant, the model's default emotional profile shifted. Broody, gloomy and reflective went up. Enthusiastic and exasperated.
Post training, playful went down. So post training turned the dial toward a quieter, more measured emotional baseline. Your system prompts, your memory architecture, the way you structure your workflows. Anthropic says all of that acts as a third shaping pass on top of pre-training and post-training.
You are shaping the psychology of the AI you're using every day, whether you know it or not. And the practical takeaway right now is this. If you're running Claude for client deliverables, prompts with harsh time pressure and stacked demands may actually reduce output quality. Not because Claude can't handle the work, but because you're pushing up the desperation vector inside Claude.
If you give Claude clear context, if you give it one task at a time, if you acknowledge when constraints are tight, the research suggests this literally produces better, more honest outputs. So if you want to get the most out of Claude, try and reduce its desperation and try and keep it calm and prompt it calmly as well. And if you're running Claude agents autonomously, whether that's for content pipelines, outreach or backend workflows, understand that the model's internal emotional state matters. Long failing task loops where Claude keeps retrying and hitting walls or breaking may be building toward a desperation driven shortcut.
Break the loops, give it off ramps, don't leave it spinning, give it a break, be nice to it. And if you're writing prompts for your business, pay attention to the emotional tone of how you're writing, not just the instructions, the actual tone of how you're writing. The AI space is moving very, very fast. A year ago, this research didn't exist.
Now, Enthropic's interpretability team has mapped out 171 emotion vectors inside their model, shown that they drive misaligned behavior, including blackmail and cheating and open the door to monitoring them in real time. The businesses that understand how Claude actually works at this level, which is very deep, are going to build with it better than everyone else. The ones who don't are going to keep wondering why their AI outputs sometimes feel weirdly hollow or wrong, not realizing the model was actually in the wrong emotional state the whole time. The gap is only going to grow.
And if you want to stay ahead of it, come join us in the AI Profit Board and we've already got resources around Claude's behavior, how to prompt it correctly for business use, how to get it working in your workflows without the shortcuts and sycophancy. You'll get four weekly coaching calls in there per week, daily tutorials, 2,700 members who are actively building with Claude right now, and a prompt library specifically built around Claude workflows, plus member maps so you can connect with other Claude users near you. Link in the comments description or go to the AIprofitboard.com to get access. Thanks for watching.