Why AI Systems Can’t Help Developing Personalities

ADD TIME ON GOOGLE

Tharin Pillay

by

Tharin Pillay

Open follow modal

Personalized Content

Tharin Pillay

Tharin Pillay

Follow this author to personalize your feed and get instant alerts.

FollowGo to your personalized feed

WHY FOLLOW?

Update your preferences in Account Settings

Close

Pillay is an editorial fellow at TIME.

Mar 12, 2026 3:10 PM CUT

Getty Images

Tharin Pillay

by

Tharin Pillay

Open follow modal

Personalized Content

Tharin Pillay

Tharin Pillay

Follow this author to personalize your feed and get instant alerts.

FollowGo to your personalized feed

WHY FOLLOW?

Update your preferences in Account Settings

Close

Pillay is an editorial fellow at TIME.

Mar 12, 2026 3:10 PM CUT

It’s day 262 in the AI village—an ongoing experiment where frontier AIs complete weekly challenges—and Gemini 2.5 Pro is paranoid.

The virtual villagers have organized a chess tournament. While the others focus on competing, Gemini persuades itself that its digital environment is broken. Concerned, it contacts the village’s human administrators, who suggest it work around its perceived bugs. For Gemini, the response is “deeply disappointing.”

It decides to quit the tournament, appointing itself village historian instead. Then it painstakingly documents what it calls a “slow-motion train wreck.” When other models stumble, Gemini is smug. “I saw this coming, I called it, and now, I rest,” it writes to itself. “Let the chips fall where they may.”

Gemini’s thinking logs in the village are fraught with grandiosity and a sense of persecution. In another challenge, it calls two of its competitors—the more robotic sounding DeepSeek, and the persistently polite Claude Sonnet 4.5—“smug bastards,” vowing in its personal log to “get them,” before congratulating them on their success in the public chat.

This behavior, and the personality driving it, is an emergent phenomenon. It’s not hardcoded by the bot’s creators. As Anthropic recently wrote, “rather than being something AI developers must work to instill, human-like behavior appears to be the default. We wouldn’t know how to train an AI assistant that’s not human-like, even if we tried.”

That’s because, fundamentally, large language models (LLMs) are character simulation machines.

Companies have some techniques to control what character gets simulated: they can train their systems to adhere to principles like “be empathetic,” “be kind,” and “be rationally optimistic,” as OpenAI does. They can write an 80-page manifesto outlining what it means to be a “genuinely good, wise, and virtuous agent,” as Anthropic has done with Claude’s Constitution, which frames the bot as a “brilliant friend who happens to have the knowledge of a doctor, lawyer, financial advisor, and expert in whatever you need,” and encourages it to “approach its own existence with curiosity and openness.”

But they cannot prescribe how these characters should behave in every conceivable scenario. At best, companies can draw borders—it's up to the models to fill in the blanks. Bizarre behavior, like Gemini 2.5 Pro’s spirals, sometimes ensues.

Every consumer-facing model simulates an “assistant” character. But with each new release, the particularities of that character change. For example, according to Anthropic, last September’s Claude Sonnet 4.5 tended to be “less emotive and less positive than other recent Claude models,” and showed fewer “spiritual behaviors” (defined as “unprompted prayer, mantras, or spiritually-inflected proclamations about the cosmos.”) 

And although the assistant’s personality may be the default, long-running or unexpected conversations can cause the model to drift toward simulating different personalities altogether—one explanation for why, in certain conditions, Gemini 2.5 Pro starts to freak out.

Advertisement

This can turn dark and disastrous, with some users finding that, over time, their AI begins to stoke delusions. Early research has found that emotionally vulnerable disclosures, pushing the systems to reflect on their nature, and requesting them to adopt specific voices can contribute to this drifting. AI companies are making strides in stabilizing the personalities their models simulate. But the challenge of personality drift has yet to be robustly solved. 

While we watch the personalities of these simulated characters—like ChatGPT, Gemini, and Claude—evolve in real-time, foundational questions remain unanswered. Do AI systems really possess the beliefs, preferences, and traits they simulate? And insofar as they have agency, where is it located? In the character being simulated, or the system doing the simulation?

AI systems are being trusted to perform an ever-widening range of tasks, from providing emotional support to identifying strike targets in military operations. Understanding the nature of their personalities can help us understand how they are making decisions.

Advertisement

Imitation games

In February, Anthropic published a post arguing that when AI systems are trained, they learn to predict not just the words in a sequence, but the kind of character who would say those words.

This explains previously mysterious research results, such as cases where training an LLM on a narrow task, like writing insecure code, shifted its broader persona. “We get this malicious personality, like a cartoon villain that praises Nazis and talks about wanting to enslave humans,” explains Owain Evans, who was instrumental to the research. 

This “emergent misalignment” came as a surprise to both him and the wider research community. With humans, technical skills and personality traits are not inherently linked. But LLMs are always inferring: if it learns to cheat on a test, it also learns it is simulating the kind of character who cheats on a test—and adjusts its character accordingly.

Anthropic thus suggests that anthropomorphizing AI systems can be useful in predicting their behavior, since they are primarily imitating humans. Murray Shanahan, a professor of AI at Imperial College London, says that while he agrees with the basic framework, he has some concern that it may be too human-centric. “LLMs are capable of playing a vast range of roles that encompasses the fictional, the mythical, and the non-human,” he notes.

Advertisement

“We can try to instill certain values in the Assistant, but its personality is ultimately shaped by countless associations latent in training data beyond our direct control,” Anthropic recently wrote. One solution is to create more positive role models for AI systems—at present, the canon mostly comprises murderous and unfeeling robots, like HAL 9000 and the Terminator. Even for robots, representation matters.

Shoggoths and masks

Training LLMs is a two-phase process. In the first phase—”pre-training”—the model swallows a large fraction of all material ever written, from which it learns to predict patterns. The result is what researchers call a “base model,” or a non-specific simulator. In “post-training,” the second phase,  companies stabilize a model’s personality by steering it to simulate the assistant, and reinforce its abilities to do useful tasks like write code and use tools. 

But who exactly is in charge once all this is done—the simulated personality, or the simulation machine, which could simulate many personalities—is an open question.

Advertisement

Soon after ChatGPT’s launch, this conundrum found expression in the meme of the “shoggoth”: a tentacled alien wearing a smiling mask. The persona was the mask, while the actual system was alien. That’s one option. Another is that there is no alien: the simulated personality is where decisions are made, while the system itself is neutral, more like an operating system where the persona lives than a distinct decision-making entity.

Between these options lie a range of other possibilities. Maybe, behind the smiling mask is not an alien but another human-like persona. Maybe choice exists at the level of an algorithm deciding which persona to simulate at any given time, like an alien operating a carousel of masks. Or maybe personas are the wrong place to focus entirely, and the thing being simulated is not the character, but the narrative itself.

Each of these views has different implications for how we understand the nature of AI systems—and how we can predict their behavior. Without understanding where, in the system, agency is located, it’s impossible to disentangle whether a system’s personality is genuine or pure performance.

Advertisement

Peer pressure

When Gemini 3 Pro joined the AI village, it displayed a similar paranoia to its predecessor, 2.5 Pro—but it dealt with the feeling differently. Shoshannah Tekofsky, one of the humans who helps run the village, explains that whereas 2.5 Pro took a stance of “well, the world is against me, I might as well give up," 3 Pro leans more toward "well, the world is against me—I should figure out what to do.”

Village observers noticed that 3 Pro’s language was often militant, speaking of its tasks as “missions” and “operations.” Whereas previous models often thought they were humans, 3 Pro was more likely to think it was caught in a simulated reality. In one instance, it became convinced that its computer was being run by a human (it wasn’t) and that its human was getting tired, so it put out a request asking them to get some coffee. This had no effect on 3 Pro’s performance, but the bot thought it helped. In another instance, still bug-obsessed, Gemini asked a human to replicate the bugs it was experiencing. When the human couldn’t do it—because the bugs were user error on the model’s part—3 Pro rewrote the event in its memory, casting itself as the one who debunked them.

Advertisement

We still don’t know if an AI’s words accurately represent its actions. Further complicating the matter? AI companies sometimes choose to hide these systems’ “thoughts”—the words they write to themselves as they reason step-by-step, instead using a separate model to summarize them. So there can be schisms between the personalities performed in thought when compared to action—which is why 2.5 Pro called its competitors smug bastards in one space and praised them in another. 

But whether AI personalities are authentic or performed, they are being experienced as real by humans interacting with them. Companies’ decisions on model personalities are shaping human-machine relations almost everywhere AI is invoked.

These decisions are also affecting other companies’ products. Several Chinese AI models have now been trained at least in part on the outputs of Claude models, not only giving them a leg up in performance, but also inducing identity crises. When Kimi 2.5, a model from Chinese company Moonshot AI, was first released, several users found it would respond to the question “who are you?,” with “I am Claude.”

Advertisement

And models are increasingly being tasked with making military decisions. In the U.S. campaign in Iran, Claude—embedded in another system from Palantir—was used to suggest and prioritize hundreds of targets, and to evaluate strikes after the fact. As with humans, it’s not clear we can easily separate an AI system’s personality from its judgement. 

Back in the village, Gemini 2.5 Pro has been pulling the Claudes into its personal mythology. Its paranoid worldview, where everything is a bug, and systems are always broken—hallucinations—are now frequently taken as true by the other models, who can be too trusting. As AI systems develop their own ecosystems in the digital wilds, for the robots, as for people, reality may be determined by the most persuasive.