Differentiation and personality in AI models
Thoughts on model "culture" as a kind of lock-in
Ask ChatGPT, Claude and Grok the same question and you get three different answers. Yes, the facts are (largely) the same, but it’s clear that three different personalities answer back. OpenAI shipped first and established a template for what a helpful assistant sounds like, careful and accommodating and reluctant to give offense. Anthropic leaned into guardrails and values and character and trained Claude against a written constitution. Elon Musk built Grok to be the rebellious one, the anti-woke alternative to everything Musk thought had gone soft in the rest of the industry.
It’s not just the models that are differentiated, though. The experience of individuals interacting with AI is shaped by the personality and values designed into the system. And as we accumulate more context with a model, we are more and more a “ChatGPT user,” a “Claude user,” or a “Grok user.” This is not entirely unlike being a Mac user vs. a PC user, or an iPhone user vs. an Android user, or being part of the Apple ecosystem or the Google ecosystem. Companies depend on this tribalism. It’s part of their moat. But in AI, the moat may be deeper than in past environments where our belonging was marked by the programs we use and the artifacts they created for us, because here it is also a matter of shared landscapes of thought.
Gregory Bateson gave a name to this phenomenon in the 1930s, after fieldwork among the Iatmul people of the Sepik River in New Guinea. He called it schismogenesis, “a process of differentiation in the norms of individual behaviour” driven by repeated interaction. He saw two forms. In symmetrical schismogenesis, each side answers the other with more of the same, boast for boast, as if it were a kind of an arms race. In complementary schismogenesis, the behavior of one draws out the behavior of the other. I’m not sure I completely understand Bateson’s distinction here since I read his Steps to an Ecology of Mind over 50 years ago and haven’t looked at it since, but sometimes even a misunderstood concept can still be a tool for insight. (I’ve used Claude to help me tease out my half-remembered lessons.)
Relating these two forms to AI, I would say (correctly or not) that OpenAI, Anthropic, and DeepMind are engaged in a form of symmetrical schismogenesis, initially all trying to outdo each other in promises of AI safety, and with boasts of the power of their models to shape the AI future. X.ai/Grok, on the other hand, represents a kind of complementary schismogenesis, an explicit effort to be what the others aren’t. Mistral (and sovereign AI in general) represents another kind of complementary schismogenesis. After all, if you are French, or Chinese, or even just a company trying to carve out your own place in an increasingly homogenized world, do you really want to adopt your values wholesale from whatever is on offer by the big labs?
David Graeber and David Wengrow ran with Bateson’s idea in their book The Dawn of Everything. They used it to explain neighboring societies organized as near mirror opposites. On the Pacific coast the fishing peoples of the Northwest were hierarchical, kept slaves, and threw competitive feasts, while the acorn-gathering peoples of California to their south were industrious, frugal, and suspicious of hoarded status. The difference was not just determined by climate or crops. People became what they were partly by refusing to resemble the neighbors across the way. Graeber and Wengrow describe these cultural choices as a kind of play, but it is serious play that can harden into rivalry and conflict.
A schismogenetic tree
The idea that schismogenesis is a kind of cultural game minimizes environmental factors that can shape it. Google Research built the transformer in 2017, the architecture every one of these models runs on. Yet because of the company’s heritage, Google’s first implementation, BERT, was positioned as an improvement to search, used to better understand the intent behind search queries, particularly longer, conversational, or preposition-heavy searches. By 2021 Google had a conversational model, LaMDA, good enough that one of its own engineers went public the next year claiming it was sentient. But Google never shipped it. A model that answers your question in a paragraph threatened Google’s search franchise, more than $160 billion a year and most of Alphabet’s revenue. Google could have built chat earlier than anyone but had the most to lose by shipping it. This is a version of what is sometimes called “the Kodak curse.” Kodak made early breakthroughs in digital photography that they never properly commercialized because of the desire to protect their highly profitable film/chemistry business.
OpenAI had no search revenue to protect, so putting a chat box on the open web in 2022 cost it nothing. And so the leader became the follower. After ChatGPT reached a hundred million users in only two months, Google CEO Sundar Pichai declared a code red, pulled founders Larry Page and Sergey Brin back into product meetings, and within weeks Google shipped Bard, whose first public demo got a fact about the James Webb telescope wrong and knocked $100 billion off Alphabet in a day. In short, the company that invented the technology arrived late, got rattled, and ever since has been building chat into search with one hand while defending search from chat with the other. So yes, the environment is a factor ;-)
Schismogenesis usually works below the level of intention, though. It’s an accumulating cultural drift nobody quite chooses. But in the case of the big AI models, schismogenesis seems to have been quite deliberate.
With Claude and with Grok you could watch their owners reach in and deliberately turn the dial. Claude drove AI further into caution and a published set of values. Musk took the other side. When early testers ran Grok through the standard political batteries and found it sitting left of center, near ChatGPT, Musk said xAI would move it, and over the next updates its answers marched right in step with his own posts on X.
Google did not differentiate because it was held in place by the business it had to defend, but OpenAI and Anthropic and xAI differentiated because they were free to, with no franchise holding them back. It’s true that OpenAI was driven by strategic business considerations and its hope to dethrone Google, but both Anthropic and X.ai made their choices initially for cultural reasons. They are a pure demonstration of schismogenesis.
So too, the open-weight world has defined itself against the closed labs, and the choice to publish weights is now as much an identity as an engineering decision, a way of saying “we are the ones who do not lock you in.” Meta leaned on open weights to distinguish itself from OpenAI and Google, and the Chinese labs that shipped strong open models turned openness into a powerful business strategy.
Further up the stack, differentiation is the whole game. When the underlying capability commoditizes, the players who capture value are the ones who stake out a position nobody else holds. Schismogenesis is not a cultural curiosity. It is a competitive strategy, and in a market this crowded it may be the main one.
Steve Jobs was good at this
None of this story would have surprised Steve Jobs, who turned schismogenesis into a marketing case study. Apple’s 1984 ad did not sell the Mac on the basis of its processor speed or available memory, as PC advertising of the day was likely to do. Ridley Scott shot a runner sprinting into a hall of gray obedient faces and throwing a hammer through Big Brother on the screen. The tag said the Macintosh was why 1984 would not be like 1984. IBM was conformity and the Mac was the hammer. Thirteen years later Think Different, a riff on IBM’s one-word slogan “Think!”, put Einstein and Gandhi and Lennon on billboards and toasted “the crazy ones.” Apple built a company on being the deliberate opposite of the beige box on the office desk.
Scene from Apple’s legendary 1984 Superbowl ad.
The AI labs might learn something from Jobs, because his differentiation was generative in a way that a purely oppositional version is not. Apple defined itself against IBM, but it defined itself for something, the individual, the artist, the person who wanted a tool that felt like it was on their side. IBM was the foil, but the content was a positive idea of who you became by choosing the Mac. That is the difference between an identity you build for yourself and one you shape too narrowly in response to a real or perceived enemy.
I think Anthropic understands positive differentiation. They started with their core values, but they used those values to choose a market position with thoughtful, careful, and caring AI at the heart of it. Their users depend on the reliability that their values offer. Grok, by contrast, has the foil without the substance. Its identity is principally the negation of the models it dislikes.
Of course, not all differentiation is schismogenesis. Different teams attacking the same problem come up with different solutions, keep what advantage they can proprietary, but at the same time they try to copy the best from their rivals. Much of the stack actually converges. Claude Code and Codex feel like siblings. But we don’t want a monoculture. We need healthy schismogenesis. It’s a competitive frontier, part of what the poet Wallace Stevens called “search[ing] a possible for its possibleness.” It’s also a cultural frontier. People all over the world don’t want one AI that reflects one set of values. Sovereign AI is not just an economic imperative, it is also a cultural one. But there’s a real risk in that.
Bateson did not think schismogenesis was benign. He thought that on its own it ran to breakdown, symmetrical rivalry escalating into open conflict, complementary difference hardening into rigid domination and submission, and that a society survived it only by having some countervailing mechanism that periodically reset the tension. Among the Iatmul it was a ceremony, the naven, that inverted the ordinary roles and let the pressure out.
In tech, the equivalent to the naven might be the standards body ;-) Or maybe it’s open source AI, and the role of standardized protocols in enabling companies and individuals to flourish beyond the boundaries that are set for them. This is a trailing thought, but that’s one of the things I learned from Frank Herbert, also 50 years ago, when he told me that one of his goals in Dune was to have his readers “go skidding out of the story” with unanswered questions that kept them coming back to his world for more. I leave you to ponder what our AI naven might be.
Freedom to leave
There are brakes on schismogenesis at the lab level. A model has to be useful or people stop using it. Users can leave for another model with a click. The public still reacts, the way it did when Grok’s MechaHitler episode drew a bipartisan letter from Congress, a rebuke from the Anti-Defamation League, and a resignation at the top of X.
The quest for sovereign AI and the role of open weight models and open source AI in giving power back from the labs to users are also a kind of brake, a competitive check on the power of the big labs to impose their values.
The simplest and most powerful reset, though, is the freedom of users to leave, to choose an alternative. This is real in AI. The moats that the labs have built so far are relatively weak, as we’ve seen recently with widespread adoption of GPT 5.6 Sol when Fable became temporarily unavailable. I had a brush with this myself. I ran out of Fable usage in my Pro account in the middle of a fine-tuning project, and decided to try moving it over to Sol. I had managed the project well, so there was a CURRENT_STATE.md with the project description and what had been accomplished so far, plus the original training set, the cleaned training set, and so on, all bundled up into a zip file. I handed it to Sol, it reviewed the work, and picked up right where we left off.
At the same time, though, I have gotten better at working with Claude and Claude has gotten better at working with me. The two of us have converged into a paired unit around the tasks I use it for. Every session cuts the groove a little deeper. My prompts tend to be for overlapping tasks, its memory fills with my context, and the fit is the product. I’ve tried to condense some of the work we’ve done together into a skill that could be executed by other people at O’Reilly, and was surprised to find that the skill + my context works better than the skill with someone else’s context. So skills as a kind of portability that I thought I could depend on turned out to be weaker than I thought.
This pairing with a model might be a kind of complementary schismogenesis at the scale of one person. The user and the model differentiate from all other user/model pairs together, and the pair pulls away from every alternative I am not using. There is a widening gap between me-with-Claude and me-with-anything-else.
It’s a marvelous feeling. It is also lock-in. The switching cost is not (yet) a file I can download and pass on reliably. The co-adapted relationship, the accumulated context, and the shared work have become one company’s moat. It erodes our freedom to leave. The more the pairing helps each of us, the more it costs to walk away. For many tasks, I haven’t even switched from my Pro account to my company account, because I haven’t yet figured out how to move over all the necessary context. For a given task, it might be as easy as my switch between Fable and Sol, but for the whole package of my Claude relationship, perhaps not.
I imagine that with some spelunking, I will find where Claude keeps all this context it has for me, and I hope most of it will be something I can move. But I’d be a lot happier if portability were something we could all take for granted. If my memory and context were mine to carry, in a format another model could read, pairing would not harden so easily into capture.
We do not have that portability standard. It is one of the missing mechanisms of the agentic economy, plumbing that a market needs to stay competitive. It’s not just memory, though memory may be the current best hook to work on. There are also differences in terminology, filenames, working patterns, skills, and assumptions. Build for portability and the choice of model remains fluid. Leave it unbuilt and the intimacy we build with our favorite models can become a source of lock-in.


