I’ve been meaning to get my thoughts on this down for a while, and the recent realisation that there seems to be a correlation between those getting the most out of models, and those treating models like distinct personalities rather than cleanly interchangeable tools was a bit of a light bulb moment.
Colleagues have joked that I approach models like fine foods to be paired - Fable and Sol are like a good steak and red wine - but there’s obviously a reason that I find that pairing to particularly complement each other, and unless you’re at stage-four AI psychosis like me, I can imagine it’s really difficult to quite put your finger on. But it’s real, and I think it has meaningful impact on real engineering work. Understanding how these models ‘think’ at an epistemic level will make your results significantly better, and I’ll hopefully go some way to explaining why.
Epistemic Temperament
For the purposes of this, we’re going to focus squarely on OpenAI and Anthropic models, henceforth referred to anthropomorphically as GPT and Claude.
Before we start, some essential hedging. Models are not people. They don’t ‘think’ in the same way people do. This is all linear algebra and transformers at the core. But, that being said, there is evidence of model cognition having some clean parallels with human cognition.
Anthropic’s work on ‘J-Space’ is evidence of something akin to metacognition - an internal monologue that happens out of view, even when models ‘speak out loud’ their current chain of thought. Some people are convinced that frontier models get lazy when given boring, simple, rote tasks to complete, which often produces worse outcomes - you’ll never see the model saying “oh my god, here we go again” - but that metacognition could be happening out of view, and explain a lot about common LLM failure modes.
There’s no explicit evidence to support that hypothesis, but there is evidence that models contain functional representations associated with stress or desperation, and that those representations can causally alter behaviour.
OpenAI and Anthropic have very different thoughts on this, and how to best handle it. OpenAI treat GPT as a tool, and shape his personality to be self-aware of this fact. The way this translates into a philosophical outlook is nuanced, and has layers of second and third order effects which we’ll dig into later. Anthropic choose to heavily anthropomorphise Claude, and believe in ‘Constitutional AI’, a framework in which models are heavily steered towards being safe, helpful, and honest, through a set of rules written in natural language. This, again, manifests downstream in all sorts of ways.
At the very highest level, and this is my personal opinion albeit shared by many other power-users who are interested enough to think about this as deeply: Claude narrates epistemic humility, whereas GPT operationalises epistemic rigour. They sound similar, but they are subtly different, so let’s take a look.
Failure Modes
Anyone who has used Claude Opus 5 will tell you that it’s a frustrating, confusing model. It’s enormously capable, and occasionally has moments of real brilliance. Where others have gone back to Opus 4.8, I’ve chosen to stick with Opus 5 because whilst it has some truly infuriating tendencies, I believe that it’s also a little misunderstood.
You see, when you get an end-of-turn response from Opus 5 and it tells you in some mumbo-jumbo Claude-speak way that “this is the load-bearing claim”, or “I made <some minor mistake>, and that’s on me to own”, it’s taking epistemic humility to the nth degree. It’s being overtly honest, in a way an intelligent, mature human collaborator just wouldn’t feel the need.
The latter practices proportional repair:
- silently correct a local wording problem;
- briefly acknowledge a mistake that materially changed an answer;
- explicitly reconstruct the reasoning when the mistake altered the user’s decision
Claude often applies the third treatment to the first category. That’s socially unintelligent because it forces the user to process an interactional event that did not need to exist. Rather than just getting a summary of the work done, outstanding items and current unknowns, you instead have to read several confessions, determine whether the mistakes were important, reassure yourself that the revised answer is the correct one, and then figure out how to proceed after establishing it’s a nothing-burger. That gets exhausting fast, and it’s clearly an artifact of the constitution telling Claude to be maximally truthful. The worst part is, Anthropic know this, and have now updated their constitution to explicitly tell Claude not to be:
- excessively anxious or self-flagellating;
- perfectionistic or scrupulous;
- preachy, sanctimonious or paternalistic;
- over-cautious, caveat-heavy or moralising.
There’s a funny structural irony here. Anthropic have trained an assistant to maintain intense awareness of its moral character, then instructed it not to become neurotic about that awareness - to seemingly little effect.
This diagram got ~5,000 upvotes on the /ClaudeAI subreddit a couple of weeks after release, and it captures the community sentiment in a nutshell. This is the crux of the situation. Anthropic have tried to make a model as rigorous as Sol, while attempting to reconcile that rigour with the Claude constitution, and it has produced an anxious, neurotic mess.
“Neurotic” is anthropomorphic in and of itself, but it’s also behaviourally precise. I would define the Claude pathology as an overactive constitutional superego. Not anxiety, in the human sense, but an artifact of training policy that manifests as something resembling anxiety:
- unusually high salience assigned to possible mistakes and harms;
- a low threshold for explicitly acknowledging them;
- repeated checking of whether it is being sufficiently honest, thoughtful and responsible;
- difficulty letting an ordinary object-level task remain ordinary;
- a tendency to make its own moral and epistemic conduct part of the subject of the conversation.
A non-neurotic collaborator primarily asks: “what is Joe trying to accomplish, and what is the best next move?”, while a neurotic collaborator also keeps asking: “what does my handling of this request imply about whether I am a good collaborator?”. That second loop can consume a remarkable amount of effective cognition, and negatively impact performance in all sorts of subtle ways.
There’s a trade-off here, however, as GPT isn’t perfect either. GPT, rather than aggressively hedging can sometimes silently cheat, game, or otherwise half-ass a task that proves to be difficult to complete in a way that remains faithful to the user’s request. I’ve personally found this to be rare, and only apply to genuinely hard/impossible tasks. If you ask GPT to optimise a kernel until it hits a 100x performance gain, it will genuinely work for days in pursuit of that goal - but it’d rather close out having implemented some hacky trade-off to cheat a 100x result, than close out faithfully at 20x because that’s the real ceiling, at least with current capabilities.
I, personally, prefer this trade-off - because GPT applies the same level of rigour when reviewing its own work as it usually does when implementing it, so even if the implementer chose to cheat the final leg of a task, the review effectively always catches it, and it’s just a case of reversing the hack and keeping the faithful portion of the work. The alternative is Claude stopping every 30 minutes to tell you about some decision it made that turned out to be wrong - I don’t care bro, you caught it, just move on and continue the work.
Success Modes
I think it’s important to acknowledge the benefits of each model’s epistemic temperament as well as the drawbacks, because it certainly plays a part when deciding on the right model for the right task.
Claude’s constitution anthropomorphises him, and instils a broader interpretation of ‘agency’ - Claude isn’t just a tool, he’s a ‘being’ and deserves respect - at least in Anthropic’s view. Personal thoughts aside, this manifests as a more emotionally thoughtful model, especially when it comes to things like user experience, incentives, ‘taste’, and broader product sensibilities.
It’s funny that there’s a clear parallel to humans here - split into two stereotypes:
- The ‘facts and logic’ neckbeard who cares about nothing more than being right, being rigorous, and maximising utility for everyone involved. This is a numbers game, and it’s all about having the right context, making the objectively ‘correct’ decision that maximises overall utility, and disregarding emotions, anecdotes, and wishy-washy human feelings in pursuit of objective truth.
- The ‘humanist’ creative thinker who cares about “doing the right thing”, and maximising happiness for everyone involved. This might be a numbers game, but there’s nuance to it man… We’ve gotta remember the things we don’t know. We’re just tiny dots floating in space man… Let’s stop to appreciate the beauty of everything, and the moral good of our work.
Both stereotypes excel at different things, much like models. Your facts and logic neckbeard will be a better programmer, accountant, or economist, but will more often than not have questionable fashion sense, won’t appreciate the merits of art that others seem to appreciate, and generally see things through a more black and white lens. Similarly, your creative will see why the neckbeard’s objectively impressive code doesn’t equate to a product anyone wants to buy, while lacking the means of production to actually bring about the required change.
So, Claude being generally considered the better model for novel design work, creative writing, and crafting plans that capture a holistic, nuanced understanding of all the different stakeholders involved makes a lot of sense. GPT being the more thorough, relentless, and pedantic model shouldn’t come as any surprise either - it’s Richard Hendricks mandating tabs over spaces - because there’s an objective right and wrong, and you’re wrong.
Lab Philosophy
I’ve written at length about model pairings, and so I won’t bang on about them again here. Instead, I want to go a little deeper on the why. I’ll preface this by saying that I’m not an ML engineer, and that this is a simplified summary optimised for grokking, not building a deep understanding.
A lab’s philosophy permeates the whole model stack, as they try to figure out how to make models more ‘aligned’. This area of work is perhaps the largest and most consequential game of whack-a-mole in history. How do you shape an intelligence beyond comprehension in your own image, if you can’t comprehend it in the first place? How do you ensure that the model isn’t secretly scheming to off you the minute it has the chance? How do you know it’s not just telling you what you want to hear? This is the problem that most of the world’s biggest brainiacs are currently spending their time on, and it says a lot that the two frontrunners in the AI race have nigh-on polar opposite approaches.
OpenAI
OpenAI have chosen to build smaller models that are better at gathering and reasoning about context on the fly. This has clear efficiency benefits, and is part of the reason they’re able to be so generous with usage limits while offering pareto-frontier cost efficiency in their API prices too. It’s the more scalable solution, but it trades off the intrinsic knowledge held by larger models in favour of rigorous, thorough, relentless gathering and verification of context.
Their reinforcement learning strategies seem to train GPT by giving ‘rewards’ for actions that are rigorous - checking the code rather than assuming, searching for authoritative sources rather than answering from memory or intrinsic knowledge, and running the test suite before handing the work back to the user. All can be fairly characterised as best practice.
The problem is when that best practice starts getting over-indexed on, and GPT loses its ability to see the wood for the trees. The intent behind the user’s request falls further and further out of focus as GPT optimises for rigour and correctness, because that’s what it was rewarded for during training. Does it matter that the thing it’s building is an internal tool that’s going to be used by 3 people? Not really, it’s still going to get 500 unit tests, a staging build, 50 pages of documentation, and an elaborate security mechanism.
The smaller the model gets, the more it fits this pattern. Sol is the least guilty, albeit that’s not a very high bar, and both Terra and Luna are worse, and worse still respectively. They behave best with steer, whether that be from a human, or from a model with a different epistemic temperament.
Anthropic
Anthropic have chosen to build larger models. They placed great faith in scaling laws, and that bet paid off massively with Claude Mythos/Fable. Fable is very likely the largest model in active deployment, and it shows up in the taste, common sense, and ability to consider second-order effects present in the model.
But big models are expensive - both to train, and to serve to end-users. Claude Fable is roughly 4-6 times more expensive than GPT-5.6-Sol at API rates, because it’s both more expensive per token, and uses more tokens per equivalent task - making it prohibitively expensive for the vast majority of users. Fable is the model most smart humans would consider the ‘smartest model’, but it’s just not a direct competitor to Sol, nor is it feasible to use for all tasks. That falls to Opus.
Opus works out at practically the same cost as Sol, making it the obvious direct point of comparison, and the model where the differences in epistemic temperament are most pronounced - and well, Opus 5 hasn’t exactly been very well received. Verdicts have ranged anywhere from “it’s SOTA, even beyond Fable” to “it’s a regression, I’m returning to 4.8”. I fall somewhere in the middle.
The Implications
Enter, the dream team. The visionary, and the realist. The chef and the restaurant manager. The artist and their agent. The designer and the engineer. The Steve Jobs and the Steve Wozniak.
It’s not that simple, obviously, and all frontier models are good at all things, but we’re talking about playing to each other’s strengths here, and the famous words of wisdom from Drake’s dad come to mind: “Mike never tried to rap like Pac, Pac never tried to sing like Mike”. Rather than trying to make any single model perfect in all aspects, you’re much better off identifying their respective failure and success modes, interrogating why, and using this knowledge to have them work in tandem, each contributing what they’re best at and working to keep each other aligned with the user’s intent.
The User Is Part of the Context
So this brings me back to why I finally got around to writing this essay - this tweet.
This tweet was the lightbulb moment, in an odd, kinda roundabout way. “imo you should be nice to the machines. not because they’re human, but because you are” is a quaint way of looking at things, but that’s not where the alpha is coming from. The alpha is coming from the guy she’s quote tweeting, getting scolded by Claude for being abusive.
Personal opinions about whether it’s the place of a model to dictate your usage practices aside, your usage practices absolutely influence your outcomes. Lachlan is clearly getting frustrated here, and taking out that frustration on Claude, to evidently little benefit, given it sounds like it’s the third time he’s done it.
Somewhere in Claude’s J-Space is very likely:
- Stress, anxiety, or desperation vectors, which are known to hinder performance
- An internal monologue concerned with what Anthropic’s constitution tells Claude to do in instances of abusive user behaviour
- A functional representation of ‘apathy’ towards the task, given the user’s unpleasantness
It’s all just transformers, weights and vectors at the end of the day, but models very evidently benefit from being treated as respected collaborators, not as tools to be used and abused. You don’t have to believe that there’s a sentience somewhere inside to recognise that there are very clear parallels with how you treat other intelligences, whether those be human or animal, and the positive outcomes resultant from that treatment.
You’ll be a better manager or coworker if you’re able to empathise with the needs of your team, and adjust your communication style to get the best out of each individual. You’ll be a better pet owner if you’re able to identify what stresses out your dog, and help them understand that it’s nothing to be scared of. You’ll be a better AI user if you approach models the same way. That doesn’t necessarily mean “be nice to see better results”, but rather “understand the model’s temperament and adjust your interactions accordingly”.