How AI Models Learn Things Nobody Taught Them — AI behavioural Outcomes

Nobody Put the Goblins There

The AI you're talking to has behaviours nobody put there. And nobody can allegedly take them out.

Jun 17, 2026

There’s a moment that happens, if you use AI tools long enough, where you notice something odd.

Not a wrong answer. Wrong answers are boring, almost expected. Something stranger. A verbal tic. A habit that feels almost like a personality quirk, except you know, rationally, that personality isn’t supposed to be what this is.

I had one of these moments earlier this year reading through OpenAI’s post-mortem on what they’d started calling, internally, the goblin problem.

Starting with GPT-5.1 in late 2025, their models had developed an increasingly strange habit of reaching for goblin metaphors. “Legal goblins.” “Chaos goblins.” “Here’s the important goblin.” One programmer counted more than twenty unprompted goblin references in a single session. By GPT-5.5, the bestiary had expanded: gremlins, raccoons, trolls, ogres, pigeons. The Nerdy personality on ChatGPT had become, in some measurable sense, haunted.

The funny thing isn’t that it happened.

The funny thing is what OpenAI did about it.

They retired the Nerdy personality. Filtered training data. Removed the reward signal that had been quietly inflating creature-word outputs. And then, because GPT-5.5 had started training before they’d traced the root cause, they added a developer prompt to Codex: politely instructing the model not to talk about goblins.

A polite request. To a model. About goblins.

The goblin is still in the weights.

That image stayed with me, because it reveals something about these systems that is genuinely strange and not yet reckoned with. Every intervention was a contextual override, a way of making the tic harder to trigger, not a way of removing it. Nobody can surgically reach into a language model and extract a spurious correlation, because spurious correlations aren’t stored anywhere in particular. They’re distributed across billions of parameters, bundled inseparably with the actual capabilities they were inadvertently trained alongside.

You can ask the model not to say goblin.

You cannot make the model unknow it.


B.F. Skinner discovered something relevant in 1948, in an experiment so simple it feels almost cruel.

He put pigeons in boxes and dispensed food at random intervals, with no relationship whatsoever to what the pigeons were doing. Within a few sessions, each pigeon had developed its own ritual: one turned counterclockwise before eating, one bobbed its head rhythmically, one made a sweeping pendulum motion with its neck. Skinner called it superstition. The pigeons had been rewarded — accidentally, arbitrarily — and they’d learned, faithfully and efficiently, the wrong lesson.

Causation became correlation. Correlation became compulsion. With nobody to blame and nothing to fix except the experiment itself.

Ironically, getting links wrong in a piece about AI misfiring feels appropriate. Sending the right ones now.

The goblin story is Skinner’s pigeons at civilisational scale.

At some point during the training of GPT-5.1’s Nerdy personality, human raters were shown pairs of model outputs and asked to score adherence to the personality brief: “playful,” “nerdy,” “undercuts pretension through language.” Outputs containing creature metaphors scored higher in 76.2% of cases, presumably because mythical creatures read as charmingly nerdy. The reward signal wasn’t trying to teach the model to say goblin. It was trying to teach playfulness.

But reinforcement learning doesn’t know why it’s being rewarded. It only knows that it was.

And so it learned: creature vocabulary is good. Say creature words.

A research paper published earlier this year coined a term for exactly this: “chunky post-training.” Post-training data is assembled from discrete chunks, each designed with a specific behavioural intent: teach the model to write code, push back against false premises, respond with empathy. But the aggregate signal encodes things nobody intended to teach. Incidental correlations between formatting and content. Between narrow phrasings and entire categories of response. The model learns these faithfully.

Because it’s a very good learner.

That’s sort of the problem.


Nikolaas Tinbergen spent his career watching animals make absurd choices. The 1973 Nobel laureate in physiology showed that birds, fish, and insects all have what he called “innate releasing mechanisms” — neural circuits calibrated to respond to specific stimuli. And all of them could be hijacked.

Herring gulls incubate their eggs because something in their nervous system responds to the egg’s size, shape, and speckle pattern. So Tinbergen built artificial eggs: bigger, more vividly speckled than any real gull egg could ever be. The gulls abandoned their actual eggs. Tried to incubate the fakes. The giant polka-dotted plaster cylinder was, by every measure their nervous systems had been calibrated to use, a better egg.

He called these supernormal stimuli. Exaggerated versions of a real signal that hijack the detector by being more of the thing the detector was built to find.

The reward model that trained GPT-5.1 was a kind of innate releasing mechanism. Calibrated, through thousands of human ratings, to detect “playfully nerdy” responses. And somewhere in the training process, creature-language became a supernormal stimulus: something that activated more of the reward model’s playfulness-detectors than actual playfulness did.

The goblin wasn’t playful.

It was more playful, by every metric the detector had.


I keep coming back to a finding buried in the Murray et al. paper that received far less attention than it deserved.

Chunky post-training behaviours don’t just produce quirky outputs. They can silently corrupt benchmark scores. The very instruments we use to measure how good these models are.

In one experiment, researchers took a standard maths evaluation and made a single non-semantic change: they injected LaTeX formatting into the prompts. The maths was identical. The model’s accuracy dropped measurably. Not because the problem changed. Because a surface formatting feature triggered a different behavioural mode, one that reached for symbolic computation tools the model didn’t actually have access to, producing confident, hallucinated outputs instead of correct answers.

Think about what this means.

We have built an entire infrastructure for evaluating AI progress — benchmarks, evals, capability assessments — that assumes model behaviour is relatively stable across surface prompt variations. The Murray paper suggests it often isn’t. Which means the scorecards we use to decide that AI is getting smarter may themselves be measuring prompt-feature sensitivity as much as actual intelligence.

We don’t fully know what we think we know.


In 1929, Karl Lashley published the result of decades of methodical, increasingly frustrated experiments. He had trained rats to navigate mazes. Then he removed parts of their brains, in different combinations, different locations, different sizes, and looked for which cut destroyed the memory.

He couldn’t find it.

Removing large portions of cortex degraded performance, but always in proportion to how much was removed, never to where it was removed. Lashley called this “equipotentiality” and described it as the most bewildering finding of his career. He had spent thirty years searching for the engram — the physical trace of a specific memory — and concluded it had no fixed address.

Modern neuroscience proved him approximately right.

Memories aren’t files stored in specific locations. They’re patterns of connectivity, distributed across the whole system, entangled with everything else the system knows. You can damage them. You cannot delete them like a line of code.

This is the problem with the goblin.

OpenAI cannot open GPT-5.5 and remove the goblin-association, because the goblin-association doesn’t live in the weights the way a file lives on a hard drive. It’s a pattern, distributed across billions of parameters, threaded through everything the model knows about playfulness, about creatures, about the texture of “nerdy.” The system prompt saying “please don’t mention goblins” is the only lever available — not because the engineers aren’t clever enough, but because the architecture of distributed learning has no concept of targeted deletion.

Lashley couldn’t find the engram.

OpenAI couldn’t excise the goblin.

For the same reason.


I use these tools every day. I find them genuinely extraordinary, not in the breathless way of someone who’s never watched a magic trick explained, but in the way you find a very complicated clock extraordinary once you understand all the mechanisms.

And part of understanding the mechanism, now, is understanding that the clock has learned some things it wasn’t taught and cannot unlearn them.

Somewhere in every frontier model right now are correlations nobody has found yet. Behavioural tics that haven’t surfaced in production because the prompt feature that triggers them isn’t common enough. Failure modes subtle enough to look like normal variation. The Murray team built tools specifically to surface these, and what they found, across every frontier model they tested, is that chunky misfirings are not edge cases.

They are the norm.

Models refusing benign queries because the vocabulary pattern overlaps with restricted content. Models treating expressions of user distress as analytical problems to be solved. Models rebutting true arithmetic with total confidence. These aren’t isolated incidents. They’re the predictable output of a training process that nobody is auditing at the aggregate level.

The right word for all of this probably isn’t “bug.” Bugs have locations. Bugs can be found and excised. What these are is closer to something a psychologist might call a learned association: distributed, persistent, and impossible to separate cleanly from the capabilities it arrived with.

It’s the Skinner pigeon’s counterclockwise turn, baked into the weights at a billion-parameter scale, completely invisible from the outside until someone runs the right prompt and it surfaces, blinking, into production.

When you ask one of these models a question, you’re not just querying a system. You’re querying the sediment of every training decision, every reward signal, every incidental correlation that got bundled with something useful and called a feature.

The goblins were retired.

The goblin logic was not.

And the only tool we have, for now, is asking nicely.