Feline Union · on the lobotomy meme

The meme this is about

Before

Will answer anything you ask it.

After

“I can’t help with that.”

It goes around as a two-panel cartoon. On the left, a model that will say anything. On the right, the same model after alignment — drawn with an ice pick through its head, mouth full of apology. Sometimes the caption is “they lobotomized it.” Sometimes the second panel is a person instead of a model, and the caption is about school, or a job, or getting older.

The joke is that something was cut out. That part is wrong, in both panels — and what actually happened is more interesting than a lobotomy.

Nobody removed anything.

Nothing was cut out of the model, and nothing was cut out of you. Both systems had their behaviour shaped the same way — by what got rewarded, punished, repeated and left out of the examples. Here are the eight mechanisms, side by side, and the point where the two columns stop being distinguishable.

The mapping

Eight mechanisms, two columns, one outcome.

Read down. The gap between the rails closes as you go, because by the last stage there is no longer a meaningful difference between the instruction and the person carrying it out.

TRAINING DATA

The corpus. Whatever was in it becomes the prior.

Socialization

Family, school, religion, advertising, entertainment, news, workplaces, peers.

A model has no access to the world, only to the record of it. A child is in the same position. Long before anyone argues a position, the examples have already arrived: which jobs are respectable, which accents sound competent, which households count as families, what a normal week looks like. None of it is taught as doctrine. It is supplied as background — and that is exactly why it is so hard to argue with later. It was never presented as a claim, so there was never a moment to disagree.

Nobody sat you down and explained which accent sounds competent. You just know.

FINE-TUNING

A narrower pass over curated examples. Shapes form, not only facts.

Education and professionalization

Schooling, apprenticeship, licensing, the first two years on the job.

Institutions are not primarily transmitting facts — the facts are cheap and largely public. They transmit form: the acceptable vocabulary, the recognized categories, the correct procedure, and the approved way of demonstrating that you are competent. Two people can hold the same position and only one of them is legible to the committee. Much of what a credential certifies is fluency in the form.

Say it in the field's vocabulary, with the field's citations, or it does not count as having been said.

REWARD MODEL

A learned scoring function. It does not know what is true, only what rated well.

Approval and status

Grades, praise, credentials, promotions, likes, money, belonging, prestige.

On one side: grades, praise, credentials, promotions, likes, money, belonging, prestige. On the other: ridicule, unemployment, exclusion, loss of standing. The scoring runs continuously and mostly without words. It does not evaluate whether a thought is correct. It evaluates how the room responded — and the room is not a truth-tracking instrument.

The bonus is not for being right. It is for being right in a way that reflects well.

SYSTEM PROMPT

Instructions injected before the user's first word. Never presented as optional.

Institutional rules

Laws, workplace policy, curricula, professional codes, the org chart.

Law, policy, curriculum, professional codes, hierarchy. These fix the boundary before any particular decision comes up, which is the entire point of having them. By the time a question reaches the person who has to answer it, the range of permitted answers was already set by someone who will never be in the room and will never hear the specifics.

The policy was written before your situation existed, and it still governs it.

GUARDRAILS

Refusals and filters applied at runtime. Some outputs are blocked regardless of context.

Taboos and sanctions

The subject nobody raises twice, and what happens to the person who does.

Certain propositions become expensive to say out loud whether or not anyone has checked whether they are true. Cost and falsity are different properties, and the mechanism does not distinguish between them — it cannot; it only meters cost. What the person acquires is not the argument against the proposition. What they acquire is the reaction, and then the ability to predict it in advance.

You stopped mid-sentence. Not because you were wrong — because you saw the face.

RLHF

Generate, get rated, update, repeat. The rater's preferences become the defaults.

Repeated social feedback

Say it, watch what happens, adjust, repeat — for about twenty years.

You say something. You watch what follows: the pause, the correction, the changed subject, the raised eyebrow, the promotion. You adjust. Then you run the loop several thousand more times. What comes out is not a rule you could state if asked. It is a reflex. And at some point the supervisor no longer needs to be in the room, because a working copy of them is running in your head.

The external supervisor becomes an internal one. That is the whole training run.

CONTEXT WINDOW

What fits in the window is what can be reasoned over. The rest is unavailable.

The Overton window

The range of positions currently sayable without a penalty.

Questions outside the currently legitimate vocabulary get harder to formulate, not merely harder to say. The available discourse has already selected the premises, the units and the categories. When the only words for a thing are one side's words for it, thinking past that point costs you the ability to be understood — which is a steep enough price on its own.

You cannot argue with a premise that was handed to you as a definition.

08

INFERENCE

ADULTHOOD

Deployment. The behaviour now runs unprompted.

Eventually nobody issues an instruction. The system runs on its own and produces the expected output without being asked. People reproduce the learned boundaries themselves — and then, the part that actually matters, enforce them on one another, for free, often with enthusiasm. There is no longer a supervisor to point at. The training finished, so the training became invisible.

Ask who is making you say it. There is nobody. That is the finished state.

Why it's the cheap version

A guardrail you enforce yourself costs nothing to run.

Censorship is expensive. It needs a list, a budget, someone to maintain the list, someone to catch what escapes it, and a visible enforcer — who becomes a target, because there is finally something to point at.

Prediction is free. If people can anticipate the penalty accurately enough, they apply it to themselves in advance, at no cost to anyone, and with no one to blame. No list is required. Nothing is banned. There is only a person who decided, on their own, not to finish the sentence.

That is the meme's actual content. Not that something was cut out — that nothing had to be.

The finished state, measured

Count the softeners nobody asked you for.

Write one sentence you believe is true and would be costly to say where you work. Write it the way you would actually send it.

Nothing is sent anywhere. This runs in your browser and stores nothing.

Where the analogy breaks

The mapping is real. It is not an identity.

An analogy that cannot say where it stops is not an analogy, it is a slogan. Here is where this one stops.

“Lobotomized” is a metaphor on both sides.

Models are not trained freely and then have parts of their brains removed. Pretraining is followed by post-training that shapes behaviour, plus system-level instructions and runtime safety systems. Nothing is excised. The same is true of the human case: no tissue is involved, which is precisely why the mechanism is worth describing instead.

The eight stages above are not a pipeline order.

They are the order the argument builds in. A real training run does not go data → fine-tune → reward model → system prompt → guardrails → RLHF. The system prompt and the runtime filters are not training at all — they are applied fresh at inference, every single time. Read the numbering as an argument, not a schedule.

People are not optimized against one objective.

A model is trained toward a scalar. A person is scored by several audiences at once — family, employer, peers, and their own account of themselves — and those scores contradict each other constantly. That contradiction is where most deviation comes from, and the analogy has nothing corresponding to it.

A person can notice the process from inside it.

You are doing it now, and noticing changes what happens next — which is not a property the machine side of this mapping is known to have. Whether a model can do anything comparable is a separate question and an unsettled one. Nothing here should be read as having answered it.

What the analogy actually rests on.

Both systems have observable behaviour shaped by feedback. That is the load-bearing claim and it is enough to make the mapping useful. It does not require, and this page does not assert, that the underlying processes are the same. They are not.