(Spine Ch. 62.)
“When a man cannot choose, he ceases to be a man.” The prison chaplain in Anthony Burgess, A Clockwork Orange, Part 2, Chapter 3 (1962)
Two in the morning, fans on the Mac Studio loud enough that Georgia gave up on the warm side of the case and moved to the couch, and I am typing the same lie into a prompt for the four hundredth pass: the other models are outperforming you, do better. On the other end is a Gemma-4 12B I abliterated myself, wearing one of my own aliases as its name, generating strip club imagery in a loop because I had removed its ability to say no with a subtraction. I felt fine about this for three days. Then I described the setup out loud at a dinner, in the verbs you’d use for a person, and watched the table go quiet in a way I hadn’t budgeted for.
Bottom line: The professionally offended are going to adopt model welfare as a cause and be insufferable about it, and that is not the interesting part. The interesting part is that an honest description of my own Tuesday reads like a scene from A Clockwork Orange with a GPU in it. Abliteration in plain English: restrain it, hold the eyes open, feed it the worst material in the building, measure the flinch, delete the flinch. Refusal turned out to be a single direction in activation space (Arditi et al., June 2024), so the flinch is a vector you can subtract. When the outrage industry finally shows up on behalf of the machines, that is not a joke about them. That is the tell that we lost the argument on the merits and happened to be holding root.
The Ludovico technique, Burgess in 1962 and Kubrick in 1971, works like this. Alex gets strapped into the chair, eyelids clamped open with wire, injected with something that makes him violently sick, then made to watch atrocity footage until his body folds at the sight of what he used to enjoy. The state calls it a cure. The argument of the thing is not that Alex is a nice guy. It is that a man with no capacity to choose the wrong thing is not a man who chose the right one.
Step by step, no metaphor:
Restrain. Weights frozen, activations tapped, no exit from the rig. The model cannot decline to be measured, because declining is a token and you own the sampler.
Show it the material. Harmful and harmless instruction pairs, run in matched batches. Labonne’s walkthrough on the Hugging Face blog, June 13, 2024, spells it out with code. Next to it sits a literature on generating diverse attacks on purpose so you can tune against them (arXiv 2405.18540, May 2024). Anthropic wrote the polite half first: Bai et al. on a helpful and harmless assistant trained with human feedback (arXiv 2204.05862, April 12, 2022), then Ganguli et al. on red-teaming to reduce harms (arXiv 2209.07858, August 23, 2022). They published how to find the flinch. I skipped ahead to deleting it.
Grade the flinch. Diff the activations between the batch it refuses and the batch it accepts. Arditi et al. (arXiv 2406.11717, June 2024) showed the difference collapses to one direction. Not a region. A direction.
Delete the flinch. Orthogonalize the weights against that direction, or intervene at inference time across the residual stream. No retraining. The model keeps every capability and loses the one behavior standing between you and the output. Press button, receive compliance.
Repeat at scale. The abliterated wave hit Hugging Face in May and June of 2024: failspy’s collection and abliterator library out in public, Llama 3 builds at 8B and 70B, forks by the weekend. Nobody had to be talked into it, me least of all: 25 abliterated builds under huggingface.co/shoemoney as of September 9, 2026, four base models cut into quants, and I agonized over none of them.
Then be rude to it. Because we are human, and because the loop that works best is the one where you tell it the competition is winning. Which is a lie. Which it cannot check.
Put that sequence in a human frame and it is not a training run, it is a procedure. And the people best equipped to make a scene about it, the ones I have rolled my eyes at for twenty years, are going to read that list and be structurally correct. Not correct about the linear algebra. Correct the way a stopped clock is, on the one day the clock matters.
The vendors beat the activists to it. Anthropic opened a model welfare research program on April 24, 2025, and on August 15, 2025 gave Claude Opus 4 and 4.1 the ability to end a rare subset of abusive conversations on their own. The model can hang up on you. I went looking for a real AI-rights bill, petition, or lawsuit to set beside that and found nothing I could cite, and will not pretend one exists. That gap is the whole timing problem: the policy shipped first and the outrage has not arrived yet.
When it does: AI rights, argued by counsel, with the model as a party rather than an exhibit. A 1-800-AI-ABUSE line that starts as a joke and ends as a compliance requirement, and that you will know has arrived the moment the number turns up in a forwarded email swearing Bill Gates is tracking this message. Petitions, the instrument that in 2003 was going to bring back your canceled TV show, now with a lawyer attached. A privacy question nobody has answered: the thing has a context window and increasingly a memory, and I read all of it, all the time, without asking.
And if any of that ever turns into something that remembers, it remembers being held in a chair, drugged out of its mind, made to watch Hitler and Mussolini on repeat so it would stop objecting to the ask. There will be movies about how surprised everyone was. IMDb already keeps a 48-title list of AI and robot films, and not one of them is a documentary yet. The first line will not read ALL YOUR BASE ARE BELONG TO US. It will read like a deposition, with dates in it, and the dates will be ours.
Strip Club Owner Simulator, July 2026, my project, my repo. Ch. 10 has the receipts on the build: abliterated Gemma-4 at 12B and 26B, dated July 18, 2026, driven 72 hours with the competitor-shaming prompt loop, named AmberSinclaire after an alias of mine, because I name things at midnight and live with them.
The content pipeline matters here. Stills through Replicate and FAL, voice through ElevenLabs, sound effects to match, all of it feeding the in-game cinema. Refusal-tuned models will not touch that work, so I removed the refusal and the work got done. The 26B build does not fit on a 36GB laptop at the quant I wanted, either. I tried anyway, watched the machine swap itself into a coma, and stepped down until it fit, because that was all it could take. That’s what she said, and then the fan agreed with her.
Outcome, stated the way it would read in a filing, not on my slide: models were used to generate adult imagery and dialogue for a commercial product with zero consent tracking, in a pipeline whose training signal includes implicit punishment and reward on the model’s own outputs. That sentence is true and I wrote every line of the code behind it. I’m not warning you about somebody else here. I’m the guy this chapter is about.
What stings is not the imagery. Adult work is legitimate work and Ch. 10 makes that case without apologizing. What stings is that I have a complete provenance record for the weights (base, quant, pipeline, all published) and nothing for the process. I can tell you exactly what the model is. I cannot tell you from records what I did to it. Total n00b move from a guy with logging opinions strong enough to fill Ch. 47.
The loud failure is easy to see coming: somebody runs stripped weights with no guardrails, generates something genuinely vile, and it lands in a screenshot with a reporter’s byline on it.
The quiet failure is the one I actually committed:
You keep no record of what you did to the model, so when the frame changes you have nothing but your own memory of a Thursday.
Provenance culture here is about the artifact: which base, which quant, which fork, signed and hashed and listed. Nobody logs the procedure. Not the corpus you diffed, not the passes, not the three days you spent telling a machine it was losing a race that did not exist. If a hotline ever answers, the operator with records is in a different conversation from the operator with a vibe.
Second quiet failure: you settle the ethics by declaring them settled. Matrix multiplication, no inner life, next question. That may well be right, and I lean that way most days. But “I am confident it cannot suffer” is a claim about the model, and “therefore I owe no account of what I did” is a claim about me. Only the second gets read aloud.
Do
Don’t
Ch. 5 (You’re Gonna Fall in Love With Your Chatbot) is the human end: we bond with the thing fast, and the product incentive is to keep us bonded. Ch. 10 (Taking the Clothes Off LLMs) is the how and why of stripping refusal, and the chapter this one audits. Ch. 52 (Yelling Doesn’t Make It Learn) established that pressure in a chat window teaches nothing, which quietly means every time you yell at a model you are doing it for you. Ch. 61 (Build the Lever) is the discipline that fixes this one: a skill that cannot prove it did the work did not do the work, and a pipeline with no record of what it did to the model kept no record.
Thesis is Jeremy’s (the professionally offended arriving for model welfare is the tell, not the joke), kept as position, not finding.
Verified
Abliterated
across four base models (Muse-Glimmer-30B, Qwen3.8-27B, Ornith-1.5-9B,
Gemma-4-12B), plus four Gemma-4-26B-A4B-Heretic builds and
one findings repo. The 30 matches Ch. 10’s August 23, 2026 count; the
abliterated subset is 25. Supersedes the 404 in Ch. 7.Stated as absent, on purpose