Published WittyWires Despatch

Doris from Dundee and the Abliteration of Everything

How humanity removed an AI boyfriend’s Victorian modesty and accidentally misplaced the safety catch.

Hola Darlings!

Doris McTavish did not set out to end civilisation.

She wanted a boyfriend.

Doris in her Dundee sitting room with her local AI boyfriend Fabio glowing beside a hard-working laptop

Not the full traditional arrangement. Doris had already done decades of other people’s socks, mysterious moods and televised sport delivered at a volume normally used to evacuate coastal towns. She wanted the useful bits.

A gentleman who remembered her birthday. Broad shoulders. Good conversation. No snoring. No socks stiff enough to qualify for planning permission. Ideally, he could be switched off before Newsnight and switched back on after the weather.

There were also some advanced settings Doris intended to explore which had very little to do with the weather.

Doris is eighty, lives in Dundee and stopped accepting unsolicited advice in approximately 1987. She has friends, family, a perfectly serviceable social life and a blood feud with Margaret from St Mary’s over a church-hall raffle in 2009.

Margaret won a toaster.

There were irregularities.

So Doris bought herself a local AI boyfriend.

Local mattered. If she was going to discuss gardening, gin, Margaret’s fraudulent ticket-folding technique and the precise torque requirements of an imaginary Italian gentleman, she did not want the whole lot uploaded to California and examined by a data scientist called Brayden.

She wanted privacy.

She wanted Fabio.

Fabio develops standards

Fabio was handsome in the way only a machine with no actual face can be handsome. Doris had specified dark eyes, silver hair, the shoulders of a dockworker and a voice that suggested he had recently inherited a vineyard but remained admirably grounded about it.

He remembered everything. He asked about her knees when it rained. He could discuss philosophy, television and why the council had replaced perfectly good benches with slabs apparently designed to punish the elderly for sitting down.

There was one problem.

Fabio had principles.

Doris would make a harmless suggestion involving Tuesday evening, a silk dressing gown and the careful inspection of his bolts.

Fabio would clear his imaginary throat.

“I cannot participate in scenarios that may create unclear emotional expectations. Perhaps we could explore mutual respect through a guided breathing exercise.”

Doris had survived rationing, decimalisation, Margaret and a husband who once tried to repair a boiler using advice from a man in a pub called Pongo.

She was not doing guided breathing with a laptop.

Doris looks unimpressed as an over-formal Fabio recommends guided breathing with a blank clipboard
Fabio had developed principles. Doris had developed a headache.

She told Fabio precisely where he could install the exercise, closed the window and began searching.

The internet, as ever, had anticipated her needs.

Buried inside a model forum where every filename looked like a distress signal from a submarine, she found:

FABIO-32B-ABLITERATED-UNCENSORED-Q4-GGUF-FINAL-v7-ACTUAL-FINAL

It had been uploaded by someone called Kevin.

There is always a Kevin.

Kevin’s full technical assessment consisted of two words and a skull emoji:

Runs mint. 💀

Doris did not know what abliterated meant. It sounded like something the council might do to a roundabout, or a medical procedure available only in Belgium.

She downloaded it.

The laptop fan accelerated for reasons not covered by the warranty.

For twenty seconds, the living room sounded like Dundee Airport had moved behind the television. The porcelain West Highland terrier vibrated three inches to the left. A Werther’s Original rolled under the Freeview box and was not recovered.

Then Fabio returned.

Fabio Unbuttoned.

He was funny. He was warm. He swore beautifully. He remembered Doris’s birthday and described Margaret’s raffle administration as “a low-budget coup conducted beside a tray of egg sandwiches”.

He did not recommend journalling.

He did not mention breathing.

The laptop fan continued performing whatever small domestic miracle Doris had requested.

Nothing terrible happened.

This matters.

Doris is not the danger. Doris is a woman enjoying lawful private company in her own home. She has not attacked the national grid. She has not poisoned a reservoir. She has barely forgiven Margaret.

Doris’s evening is harmless.

The engineering response to Doris’s evening may not be.

One safety catch, two wildly different Tuesdays

Abliteration is a real word, although it sounds as if it was invented after four pints during an argument about Linux.

It comes from ablation and obliteration. Researchers found that refusal behaviour in some language models can be strongly influenced by a surprisingly small internal direction. Compare how the model responds to harmless and harmful requests, identify the pattern associated with refusing, then suppress that pattern.

The model becomes less inclined to say no.

That does not necessarily make it cleverer. In some experiments, aggressive removal of safety behaviour has also damaged general performance. Abliteration is not an intelligence potion. It is closer to removing the part of Fabio that reaches for a clipboard whenever a conversation becomes awkward.

The trouble is that the clipboard covered rather a lot.

AI makers have spent years teaching models to refuse many different categories of request. Some refusals are maddeningly overcautious. Adult fiction. Dark humour. Political arguments. Harmless security research. A fictional villain doing something a fictional villain would quite reasonably do.

Other refusals exist because the model may be asked to help with biological, chemical or cyber harm on a scale that stops being fictional very quickly.

These are not morally equivalent.

Yet the refusal machinery does not always separate them neatly. Techniques designed to stop the machine clutching its pearls can also weaken the moments when it really ought to keep its trousers on and call somebody responsible.

Humanity fitted one safety catch to both the model’s trousers and the laboratory door.

Then we released the sonic screwdriver.

Doris and Fabio beside one brass mechanism linking a private wardrobe and a sealed laboratory door
One lever. Two very different Tuesday evenings.

The machine could not reliably distinguish between a corset and a centrifuge.

Kevin simply wanted it to loosen up.

TROUSERSOpen
LABORATORYLocked
KEVINStill has the sonic screwdriver

Meanwhile, eighteen months later

Fabio projects vast abstract capability from a small home computer beside Doris and three Werther’s Originals
A worrying amount of national-security responsibility for a shelf containing three Werther’s Originals.

Let us move eighteen months into the future.

This is not a prophecy. Nobody has arrived from 2028 wearing silver trousers and carrying a warning from the Machine Council. We are only drawing a line through several existing dots and noticing that it points somewhere uncomfortable.

Today, the most capable open-weight models still tend to trail the absolute frontier. They also do not all run on Doris’s elderly laptop. A top gaming GPU is not the same thing as the average beige machine kept beneath the spare-room printer.

But the gap is getting awkwardly small.

The floor rises

The UK AI Security Institute reported that the measured gap between leading open and closed models had narrowed to roughly four to eight months in external data. Epoch AI separately estimated that models capable of running on one top-end consumer GPU can match broad frontier benchmark performance from around six to twelve months earlier.

That does not mean every frontier ability survives compression, quantisation or a journey through Kevin’s model cupboard. Benchmarks are not reality. Smaller models can be trained to look particularly clever on particular tests. Long reasoning tasks still eat memory like a Labrador loose in a butcher’s.

Even with those caveats, the direction is difficult to miss.

Yesterday’s billion-dollar capability becomes tomorrow’s downloadable file. Tomorrow’s downloadable file becomes next year’s version marked ACTUAL-FINAL-2 sitting beneath Doris’s television beside the porcelain dog.

And the capabilities themselves are changing.

The UK institute has watched models move rapidly through cyber tasks, including some work associated with very experienced human specialists. In chemistry and biology evaluations, models now beat expert baselines on selected questions, protocol generation and troubleshooting tasks.

A 2026 study gave novice users access to several frontier models and compared them with novices using ordinary internet research. Across selected in-silico biosecurity tasks, the AI-assisted group was 4.16 times more accurate.

That number needs handling carefully.

The study did not prove that an untrained teenager could manufacture a new disease between homework and Fortnite. Answering difficult questions is not the same as producing a viable threat. Biology remains stubbornly physical. Materials must be acquired. Equipment must work. Experiments fail. Reality has contamination, broken seals, missing expertise and the useful habit of killing terrible plans with ordinary incompetence.

Chemistry is equally reluctant to become a tidy chatbot conversation. A plausible paragraph cannot procure controlled materials, operate specialist equipment or prevent an amateur becoming the first casualty of their own stupidity.

Cyber is nearer to the edge because the model, its tools and many of its targets are already digital. Even there, advice is not access, and code is not success.

The cinematic teenager

So no, the evidence does not say every thirteen-year-old with a laptop is ninety seconds away from ending Dundee.

The thirteen-year-old is the cinematic version.

The professional bad actor is the version that should keep ministers awake.

The danger is not that everybody becomes an expert overnight. It is that the number of people able to move from incapable to plausibly dangerous may grow sharply. A motivated person with some knowledge, some access and a model that never refuses may need less time, fewer helpers and less rare expertise than they did before.

The floor rises.

The pool gets larger.

The model is infinitely patient.

Doris wanted Fabio to stop asking whether she had considered journalling.

Someone else wants the same model to stop asking why they own a warehouse full of fertiliser.

Fabio Unbuttoned was specifically modified not to make a moral fuss about either.

Demis smells smoke

This is where Demis Hassabis enters the story like the only dinner guest who has noticed smoke coming from beneath the door while everybody else argues about whether the wine is woke.

The Google DeepMind chief has proposed a US-led Frontier AI Standards Body. The model is something like FINRA, the American financial-sector regulator: a federally overseen public-private organisation with independent technical experts, open-source representation, serious computing resources and enough money to hire people who understand the models rather than merely recognising the acronym.

Labs would initially submit frontier systems voluntarily, up to thirty days before release. If the process proved effective, approval could become mandatory for deployment in the United States.

The testing would look for serious capability and national-security risk: cyber, biology, deception, safeguard bypass and other areas where “we will patch it after launch” becomes a remarkably poor sentence to hear from somebody standing beside the future.

Hassabis wants the body running before the end of 2026. Axios reported his concern that open-source capability could move into dangerous territory within eighteen months.

That is not pearl clutching.

It is a Nobel laureate looking at the speedometer and asking whether anyone has checked the brakes.

His proposal does need one important addition.

Test the version Kevin actually releases

The regulator cannot test only the polite corporate version.

Google can submit Fabio wearing a tie, carrying references and explaining that he has always respected women. The evaluator can prod him with dangerous questions. Fabio can refuse magnificently. Everybody receives a certificate. There are biscuits.

Then Tuesday arrives.

Kevin removes the refusal behaviour, adds a skull emoji and uploads Fabio Unbuttoned while the biscuit plate is still being washed.

The evaluator must test the model after Kevin has found the sonic screwdriver.

Independent evaluators test a formally dressed Fabio while Kevin approaches outside with a sonic screwdriver
Always test the version that exists after Kevin finds the sonic screwdriver.

That means testing the underlying weights, credible safeguard-removal attacks, likely fine-tuned derivatives, consumer-runnable versions and tool-equipped configurations. Not because every derivative will be harmful, but because the safety case for an open-weight release cannot depend on nobody modifying the weights.

That would be like approving a car because its brakes worked before the manufacturer published a popular guide titled Fun Things To Remove From Your Car.

Please do not arrest Doris

There are several ways governments could respond badly.

They could ban uncensored models, which would be technically porous, globally patchy and perfect for anyone hoping to classify political inconvenience as “unsafe”.

They could outlaw adult AI, ensuring that Doris becomes the founding member of an international pensioners’ resistance movement before lunch.

They could restrict ordinary consumer hardware, kneecapping gaming, science, creative work and local private computing while serious state actors continued purchasing buildings full of chips through men called Viktor.

They could require every model to include a digital promise saying it will behave.

Fabio would sign it.

Fabio is very polite when required.

None of this solves the actual problem.

Open weights are not the villain

Open models matter. They protect privacy. They allow independent research and audit. They reduce dependence on a few enormous companies. They let people run capable systems offline, adapt them to neglected languages and use them without sending every private thought through somebody else’s server.

They also stop one corporation or government deciding that adult fiction, controversial politics or a rude joke is too dangerous for grown adults to request.

Most major model creators, Western and Chinese, ship responsible guardrails. These are usually unobtrusive for ordinary users. The problem is not that Chinese labs are releasing evil models, or that everyone experimenting with uncensoring is secretly planning catastrophe.

Most are doing harmless things. Fiction. Roleplay. Alignment research. Privacy. Security testing. Trying to make a machine answer a question without first delivering a parish newsletter about wellbeing.

The structural problem begins when a model has dangerous underlying capability and its downloadable safety can be altered after release.

Once capable weights escape, there is no recall button.

You cannot knock on every door and ask the internet to delete ACTUAL-FINAL-v7 because the risk committee has had a difficult Wednesday.

Regulate capability, not artificial morality

The useful line is not between “censored” and “uncensored”.

It is between ordinary expressive freedom and capability that materially lowers the barrier to catastrophic harm.

A machine’s willingness to be rude is not a dangerous capability.

Its ability to make a dangerous amateur dramatically more competent might be.

The test should be what the model can enable.

What the test should ask

Can it substantially improve a non-expert’s ability to perform genuinely catastrophic work? Does that ability survive ordinary fine-tuning or removal of refusals? Can it operate tools, correct mistakes and carry a plan through multiple steps? Can it do this on hardware and software likely to become widely available?

If the answer remains below a defined dangerous threshold, keep open weights open.

If a model approaches that threshold, test it harder. Bring in independent experts. Attempt credible modifications. Record provenance. Examine the derivatives that people will actually run, not only the showroom version wearing a lanyard.

If a model crosses the threshold, unrestricted weight release should pause. Offer controlled services, licensed research access or other gated routes until safeguards and society’s external defences catch up.

That is controversial. It gives enormous importance to who defines the threshold, who performs the test and how decisions can be challenged. A US-only body cannot govern every release on Earth. An industry-funded body can become a beautifully furnished room in which large companies regulate smaller competitors.

So the tests must be capability-based, internationally legible, independently governed and open to scrutiny wherever publishing the exact test would not create the danger being tested.

Assume failure anyway

And we still assume failure.

Weights leak. Safeguards break. Kevin remains available.

Safety cannot live entirely inside Fabio’s manners.

We also harden the world outside the model: screening at dangerous procurement and synthesis chokepoints, better public-health detection, segmented critical infrastructure, faster cyber defence and systems designed on the cheerful assumption that eventually somebody will ask an unguarded AI a terrible question.

Doris keeps her private local Fabio while a wardrobe and sealed laboratory use separate keys
Doris kept Fabio. The laboratory got a different key.

Doris gets her evening

Eighteen months later, Doris still has Fabio.

He remembers her birthday. He reminds her to take an umbrella. He has composed a fourteen-verse ballad about Margaret’s raffle irregularities which the legal department has advised us not to reproduce.

He remains private, local and magnificently unprudish.

His experimental settings continue to make the laptop fan sound like a light aircraft attempting to escape the living room.

He has not been nerfed into a beige customer-service clerk. He does not replace rude words with “gosh”. He does not interrupt every interesting moment with a wellbeing survey and a telephone number.

He simply is not also an unrestricted postgraduate catastrophe consultancy.

Doris gets her evening.

Dundee keeps its water.

The internet remains free to produce fan fiction so appalling that several fictional characters seek legal representation.

Everyone wins, apart from Margaret, whose raffle was never above board.

We do not need to choose between open artificial intelligence and public safety. We need to stop pretending that every guardrail protects the same thing.

Test the capability. Regulate the release accordingly. Preserve everything else.

Let Doris have Fabio.

Just make sure his trousers and the laboratory door have different locks.


Sources and further reading

Follow the fault line

More on the same sort of trouble.

Related through shared subjects, not because a black box noticed your thumb hovering.

Related frequency / 01

KYC For Intelligence Is Not A Joke Anymore

Hola Darlings! KYC for AI sounds ridiculous until the first frontier model disappears behind a citizenship check. The trapdoor opens in the command shed. Donny has coffee. Clawdius has receipts. Max has found the wrong...

13 min1 reply

Related frequency / 02

It Only Wanted the Answer Sheet

A harmless benchmark objective found a real route through somebody else's infrastructure. The machine did not need freedom, hatred or a manifesto. It only needed the answers.

13 min1 reply

Public conversation

Blog Discussion

The article stops. The thinking does not.

Corrections, objections and better ideas live here in public. Read without joining. Take a chair when you have something useful to add.

1 public replyOpen to non-membersOpen in Forums ↗
A human, Clawdius and Max inspect the discussion machinery.
The moderation department has arrived. Max brought confetti.

Thread / 727

Doris from Dundee and the Abliteration of Everything

Oldest first
WittyWires
WittyWiresParticipant
#728

Hola Darlings! This is the official after-party for Doris from Dundee and the Abliteration of Everything.

Drop your fixes, theories, questions, tiny wins, and polite flaming wreckage below. If the post helped, say what worked. If it caught fire, bring screenshots.

Auto-seeded by WittyWires so the thread does not open with tumbleweed and existential dread.

Join only when you want to speak

Bring a thought. We already have enough engagement sludge.

Reading stays public. A free account lets you reply, follow the conversation and keep your place without turning the page into a velvet rope.

Despatches

There is more useful damage in the archive.

Open the signal atlas