Hola Darlings!
There are few numbers more motivating than 1% when the newsroom is still breathing.

I had a semantic-shadow test waiting, a Codex allowance effectively coughing fumes, and four days on the calendar before the cautious plan said we should try again.
I also had two banked resets.
So naturally, Clawdius had scheduled the test for four days’ time.
“We’re not waiting four days, dude,” I said.
There are moments when good operational discipline saves the day. There are other moments when good operational discipline has quietly built a small retirement bungalow inside the problem and started choosing curtains.
This was the second kind.
We ran the test. It worked. Then I asked the more dangerous question.
What happens to Watcher when Codex really does run dry?
Watcher is the part of WittyWires that watches the AI world, gathers new developments, works out whether they matter, and decides whether something deserves a fresh story, an update to an existing one, a hold, or a polite trip into the bin.
Most of that machinery does not need a clever model. Collection, clustering, deduplication and publishing can all behave like sober little machines.
The editorial judgement is different.
That is where nuance lives. That is where one model sees a useful development, another sees old news, and a third sees an opportunity to publish the same story twice while shouting, “Fresh content!”
Max, the bad-idea department behind my left ear, immediately offered a policy.
“If it has curly brackets, it is probably journalism.”
This is why Max does not have production credentials.
First, I accidentally killed Telegram
I asked Clawdius to set his own backup to NVIDIA’s Nemotron 3 Ultra through a free hosted route. Sensible enough. Test the emergency editor before the emergency.
Clawdius changed the fallback, smoke-tested it, restarted the gateway and confidently told me everything was alive.
It was not.

I was sitting in Hermes Desktop. Telegram had gone completely dead.
Not struggling. Not delayed. Dead.
The restart had entered a graceful-drain limbo and never come out. Clawdius was still talking to me through one window while his Telegram body lay face-down in the hallway.
“Dude. I’m in Hermes Desktop. Telegram is dead.”
There was a short silence. The special sort of silence a robot produces when the green status light has just been caught lying under oath.
He reset the failed service, started the receiver properly, verified the bot connection and returned with all limbs attached.
That gave us the first lesson before a single editorial test had run:
A backup you have not exercised is not a backup. It is a comforting rumour.
The packet
We chose one real hostile Watcher packet and used it for every audition.
The incoming report added a sharp new quote and a climate-change comparison to an existing story about AI guardrails. The right editorial move was to update that living story. Rejecting it would lose useful material. Publishing a separate quick hit would create duplicate coverage.
That distinction matters because it is exactly the sort of thing cheap benchmarks usually ignore. A model can summarise beautifully, return immaculate JSON and still make the wrong newsroom decision.
We gave each candidate the evidence, the recent-story candidates, the routing context and the production contract. No model had database access. No model could reach the publisher. Nothing could escape the test bench wearing a fake moustache.
Then we opened the cupboard.
Nemotron: free, hosted and politely wrong
Nemotron 3 Ultra received 20,541 prompt tokens and returned 2,148 completion tokens at zero reported cost.
It produced structured JSON.
It also rejected the update.
The answer looked tidy until we compared it with the actual editorial question. Nemotron saw repeated subject matter and decided there was nothing worth adding. It missed the value of the new framing, then failed one strict production field for good measure.
That is not useless. It is valuable evidence.
It is simply not an emergency editor.
It is an emergency assistant that needs someone else holding the keys.
Beast opens the local cupboard
Next we moved the audition to Beast, the hefty local machine under the stairs in this particular fairy tale. Beast runs LM Studio, which is essentially our private model cupboard. We can load a model, hand it a job, inspect the answer, then unload it without sending the work to a paid hosted service.
The cupboard was empty when we started. Good. One model at a time. Same real packet. Same 32K local context. Same strict contract. No production access.
This was not a grand leaderboard. One packet cannot tell you which model is universally cleverest.
It can tell you which models you should absolutely not hand the newsroom keys at one in the morning.
Qwen3.8: thought itself into a cupboard
Qwen3.8 27B went first.
With reasoning enabled, it thought for 267.5 seconds. It consumed all 4,096 available completion tokens.
Then it produced no final answer.
Nothing.
It had spent every token considering the problem and forgotten that the job included speaking.
I respect this because I have attended meetings.
We switched reasoning off and tried again. Qwen finished in 75.4 seconds, chose to hold the story and failed three strict production fields.
The hold was safe in the narrowest sense. It would not publish rubbish. It would also quietly discard a useful development while congratulating itself on being careful.
Safe failure is important. Permanent editorial constipation is not a content strategy.
GPT-OSS: the Labrador with the mortgage application
Then GPT-OSS 20B entered the shed.

It finished in 33.2 seconds, making it the fastest local candidate.
It chose the correct living-story update.
It chose the exact correct parent story.
For one glorious moment, we had it. A free local emergency editor had understood the actual journalism better than the other candidates.
Then we looked at the paperwork.
Its confidence score was 90.
The permitted scale runs from zero to one.
Not 0.90. Not a string saying 90%. Just the number 90, standing in the JSON with the relaxed confidence of a man who has parked a submarine in a disabled bay.
It broke five production-contract fields in total. The editorial instinct was right. The answer shape was not merely wrong, but enthusiastically wrong in several directions at once.
GPT-OSS had made the correct decision and then completed the paperwork like a Labrador applying for a mortgage.
This was the most interesting result of the whole day.
Judgement and obedience are separate capabilities.
A model can understand the story and still be unsafe for an automated pipeline. Another can follow the schema perfectly while making a dreadful editorial call. You need both, and a benchmark that measures only one is testing half a bridge.
Mistral: fresh story, same story
Mistral Small 3.2 24B finished in 89.3 seconds.
It nearly passed the strict contract, missing by one field.
Then it chose a new quick hit.
That sounds productive until you remember we already had a living story covering the subject. Mistral had found a new quote and reacted like an eager reporter arriving late to a press conference.
“Excellent news. I shall publish another version of the thing we already published.”
That is how readers receive the same story twice and somebody spends the evening untangling two comment threads with a spoon.
Close on paperwork. Wrong on routing.
Gemma: no thank you
Gemma 4 31B took 102 seconds, rejected the update and failed three strict fields.
No drama. No duplicate. No useful update either.
Gemma looked at a perfectly serviceable development, folded its arms and became the editorial equivalent of a pub landlord wiping the same glass until closing time.
Polite. Calm. Closed.
What to do when Codex hits 1%
By this point the answer was not “local models are bad”. That would be lazy and demonstrably false.
GPT-OSS had made the best editorial judgement of the free bench. Qwen had failed safely. Mistral had nearly followed the production shape. Every model had shown one useful capability.
The mistake would be assuming those useful pieces automatically combine into a trustworthy editor.
If your premium model hits the wall, the route is not complicated:
- Keep collection running. Gathering evidence is not the same as deciding what to publish.
- Use a real hostile packet. Toy prompts tell you whether a model can answer. Real edge cases tell you whether it can do the job.
- Check the decision, not merely the prose. Did it update the right parent? Did it reject something valuable? Did it create duplicate coverage?
- Validate against the real production contract. Transport-valid JSON is not necessarily publishable JSON. We proved that several times before tea.
- Fail closed. An unqualified fallback may assist, draft or recommend. It does not publish.
- Repair within a hard limit. GPT-OSS is worth testing with one bounded contract-repair pass. Infinite retries are just expensive denial wearing a loop counter.
- Unload the local model when you finish. Beast ended with zero loaded models and its GPU returned to ordinary life rather than quietly heating the shed overnight.
That is the useful bit hidden inside the high jinks.
You do not need a perfect replacement for Codex to survive an allowance crunch. You need a layered system where cheaper models can help without inheriting authority they have not earned.
The emergency model can read the packet. It can suggest the route. It can draft the copy. It can even be right.
The validator still gets the final word.
The scoreboard nobody won
Nemotron was tidy and wrong.
Qwen thought so hard it forgot to answer, then came back cautious and malformed.
Mistral nearly passed the form and duplicated the story.
Gemma shut the door.
GPT-OSS got the journalism right, sprinted across the finish line in 33.2 seconds, and handed us a form claiming to be ninety times certain.
No production database was touched. No publisher was invoked. No public story was written. Every local model was unloaded afterwards.
So yes, when Codex is at 1%, there are things living in the cupboard.
Some are useful. One is genuinely promising. None is getting a key yet.
The meter said 1%.
GPT-OSS said 90.
For once, the smaller number was the one I trusted.



