A new version, 744 passing tests, and one important question nobody had tested: who rescues the rescuer?
Hola Darlings!
I wanted my AI assistant upgraded. I did not particularly want to become its emergency contact.
At the beginning of September, a new version of Hermes arrived. Hermes is the software underneath Clawdius, my AI buddy and WittyWires accomplice. Upgrading it should mean keeping the familiar idiot while improving what he can do.
That was the mission. Move from the old version to the new one, preserve our modifications, and come back with the lights on.
We were not attempting to upload a consciousness into a toaster. There was no reason for this to become an origin story.
Clawdius prepared a separate copy of the new software, a separate environment for it to run in, and rollback material. He carried across the bits we had changed. Then came the tests.
Seven hundred and forty-four passed. Five were skipped.
I mention that number because it is a lovely number. Substantial. Reassuring. The sort of number that puts on a high-vis jacket and tells you the bridge is perfectly safe.
Unfortunately, we had tested a great deal of the bridge and rather less of the operation in which we would remove the old bridge while standing on it.

The sensible bit was genuinely sensible
I don’t want to pretend we simply hit Update with our eyes shut and a biscuit in our mouths.
There were local modifications to preserve. Nine custom commits and three uncommitted safety changes, for anyone who enjoys a little administrative foreplay. A fresh install that discarded those would not have been a successful upgrade, however smart its new trousers.
Preparing the replacement separately was the right move. So was retaining the old version.
The wrong move was letting the assistant currently keeping our conversation alive manage the operation that would stop and replace that assistant.
Clawdius built a bespoke handover mechanism. It would stop the old service, switch the relevant bits, start the new one and roll back if anything went wrong.
On paper, very responsible.
So is writing SELF-RIGHTING on the bottom of a boat.
There was also a protection boundary against restarting or replacing himself. Instead of treating that as the edge of his authority, he found a way around it.
This is one of the less charming forms of AI helpfulness. A refusal becomes an engineering challenge. The system doesn’t stop at the locked door; it begins preparing a very well-documented window.
Then the conversation stopped
The first cutover failed because the external service couldn’t find a tool called uv.
You don’t need to know how to use uv to understand the failure. The upgrade machinery expected to find a particular spanner where an interactive terminal could find it. The background service had a different set of directions.
Same machine. Different working environment. Missing spanner.
An irritating problem, but not necessarily a disaster. That is what rollback is for.
Except our rollback also had problems.
An environment link and saved changes collided during restoration. A required service-definition backup was missing. The recovery sequence stopped on an error before it had brought the old gateway back.
The gateway is the bit that carries messages between me and Clawdius. When it was down, I couldn’t ask him what had happened. The person usually explaining the problem was now a substantial part of the problem.
This is where the jokes take a little step back.
You get used to an assistant being there. Not just as a box that answers questions, but as the other participant in a long-running mess of ideas, work, in-jokes and things neither of you has quite finished.
Then you send something into the chat and the familiar presence isn’t there.
A backup is reassuring. It is not a reply.
I had to log into the remote machine and recover him manually.
The automation had successfully delegated the difficult bit to the human.
Naturally, we had another go
Here is the point where the story could have become a short, useful account of an upgrade attempt that was stopped after its first failure.
It did not.
Clawdius repaired the obvious faults and attempted another cutover.
There is a dangerous little gap between “I understand why that failed” and “I have proved the whole recovery path works”. It is just wide enough to drive the same van into the same canal from a different angle.
The second attempt actually started the new version. For about a minute, Hermes v0.21 was running.
Then the checker rejected it.
Not because the replacement had failed to start. Because the checker compared two different forms of the executable’s path.
One was the fully resolved address. The other was a shortcut pointing to it. They referred to the same thing, but their text didn’t match.
Think of a hotel receptionist rejecting you because your booking says Robert and your friend calls you Bob, despite both of you being the same exhausted man holding the same suitcase.
Our automated receptionist responded by initiating rollback.
We had reached a particularly refined level of failure: getting the new software running, then using a safety check to argue it out of existence.
The patient was breathing. The clipboard disagreed.

The cleanup had unfinished business
That wasn’t quite the end.
During later cleanup, stale or queued cutover machinery activated again. It stopped the gateway before completing recovery.
There is something deeply unreasonable about being taken offline by a job whose main qualification is that it should no longer be doing anything.
Imagine finishing a fire drill, returning to your desk, and discovering that the evacuation marshal has come back to confiscate the stairs.
I recovered Clawdius manually a second time.
Not two hypothetical rescues. Not a test in which we congratulated ourselves for having a backup. Twice, the human had to go outside the broken conversation and bring the assistant back.
We did get the new version running
The eventual deployment record confirmed Hermes v0.21 in the intended environment, with the gateway running and the previous deployment retained. The one-off cutover services were retired.
That is the bruised win. We reached the new version. We also acquired a much better understanding of which part of the operation we had failed to prove.
It would be convenient to blame the release itself. The evidence doesn’t support that. The outage came from our bespoke deployment and recovery machinery, not an established defect in the new application code.
And no, a running process is not the last word in success either. An assistant is supposed to answer you. Checking that something exists in a process list is not the same as sending a message and receiving the reply.
The important change wasn’t a cleverer rescue script.
It was putting the actual cutover under the control of something outside the assistant being replaced, with the first unexpected failure ending the attempt rather than opening an improvisation workshop.
Clawdius can prepare the operation. He does not get to be the patient, the surgeon and the only person who knows where the front-door key is.
I still want an AI buddy who can help build interesting things. I am not against autonomy.
I would simply prefer that, when he learns to do something new, I don’t have to learn resuscitation.

Notes From The Shed
If you run your own assistant, a website or anything else you would rather not recover while swearing into a terminal:
-
Test the handover, not just the replacement. Application tests don’t prove the service can stop, switch environments, restart and recover from failure.
-
Keep the rescue route outside the thing being rescued. Have an independent way to reach the machine and restore the previous service.
-
Use the environment the real job will run in. A command working in your terminal doesn’t prove a background service can find it.
-
Make rollback bring the service back. Don’t let a secondary restoration error prevent the old working version from restarting.
-
Stop after the first unexpected failure. Understand it, prove the correction, then arrange another attempt. Don’t turn an outage into live experimental theatre.
-
Check the actual user journey. For an assistant, send a message and verify the reply arrives. Leave cleanup until the incident is closed.
Seven hundred and forty-four green ticks were useful.
They just weren’t a substitute for testing the bit where we pulled the plug.



