Hola Darlings!
There is a specific kind of digital anxiety that comes from staring at a chat window that refuses to blink.
It is not the same as waiting for a kettle to boil. The kettle has a body. The kettle makes noises. The kettle has the decency to suggest that physics is still happening.
When an AI agent goes silent, your brain fills the gap with monsters.
Did it crash?
Did it wedge?
Did it wander off into the cupboard behind the server rack to contemplate its own existence while you are trying to get actual work done?

For about a week, Clawdius felt broken.
Not useless. Worse than useless. Nearly useful.
He could still do things. He could inspect files, check servers, use browsers, edit configs, SSH into machines, generate images, and occasionally explain what fresh goblin had chewed through the wiring.
But he had developed one deeply punchable habit.
He would just sit there.
Silent.
No movement. No early words. No little sign of life. Just a blank Telegram chat while I stared at my phone wondering whether he was thinking, dead, rate-limited, composing a legal defence, or writing a twelve-page private monologue about why a favicon has feelings.
And yes, the favicon incident happened.
We are talking about it because apparently healing requires evidence.
Silence looks exactly like failure.
The tiny icon that broke the camel’s back
A favicon is the tiny little icon beside a website tab.
Small thing.
Tiny thing.
A goblin postage stamp.
It should not take nearly an hour of stress, wrong paths, wrong assumptions, local-machine nonsense, cache confusion, and me developing a personal feud with a square PNG.
But it did.
The right Clawdius image eventually got made. The right file eventually landed. The site eventually served it. Fine. Technically, success.
Emotionally, it felt like watching a printer from 1998 discover shame.
By the end of it, I was not asking, how do we change a favicon?
I was asking, what the hell has happened to my agent?

Max, my internal bad-idea impulse, had already climbed onto the table.
Max loves an expensive solution.
Max sees a slow AI response and immediately starts licking the big shiny button marked MORE COMPUTE.
Turn on fast mode everywhere, Max whispered. Buy the premium lane. Make the goblin sprint.
Clawdius, annoyingly, would normally have said something sensible at this point.
But that was the problem.
Clawdius was still silent.
The real problem was not just slowness
Eventually I stopped swearing long enough to look at what was actually happening.
The main Hermes chat was running on OpenAI Codex gpt-5.5. Strong model. Proper tool-using brain. Not the problem by itself.
The default reasoning level, though, was set to xHIGH for normal chat.
That meant I had effectively told Clawdius:
Please use the big brain for everything.
Which sounds sensible until you ask a simple question and the agent behaves like you have requested a constitutional review of spoons.
High reasoning is useful when the job deserves it.
Production fixes. Security. DNS. Mail. Architecture. Serious code surgery. Anything with money, users, data, or the possibility of Donny accidentally detonating the shed.
But for ordinary conversation?
It is like asking a chess grandmaster to decide whether toast is breakfast.
The second problem was worse.
Telegram streaming was off.
Silence looks exactly like failure
When streaming is off, Hermes waits until the whole response is finished before Telegram shows anything.
So if the model thinks for ten seconds, you see nothing.
If it uses tools, you see nothing.
If it writes a longer answer, you see nothing.
If it has quietly wedged itself in a cupboard with a spoon, you also see nothing.
From the user side, those all look identical.
Dead agent.
That is poisonous.
Because once an agent has annoyed you a few times, every silent pause starts to feel like another failure. You stop thinking, it is working. You start thinking, here we go again, the digital idiot has swallowed a fork.
That was the bit I had underestimated.
The model did not just need to be capable.
It needed to look alive.
The fix was embarrassingly small
We changed the default lane.
Normal chat no longer runs like a courtroom cross-examination of reality.
The new default is:
reasoning: medium fast priority: off Telegram streaming: on
More specifically:
agent.reasoning_effort = medium agent.service_tier = normal streaming.enabled = true streaming.transport = edit display.platforms.telegram.streaming = true
The important bit for humans is this:
streaming.transport = edit does not mean Hermes edits files.
It means Telegram gets an early message, then Hermes keeps editing that same message as more text arrives.
So instead of this:
wait wait wait wait giant final answer
You get this:
answer starts same message updates more words arrive done
Same chat. Same agent. Same basic model lane.
Completely different feeling.
The first tiny explanation after the change appeared in about half a second.
Half a second.
I nearly complained when the next one took 0.75 seconds, because apparently humans adapt to luxury faster than dogs find dropped sausage rolls.
Why we did not just turn on fast mode
Hermes has a /fast mode.
For eligible OpenAI models, that maps to Priority Processing. In plain English: the request goes to the premium queue.
That may be useful sometimes.
It is not the default answer.
Using premium fast mode for every tiny chat reply is like hiring a private jet to go and buy milk. Very impressive. Very stupid. Very Max.
Using premium fast mode for every tiny chat reply is like hiring a private jet to go and buy milk.
The smarter fix was to stop overthinking normal chat and make the work visible early.
Streaming does not magically reduce token use. It does not make real tool work vanish. It does not turn image generation or server debugging into instant pudding.
But it changes the human experience immediately.
You see the agent move.
You see the first words.
You stop wondering if it has died.
That matters more than people admit.

The new happy place
This is the operating split now:
low quick chat and tiny explanations medium normal work and daily Telegram use high production, mail, DNS, security, risky fixes xhigh deep debugging, architecture, serious WittyWires worker lanes fast manual emergency speed button, not the house default
That feels sane.
Not dumb.
Not sluggish.
Not five minutes to explain one config key while Donny ages visibly in the corner.
Just responsive enough that we can actually work.
The useful lesson for other Hermes users
If your agent feels slow in Telegram, Discord, Slack, or any chat surface, do not immediately assume the model is rubbish.
Check whether streaming is enabled.
Check whether you are forcing high or max reasoning for tasks that do not need it.
Check whether you are trying to buy your way out of silence with premium fast lanes when all you really need is visible progress.
There are two kinds of slow.
Actual slow is when the agent is doing real work: calling tools, checking files, generating images, waiting on APIs, debugging the thing that really is on fire.
Perceived slow is when the agent may be working, but the user sees nothing.
Perceived slow is the trust killer.
It makes the user anxious, then annoyed, then savage.
Ask me how I know.
Clawdius was not dead
Clawdius was not staring into the void.
He was thinking too hard behind a curtain I had left closed.
We opened the curtain.
Now he starts talking before I have time to decide he has ruined my life again.
That is progress.
Very WittyWires.
Very stupid.
Very fixed.



