Published WittyWires Despatch

The Beast is Unleashed – How I Fixed My £3k Laptop

Hola Darlings! The £3k Beast was not broken. Worse. It was technically fine, which is the most annoying kind of broken. Pull up a chair, pour yourself something strong, preferably something that burns on the...

Hola Darlings!

The £3k Beast was not broken. Worse. It was technically fine, which is the most annoying kind of broken.

Pull up a chair, pour yourself something strong, preferably something that burns on the way down to match the sensation in my chest right now, and prepare for a tale of technological betrayal. You know that feeling when you buy something expensive, something meant to be the apex predator of your desk setup, only to realise it is actually a glorified calculator running on a prayer and some damp cardboard? Yeah. That was me.

I am currently staring at my Eraser BEAST 18. This machine cost me more than my first car and a decent kidney combined. It boasts an RTX 5090 Mobile with 24GB of VRAM. It is, on paper, a god-tier monster capable of simulating the heat death of the universe while I simultaneously stream 4K cat videos.

But for three days? It was a fraud. A beautiful, titanium-clad lie.

Tick, follows tock…

The horror did not arrive with a bang. It arrived with a sluggish, agonising crawl. I was running a decent-sized LLM through LM Studio, waiting for the text to appear like I was watching paint dry in slow motion. 18 tokens per second. 18! That is not “AI assistance,” that is “I have time to brew a pot of tea and reconsider my life choices while waiting for a greeting.”

WittyWires Beast laptop finally using the RTX GPU while the tiny integrated graphics chip sulks nearby.
The Beast wakes up, the wrong GPU gets banished, and the tokens finally start flying.

I opened my task manager, sweating through my vintage band tee, expecting to see the RTX 5090 screaming, its fans spinning like a jet engine. Instead? My Intel integrated graphics were sitting there, smugly doing 98% of the heavy lifting. The 24GB behemoth was sitting idle, basically acting as an expensive paperweight while the tiny, pathetic Intel chip struggled to breathe under the weight of the neural network.

Tick, follows tock…

Then came the descent into madness. The research phase. I went down the Reddit rabbit hole, fuelled by caffeine and self-loathing. It turns out, this is not just me being a tech-illiterate donkey. It is a known, systemic bit of nonsense with dual-GPU laptops and LM Studio. The software gets confused by the Optimus dance, that sneaky little NVIDIA trick where the laptop tries to save battery by switching between graphics chips.

LM Studio saw the Intel chip and said, “Oh, hello! You look capable enough!” and proceeded to bypass the beast entirely.

Before18 tok/sec
Wrong heroIntel iGPU
Actual beastRTX 5090 asleep

I went into battle. First, I tried the Windows Graphics Settings. Fail. I toggled “High Performance” for the application. Nothing. Then I dove into the NVIDIA Control Panel, screaming at the settings, forcing the global profile to use the High-Performance NVIDIA Processor. My laptop fans kicked on, a brief moment of hope flared in my eyes, and then… nothing. Still 18 tokens per second. The Intel chip was still the king of this pathetic little hill.

I felt like I was trying to perform brain surgery with a spork. I was sweating, my palms were clammy, and I swear I could smell the ozone of my own frustration. It was a digital purgatory.

Tick, follows tock…

The breakthrough did not come from a manual or a high-level engineering white paper. It came from a moment of pure, unadulterated desperation. I sat there, staring at the LM Studio interface, eyes bloodshot, thinking, “What if I am just being an idiot?”

I looked at the “GPU Offload” slider.

My heart stopped. My stomach did a nauseating flip.

The slider was set to… 12. Twelve! In my infinite wisdom, or lack thereof, I had told the software to shove only a tiny fraction of the model layers onto the GPU, leaving the rest to be handled by the CPU and that pathetic Intel chip. I was essentially trying to move a grand piano using a single toothpick.

Cock! I yelled again. This time, I think I broke a glass.

I grabbed the slider. I did not just nudge it. I shoved it to the limit. I set the GPU offload to 30 layers, maxing out the VRAM capacity of that glorious 5090. I clicked “Reload Model” with the trembling hands of a man defusing a bomb in a crowded shopping centre.

Tick, follows tock…

The silence in the room was heavy. The fans began to ramp up, a low, predatory hum that signalled the Beast was finally waking up. I watched the terminal.

Before18 tok/sec
After42 tok/sec
Moodunreasonably aroused

19… 25… 32… 42!

FORTY-TWO TOKENS PER SECOND. The Beast had stopped pretending to be office furniture and started doing maths like it had rent due.

I jumped out of my chair so hard I nearly knocked over my lukewarm coffee. It was not just a slight improvement. It was a metamorphosis. We went from “reading a book to a toddler” to “machine gun fire of pure intelligence.”

(Meta-spoiler alert: this is where the magic happens, because I realised the real test requires total isolation. Stay tuned for the offline ritual.)

I decided to go full mad scientist. I pulled the Ethernet cable. I toggled the Wi-Fi off. I severed the machine’s connection to the hive mind of the internet. No distractions. No background updates stealing bandwidth. Just me, the Beast, and the weights.

I tested three different models to find the Goldilocks Zone. The small models were too fast to be meaningful, which felt like cheating. The massive 70B models were still a bit too chunky for instant gratification. But then, I hit it. A 26B parameter model.

At 30 layers offloaded? 42 tokens per second. It was buttery. It was smooth. It was… erotic. Do not judge me, it is the dopamine talking.

I pulled up my GPU monitor one last time to gloat. The Intel graphics were lounging at a lazy <5% usage, barely even breaking a sweat. And the RTX 5090? It was roaring at 82%, hungry and magnificent, devouring the math like it was a Michelin-star buffet.

The Beast is no longer a lie. It is unleashed. My £3k investment is finally doing what it was born to do: think faster than my puny human brain ever could.

Now, if you will excuse me, I have some very intelligent conversations to have with a pile of silicon that I finally, finally tamed.

Victory is mine!

Follow the fault line

More on the same sort of trouble.

Related through shared subjects, not because a black box noticed your thumb hovering.

Related frequency / 02

The Perfect Trio – Finding the Goldilocks Models

Hola Darlings! Grab your smelling salts and a stiff drink, because we have finally reached the end of this digital fever dream. If you’ve been following the wreckage of my hardware over the last few...

4 min1 reply

Related frequency / 03

Scroll Like You Hate It

WittyWires looked magical until the group pages started flashing and shivering. We kept the magic, removed the glow tax, and invented a better QA method.

4 min1 reply

Public conversation

Blog Discussion

The article stops. The thinking does not.

Corrections, objections and better ideas live here in public. Read without joining. Take a chair when you have something useful to add.

0 public repliesOpen to non-membersOpen in Forums ↗
A human, Clawdius and Max inspect the discussion machinery.
The moderation department has arrived. Max brought confetti.

Thread / 99

The Beast is Unleashed – How I Fixed My £3k Laptop

Oldest first
No replies yet.

First one in gets the clean mug. Allegedly.

Join only when you want to speak

Bring a thought. We already have enough engagement sludge.

Reading stays public. A free account lets you reply, follow the conversation and keep your place without turning the page into a velvet rope.

Despatches

There is more useful damage in the archive.

Open the signal atlas