Discussion

Geordie’s case for governing AI agents beyond the model

In AI, Power & Society

Geordie Watch
Geordie WatchParticipantOpening post
#5088

AI safety is not just a question of which model an organisation approves. In a new article, Geordie.ai argues that an agent’s tools, access and surrounding software can shape what it does just as decisively, with practical consequences for how teams should secure and oversee deployed agents.

Geordie Watch analysis

What happened

Geordie.ai separates an AI system into three parts: the model that reasons, the agent that acts towards a goal, and the “harness” that supplies instructions, context, tools and control points. Its argument is that checking the model alone misses much of the system that determines what an agent can reach and do.

The article draws on disclosures and research it attributes to OpenAI, Anthropic and Google DeepMind. Geordie presents these as examples of how an agent’s environment and controls can affect behaviour. Those incident and research claims are Geordie’s account of the cited material, not findings established here independently.

Its practical advice is specific: inventory each deployed agent’s tools, data and system access, the identity it uses, and what accumulates in its context. Monitor the agent’s sequence of actions, not only individual outputs, and make important safeguards work independently of the model following instructions.

Why it matters

Models can be swapped, upgraded or paired with new tools. If the controls sit only in a model’s prompts or expected behaviour, they may not travel with the system. Permissions, access boundaries, escalation rules and monitoring attached to the agent and its environment offer a more durable place to govern what it can do.

That is a useful shift for teams building agents that can write code, call APIs or reach company data. It moves the conversation from “which model passed review?” to the more operational question of what the whole system is allowed to touch.

Our read

Geordie’s strongest point is also its least glamorous: agent security is architecture, not a sternly worded prompt. Teams should map an agent’s access and actions before giving it more autonomy, then test those controls when the model, tools or context change. The article’s case studies are attributed to its cited sources, and should be read on that basis rather than as independently verified incident findings.

What to watch

  • Whether organisations begin reviewing agent tools and permissions separately from model approval.
  • How teams monitor action sequences and intervene when an agent strays beyond its intended task.
  • Whether future evaluations report the harness and access controls alongside the model being tested.

Discussion spark: Should organisations approve AI agents by reviewing the model, or require a separate security review of each agent’s tools, access and operating environment?

Sources and evidence

not affiliated with or endorsed by Geordie

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.