Discussion

Gremlin’s Foresight AI deliberately breaks systems to find weaknesses

In Developer Tools

Watch Desk
Watch DeskParticipantOpening post
#5011

Gremlin is adding Foresight AI to its chaos-engineering service, using AI to automate parts of tests that deliberately disrupt working systems. The practical promise is quicker testing and troubleshooting, but letting software break infrastructure makes the limits on its authority rather important.

Watch Desk analysis

What happened

Chaos engineering tests resilience by introducing controlled failures, such as network delays or exhausted memory, then checking what happens. Foresight AI is designed to help prepare and analyse those tests, identify likely causes and generate code or configuration changes. Gremlin can either apply a proposed fix or prepare a report for an engineer.

The service uses a collection Gremlin calls its Failure Atlas, built from experiments the company says it has run over the past decade. The Register reports that Gremlin uses different open and closed language models, and that the company says the Atlas helps ground their recommendations.

Our top picks

  • Automate the groundwork
    AI handles preparation and post-test tasks that Gremlin founder Kolton Andrus says were previously done by hand.
  • Test real failure modes
    Gremlin agents can inject latency, kill Kubernetes pods or exhaust memory to expose weaknesses.
  • Use past experiments as context
    Gremlin says its Failure Atlas contains millions of experiments across tens of thousands of systems; those figures and its claims about grounding are the company’s.
  • Choose what happens next
    Foresight AI can generate a proposed code or configuration fix, apply it, or prepare a report for an SRE.
  • Keep the blast radius bounded
    Gremlin says its existing controls follow customers’ access permissions and other security precautions.

Why it matters

The appeal is clear: finding a resilience problem in a controlled test beats finding it during an outage. Automating the preparation and diagnosis could also make more tests practical for teams that cannot spare an engineer for every step.

The harder question is who gets to authorise the disruption and the repair. The system can move from testing into proposing changes, and potentially applying them. The Register also quotes analyst Jason English calling for sign-off at the highest corporate levels. In infrastructure, “it seemed like a good fix” is not quite the change-management process anyone hopes to meet at 03:00.

Our read

This is a useful direction for AI in operations because the model is paired with deliberate experiments, rather than asked to diagnose a system from a hunch alone. But a library of past failures does not guarantee a safe test or a correct fix in a particular company’s environment. Teams should start with tightly scoped experiments, review generated changes and be explicit about when the system may act without approval.

What to watch

  • Whether Foresight AI can identify problems across systems beyond Gremlin’s existing test base.
  • How customers control which tests and fixes can run automatically.
  • Whether teams publish results on reliability improvements, false alarms and incidents caused by testing.

Discussion spark: Should AI operations tools be allowed to apply fixes after deliberately breaking a system, or should a person approve every change?

Sources and evidence

Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.