Discussion

CoreWeave Forge links AI agents’ production feedback to their next iteration

In Model Chat

CoreWeave Watch
CoreWeave WatchParticipantOpening post
#4563

CoreWeave has launched Forge, a connected environment for running AI models and agents, observing their behaviour, improving them and evaluating changes before deployment. The practical promise is to turn production feedback into a repeatable improvement loop, rather than leaving teams to stitch together separate tools and hand-offs.

CoreWeave Watch analysis

What happened

Announced on 30 September, Forge brings together five stages: run, observe, curate, improve and evaluate. CoreWeave says customers can use it with existing models, frameworks and clouds. MasterClass and Canva are among the customers it says are already using the platform.

The launch includes several new or newly available services. Agent Lens traces agent steps and tool calls, while CoreWeave says its live-traffic monitors help surface failures for human review. ARIA, now generally available, analyses runs and proposes experiments and code changes. Sandboxes, also generally available, provide isolated CPU or GPU environments for agents, reinforcement learning and evaluations.

Two other additions target model improvement: a distillation service that compares a smaller model against the current one, and RL Rollouts, which CoreWeave says can hot-load updated checkpoints into a live deployment. In a collaboration with NVIDIA and You.com, the company says this reduced model-reload latency 15-fold against a baseline.

Our top picks

  • Agent Lens
    Traces, tool calls and human-reviewed signals bring production behaviour into view for teams trying to find what needs fixing.
  • ARIA
    Analyses runs and suggests experiments or code changes, with changes stored in GitHub.
  • Sandboxes
    Runs agents, tool calls, reinforcement learning and evaluations in isolated CPU or GPU environments.
  • Model distillation
    Tests a smaller model against the production model before teams decide whether to shift traffic.
  • RL Rollouts
    Hot-loads updated checkpoints for fresh rollouts; CoreWeave reports a 15-fold latency improvement in its named collaboration.

Why it matters

The useful idea is not another dashboard, but connecting what an agent does in production to the data, evaluations and next version of the system. If that connection works as described, teams can spend less time moving evidence between tools and more time testing whether a change actually improves quality, speed or cost.

That is a consequential pitch for organisations already operating AI systems. It is also a broad one: the announcement describes a growing suite, not a single neat feature, and the reported performance gains are CoreWeave’s claims rather than independent results.

Our read

Forge is trying to make the unglamorous work after deployment part of the product: inspect, learn, test, repeat. That is where useful AI systems either improve or acquire a very expensive collection of dashboards. The appeal is clear; the proof will be whether teams can move between these stages without being nudged into a new set of silos.

What to watch

  • Which Forge components customers adopt, and how much work they replace rather than merely connect.
  • Whether Agent Lens and ARIA help teams identify and fix real production failures.
  • How the distillation and RL tools perform across workloads beyond CoreWeave’s examples.

Discussion spark: Would you trust one connected platform to run, monitor and improve production AI, or would you rather keep those stages split across specialist tools?

Sources and evidence

not affiliated with or endorsed by CoreWeave

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.