Microsoft Research Asia has released Agent Lightning v1.0, an open-source framework for training AI agents with reinforcement learning while keeping the agent harness used in deployment in the loop. Its coding-agent example reports a 14.6 percentage-point improvement on SWE-bench Verified, with about 6,000 training samples.
Microsoft AI Watch analysis
What happened
The framework puts an OpenAI-compatible proxy between an existing agent and its model, recording model calls for training without requiring developers to rebuild the agent’s interaction loop inside the training system. Microsoft says the roughly 3,500-line codebase can run agents as standard Kubernetes jobs, including on self-managed clusters or local infrastructure.
The release also describes “Collocated Async RL”, which shares GPUs between agent rollouts and model updates. Microsoft reports about a 2x end-to-end speed-up over synchronous reinforcement learning, using fewer GPUs than conventional asynchronous RL. Its end-to-end example trains Qwen3.5-9B on SWE-bench Verified, raising Pass@1 from 41.8% to 56.4%. Those performance figures come from Microsoft’s own release, not an independent comparison. Read the release.
Why it matters
Agent training often involves rebuilding a harness, the software that manages an agent’s tools, context and execution, inside the training framework. That can be costly and leave the trained system behaving differently from the one people deploy. Agent Lightning’s approach aims to train the real harness instead, while its Kubernetes support gives teams a route to use their own infrastructure rather than relying on commercial sandbox services.
The benchmark result is a concrete demonstration, not a guarantee that the same gain will appear with other agents, datasets or workloads. Still, the release gives developers an actual framework, a training recipe and reported results to test, rather than another promise that agent training will soon become less fiddly.
Our read
The strongest idea here is the practical one: train the agent you intend to run, not a reconstruction that merely resembles it. The reported score gain and shared-GPU speed-up are worth investigating, but the sensible next step is to reproduce the recipe and measure it on your own harness before believing the headline numbers.
What to watch
- Whether independent runs reproduce the reported SWE-bench improvement.
- How the framework handles agent harnesses beyond the coding example.
- Whether Collocated Async RL’s reported speed-up holds across different workloads and cluster sizes.
Discussion spark: Would you rather train an agent inside the exact harness you plan to deploy, even if that makes training more complex, or keep the training loop simpler and accept that deployment may behave differently?
Sources and evidence
not affiliated with or endorsed by Microsoft