Agent Lightning v1.0: Microsoft’s 3,500-Line Framework Trains the Agent Harness You Already Deploy
The Problem With Traditional Agentic RL
Early agent RL systems, including verl, AReaL, and slime, were built on a shared assumption. The training framework owns the interaction loop with the environment. A model emits an action, the environment returns an observation, the context updates, and the whole sequence forms one continuous token trajectory that the trainer can score and optimize.
That assumption held while agents were simple ReAct-style loops. It stopped holding when coding agents got good.
Real harnesses such as mini-SWE-agent, OpenHands, OpenCode, Claude Code, and Codex each bring their own context management, tool protocols, execution logic, and dependency stacks. None of them are shaped like a single token trajectory. To train one of these agents under the traditional model, a developer has to reimplement it inside the training framework, re-creating the context compaction rules, the tool-call parsing, the retry behavior, and everything else that makes the harness work.
The cost is not just engineering time. It is fidelity. A reimplementation drifts from the deployed system, which means the policy you optimize may not be the policy your users interact with. Agent Lightning’s core move is to stop reimplementing anything.
How Harnessed Agentic RL Works
The mechanism is an LLM proxy. Developers take the endpoint that previously called the model API and point it at Agent Lightning instead. The framework sits between the agent and the model, observing and recording model calls as they happen.
From there the inversion follows. The agent harness keeps ownership of the environment interaction loop. The training system never sees a ReAct loop or a tool schema. It sees a series of LLM request/response pairs, and that is all it needs.
The consequence is worth pausing on. Because the trainer observes calls rather than a unified trajectory, a single rollout can split into a variable number of training samples. One agent episode might produce three model calls or thirty, depending on how the harness decides to compact context or retry a failed tool call. The researchers describe this as introducing four key challenges, though the released material names them without detailing them. It is also the clearest sign that the abstraction has genuinely changed rather than been renamed.
Inside the Architecture: Three Components
Agent Lightning v1.0 is built from three pieces.
The API Gateway does the heavy lifting. It serves as an OpenAI-compatible LLM proxy, stores rollouts, models, and events, and links every model call from the harness back to its rollout. It also records the prompts, responses, and log probabilities that training requires. Because it speaks the OpenAI API format, most harnesses can connect without code changes.
The Rollout Controller starts and manages agent execution, running agents either as local processes or as standard Kubernetes jobs. This keeps agent execution separate from the trainer, which matters for both isolation and scaling.
The Customized Trainer is built on verl. It creates rollouts, waits for completion, collects samples, and assembles final training samples through a sample adapter. The adapter is where the variable-sample problem gets resolved.
Collocated Async RL and the 2x Speedup Claim
The performance story centers on a technique the team calls Collocated Async RL.
Synchronous RL has a well-known inefficiency. The trainer waits for the slowest agent in a batch before updating, and GPUs sit idle during that wait. Fully asynchronous RL fixes the utilization problem but typically demands separate GPU pools for rollout and for training updates, which raises the hardware bill.
Collocated Async RL lets rollout and model updates share the same set of GPUs. In the team’s experiments, this reportedly delivered about a 2x end-to-end speedup over synchronous RL while using fewer GPUs than conventional asynchronous setups. The comparison comes from the team’s own benchmarks rather than independent testing, and the published material does not specify the hardware configuration behind the figure. The claim is notable because it targets both axes at once: faster and cheaper, rather than trading one for the other.
The Transparent Update Cycle
Sharing GPUs between rollout and training raises an obvious question. What happens to an agent mid-episode when the model weights change underneath it?
The API Gateway handles this. During an update, it pauses new requests and waits for in-progress requests to finish. Once the update completes, rollout resumes. The entire state transition is invisible to the external agent harness, which simply experiences a brief delay in model responses rather than a broken connection or a corrupted episode.
This is the kind of detail that determines whether a framework is usable in practice. Weight updates that leak into running episodes would silently poison training data, and the failure would be hard to diagnose.
Integration: A Near-Zero Setup Cost
For an existing agent harness, connecting to RL training usually means pointing the model endpoint at the Agent Lightning proxy. That is the whole integration surface for many setups, which is a striking contrast to the reimplementation work traditional systems require.
The contrast with other Harnessed Agentic RL frameworks is also worth noting. Some host agents on commercial infrastructure, which introduces a dependency between the training loop and a third-party service. Agent Lightning’s design keeps execution on local processes or Kubernetes, which gives teams more control over where agent code and data actually run.
Adoption Caveats Worth Scrutiny
Two things deserve attention before adoption. First, reproducibility. When the trainer observes request/response pairs rather than a canonical trajectory, reproducing a run depends on the harness behaving identically across executions, including its context compaction decisions and retry logic. Any nondeterminism in the harness becomes nondeterminism in the training data. Second, deployment. The proxy sits on the critical path between agent and model, so its latency and reliability characteristics become part of the agent’s own. Teams should treat the gateway as production infrastructure, not a training-only convenience.
The open questions extend past these two. The four challenges that variable-length rollouts introduce are named but not detailed in the released material, and how well the sample adapter handles them across diverse harnesses will determine whether the approach generalizes or only works for the agents the team tested. Reproducibility under harness nondeterminism remains an unsolved engineering problem rather than a design flaw. And the 2x speedup figure comes from the team’s own experiments, so independent replication on different hardware and agent workloads is the number to watch.
Where Harnessed Agentic RL Fits
Agent Lightning v1.0 makes a structural argument rather than a benchmark argument. By handing the environment loop back to the agent harness and reducing the trainer to an observer of LLM request/response pairs, it removes the reimplementation step that has kept production agents out of RL pipelines. The 3,500-line codebase is small on purpose, and the OpenAI-compatible proxy means the integration cost is close to zero for many teams.
What the framework does establish is a cleaner default. If the agent you train is the agent you ship, then the harness stops being an obstacle to RL and becomes the thing being optimized. That is a meaningful shift in where the boundary between agent and trainer sits, and it is likely to influence how the next generation of agent training systems is built.