8–12 min read · 2026-08-18
Lego-RL: Harness-Native Reinforcement Learning for Coding Agents
TL;DR
The ideal way to run reinforcement learning on a coding agent is clear: whatever agent you deploy in production, train with that same agent. However, this is non-trivial in practice. You will hit an engineering wall that has nothing to do with the algorithm:
- the agent framework quietly rewrites its own conversation history while it runs;
- the model finds shortcuts that reach the reward while bypassing the task itself;
- more than 90% of a rollout's time is spent inside the agent's interaction loop, while the trainer can do nothing but wait.
Lego-RL makes this possible: train the coding agent you actually deploy. Without modifying a single line of native control flow in OpenHands SDK, Claude Code, or OpenCode, we use all three harnesses directly for reinforcement learning. This delivers large gains for Qwen3.5-35B-A3B-Instruct across all three harnesses: from 64.0 to 70.4 in OpenHands SDK, from 62.4 to 68.2 in Claude Code, and from 57.2 to 66.6 in OpenCode.
- SWE-bench Verified (%)
- OpenHands SDK64.0 → 70.4 (+6.4)Claude Code62.4 → 68.2 (+5.8)OpenCode57.2 → 66.6 (+9.4)
- Train/Inference Correlation
- ≥ 0.998
- Median log-probability correlation
- Training Speed-up
- 2.5×
- Async + partial rollout, 2.5× faster than synchronous training
- Training task pool
- 2,699
- Tasks retained after difficulty screening
Why insist on training inside the native harness?
The mainstream recipe for agentic RL reshapes the agent into a form the training framework accepts: rewrite its initialization logic, swap in the framework's own tool set, then bolt on termination logic so the reward can be read back out. The recipe works, but it quietly changes the optimization target. The policy you train is optimal under that remodeled control flow, not necessarily under the harness users actually deploy.
How much does the harness matter? The starting model is the cleanest evidence we have: the same Qwen3.5-35B-A3B weights, dropped into three different harnesses, score 64.0, 62.4, and 57.2 on the same SWE-bench Verified suite. Swapping the harness alone opens a gap of nearly 7 points, larger than the headline gains of many post-training methods.
| Model | OpenHands SDK | Claude Code | OpenCode |
|---|---|---|---|
| Qwen3.5-35B-A3B (starting point of this work) | 64.0 | 62.4 | 57.2 |
| Qwen3.6-35B-A3B (next-generation base) | 67.4 | 63.4 | 60.6 |
| KAT-Coder-V2.5-Dev (post-trained from Qwen3.6) | 67.0 | 66.8 | 64.8 |
| Lego-RL-Qwen3.5-35B-A3B (this work) | 70.4 | 68.2 | 66.6 |
SWE-bench Verified (%), all measured under one unified configuration: temperature 0.7, 200 turns, 200k context.
Lego-RL takes the top score in every column. It even beats the newer Qwen3.6-35B-A3B by more than the entire 3.5→3.6 generational jump. For a fair comparison, all three runs use the same checkpoint, 2,699 tasks, a 200k context window, and 126 training steps.
Reward rises across all three runs without an entropy collapse, but the model does not behave the same way in every harness. Mean response length nearly doubles under OpenHands SDK (43.5k → 90.9k), while it grows from 41k to 51k under Claude Code. The harness shapes the model's behavioral personality, not just its final score.

Overall architecture: one infrastructure, shared by all harnesses
Lego-RL builds its trainer on verl and its sandboxed execution on Harbor. Adding a new harness only requires a lightweight adapter to launch the agent, point it at the inference service, and pass the interaction data back. Everything downstream is shared.


The optimization objective: a formal definition
A task instance consists of a problem statement, an initialized repository environment, and a task-specific executable verifier. The harness belongs to the environment, not to the policy. At turn , the harness maps the current interaction and repository state to a context , the policy generates an action , and the harness executes whatever tool actions were requested, producing . A rollout is the sequence of prompt/response pairs actually exchanged at the model API, and the verifier ultimately outputs a single bit:
Only policy-generated tokens participate in training. Writing for those positions:
Every turn conditions on the harness-supplied context , not on the raw history. We maximize the expected verifier reward with group-relative advantages. Each task samples trajectories and takes , with . All three production runs use GSPO's sequence-level surrogate:
where is the policy version that generated the group, and zeroes out trajectories terminated by infrastructure failures. The asymmetric bounds give the sequence-level ratio more room to move up than down. Replacing with the per-token ratio recovers PPO or GRPO, which the trainer also supports.
Two properties shape the engineering choices that follow. When rewards within a group are identical, : the group stays in the batch but contributes no gradient, so task difficulty relative to the current policy is on the critical path. And comes from executing code, so the signal is only as credible as the sandbox that produced it.
Pillar 1: Faithful optimization
The core difficulty: the archived transcript ≠ the token sequence at sampling time
The most intuitive way to log training data for an agent is to save the conversation transcript and re-tokenize it at training time. That is enough for SFT. On-policy RL needs the log-probabilities of the exact token sequence generated at sampling time, and the quantity recomputed from a transcript is not that quantity.
Real harnesses do rewrite their own history:
- Claude Code injects
<system-reminder>blocks mid-conversation; - OpenHands performs history compaction once the context window fills;
- tool-call arguments get re-serialized, sometimes with keys in a different order;
- sub-agents share a session with their parent agent.
Any one of these plants an error in the importance-sampling ratio that never raises an exception. It just drifts.
Lego-RL's answer: an in-process proxy at the serving boundary
The proxy speaks both the Anthropic and the OpenAI protocol, so the harness sees nothing but a changed base URL. At the moment of generation it records token IDs, response masks, and log-probabilities along with MoE expert routing. Context alignment then runs turn by turn:
- match tool calls by ID, not by serialized arguments;
- isolate sub-agent sessions to avoid cross-contamination;
- discard history the harness compacted away rather than reconstructing it, since reconstruction would itself introduce bias.
MoE routing replay: a detail that is easy to miss
MoE (mixture-of-experts) models add one further requirement: identical token IDs are not sufficient. If vLLM routes token to experts {3, 17} at sampling time but the training-time forward pass selects {3, 41}, the two are not computing the same probability.
Replaying rollout-time routing (R3) lifts the correlation from 0.9946 to 0.9993. Two silent bugs showed why this has to be checked, not assumed: a one-position routing misalignment scored worse than no replay (0.750 vs 0.995), and an undersized capture buffer for hybrid-attention models wrote out-of-bounds entries as 0, so coverage decayed to 24% before the guard started raising. After the fix, coverage stays above 99.8%.
Formally, write for the log-probability recorded at generation time and for the value the trainer recomputes. Faithful optimization requires:
That identity is on the same weights . Under async training the trainer's runs ahead of — bounded off-policyness, which corrects. A violation of the identity is a capture defect, and nothing corrects for that. Across the three production runs, median train/inference correlation never falls below 0.998.
Pillar 2: Reliable execution
Reward integrity
Reward is the task's own tests: 1.0 resolved, 0.0 otherwise. No reward model, so no reward-model drift — and every shortcut to a score without a fix has to be sealed off, because a strong coding model will find them.
| Cheat path | Incidence | Notes |
|---|---|---|
| Reading the fix from local git history | 4.6%–20.5% | log -p, show, checkout; no command blacklist covers them all |
| Modifying test files | 2.4%–19.4% | tests pass by construction |
| Downloading the reference patch from GitHub | ~1.9% | one request away if the network is open |
| Grader applying the reference patch itself | ~2.5% | score is 1.0 regardless of the agent |
Defenses sit in the task environment, on by default, and flip with Harbor's phases: denied while the agent runs, restored for grading.
| Resource | Agent phase | Grading phase |
|---|---|---|
| Network | Egress firewall in a privileged sidecar; public traffic dropped. The main container has no NET_ADMIN. | Relaxed so graders can still install PyPI deps. |
| Git history | Rebased to a single commit; fix objects no longer exist. | .git.orig restored for git apply. |
| Test files | Not provided; edits to test paths are rolled back. | Restored, then graded. |
Difficulty is relative to the current policy
Each task samples 8 rollouts. All-success or all-fail groups contribute no gradient, so a fixed pool gets less informative as the policy improves: on OpenHands SDK, zero-variance groups climb from 44.7% to 51.4%; OpenCode holds around 43.3%.

We begin with 36,884 OpenSWE candidates and screen them for validity, executability, and difficulty. A task must be solved 1–3 times out of 4 by Qwen3.6-27B on OpenHands SDK. That leaves 2,699 tasks, which also transfer to Claude Code and OpenCode. An ablation across four 951-task pools supports this cutoff: the selected band and its upper half improve, while the lower half and an unscreened sample do not. In fact, 72.7% of the unscreened pool was never solved.


When infrastructure fails, we mask the affected trajectory from the loss: 7.1% for Claude Code, 2.4% for OpenHands SDK, and 6.4% for OpenCode. We still keep these trajectories in the batch, but give them zero weight. Trajectories that simply hit the turn or token ceiling continue to count.

91% of a trial is the agent executing
An OpenHands SDK trial takes 920 seconds on average, and agent execution accounts for 91.3% of that time. Under Sync, this imbalance caused screening to stall 31 times at batch boundaries, with a median pause of 38.7 minutes. Async reduces per-step time by 2.5×. Even then, generation remains the bottleneck: the trainer still spends 40.8–66.1% of its time waiting.

Partial rollout avoids throwing away a ten-minute trial when the weights need to sync. It interrupts vLLM, loads the new weights, and resumes on the same replica using the prefix KV cache. From the harness's perspective, the entire trial is still a single HTTP request.
| Optimization | Stage | On | Off | Median speedup |
|---|---|---|---|---|
| Prebuilt task images | sandbox startup | 1.04s | 36.2s | 33.2× |
| Mounted agent runtime | agent startup | 0.51s | 7.82s | 15.4× |
| Lazy image pull | sandbox startup | 1.57s | 2.66s | 1.7× |
| Packaged grading toolchain | grading | 3.81s | 2.72s | 0.71× |
Lazy pull's gain is in the tail (23× worst-case, 21.6 GB → 1.59 GB network). Baking the grader into the image is overhead we accept for reward reproducibility.
Pillar 3: Observable training
Failure attribution: quickly telling policy problems from infrastructure problems
Several failures looked similar on the reward curve, but their causes were completely different. One run died at step 3 because the tool-call parser was set to hermes instead of qwen3_coder. In another, validation reward fell from 0.556 to 0.150 — it looked like policy collapse, but only 60 of 172 trajectories had reached grading; the rest never finished environment setup.
A third run killed all 1,024 trajectories on turn one, again because of an incompatible parser. One run did collapse for real. Trajectory-level analysis separated that policy failure from the infrastructure failures and produced an early-stop condition that would have fired 8 steps earlier.


What did training actually change? Evidence at the behavioral level
A rising mean reward cannot say what the model learned; behavior can. We drew 420 trajectories from each end of the OpenHands SDK production run and analyzed the agent's behavioral patterns.
The clearest change is in self-verification:
| Behavioral metric | Before training | After training | Change |
|---|---|---|---|
| Re-reading a file to confirm after modifying it | 73.6% | 98.1% | +24.5pp |
| Files examined before the first edit | 3.45 | 6.92 | 2× |
| Proactively running the test suite | 85.0% | 93.6% | +8.6pp |
The model picked up a working habit: look around before editing, and check the result after every change. Error recovery barely moved — among trajectories that hit a failed command, the share that still solves the task rises only from 63.9% to 66.8%. A terminal binary reward only sees the final outcome, so verification gets reinforced while mid-course correction stays flat.
The clearest gain is reliability rather than coverage. pass@8 improves by 4.7 points (83.2 → 87.9), while pass8 improves by 11.1 points (28.3 → 39.4). Malformed tool calls also fall from 1.07% to 0.15%. The model gets there by taking more turns, not by making each turn much longer: the average rises from 46.6 to 83.1 turns, while tokens per turn increase by 17%. This makes the turn limit just as important as the context window; otherwise, late-training trajectories will be cut off at the ceiling.

What's next
Several of the gaps in this release are already on the near-term roadmap.
- More tasks. The pool will grow beyond Python SWE-bench-style repair into a mixed set of other verifiable software tasks: multilingual work across more programming languages, NL2Repo (standing up a repository from a natural-language spec), and agent–user interaction, where the agent has to ask, clarify, and iterate with a human rather than close a ticket in isolation.
- Joint scaffold training. The three production runs each trained one policy on one harness; next we will train a single policy jointly across scaffolds, so the model has to stay useful under more than one control flow—the same harness dependence that opened a nearly 7-point gap on the starting checkpoint.
- PPO critic. A terminal binary reward cannot credit mid-course recovery, which is why self-verification moved and error recovery did not; we will add a PPO critic so value estimates can carry denser credit assignment than a single pass/fail bit at the end of a trajectory.
Framework updates, harness adapters, checkpoints, and task indices will keep landing in the open.
BibTeX
@misc{du2026legorlharnessnativereinforcementlearning,
title={LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents},
author={Yiming Du and Yuxin Jiang and Tao Yuan and Jianbo Dai and Shaowei Wang and Jierun Chen and Chaofan Tao and Xianzhi Yu and Lifeng Shang and Kam-Fai Wong and Xiaohui Li and Haoli Bai},
year={2026},
eprint={2608.17393},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2608.17393},
}