Building RL Training Tasks for Coding Agents: Planted Bugs, Rollouts, and Grader Bugs
Notes from a spring 2026 contract building RL environments and training tasks for coding agents: what made a planted bug hard, how rollouts set difficulty, and why I audited the grader before believing a low score.
This spring I worked on a contract building RL environments and training tasks for coding agents. The environment was a synthetic ML-platform codebase: registry entries, experiment logs, drift reports, policy files, and the Python that reads them. Each task was a container holding a copy of that repo with bugs planted in it, a prompt describing the symptoms, a reference solution, and a grader that checked whatever the agent left behind. We ran each task many times against a coding agent and looked at the scores.
My job was to plant bugs the agent would miss and then make sure the grader rewarded only real repairs. The second half took more of my time than I expected.
My first nine tasks were far too easy
My first batch had nine tasks, from easy to hard: a mismatched config key, a model registered at the wrong stage, a significance test that assumed an even traffic split when the config declared an uneven one. I ran each one ten times. Eight averaged a perfect score and the ninth averaged 0.96. In one local run the agent finished a task in four turns.
A later task shows what I had to unlearn. It asked the agent to repair a tamper-detection audit:
- Version 1 had five bugs, and I had left
# BUG 1:-style comments on them. It scored 1.0 on all ten rollouts. - Version 2 removed the comments. It still scored 1.0. The agent compared the policy file to the code and fixed every line that disagreed.
- Version 3 added two functions the agent had to implement from the policy, plus three quieter bugs. The average dropped only to 0.92.
A drift-check task built on four textbook statistical mistakes, such as a logarithm in the wrong base, still averaged about 0.95. Hardening could also backfire. On three moderate tasks, a second round of hardening pushed scores up, not down, so I reverted all three. One of them was a "make the code match the policy" task, and it reached 1.00 even after I added new bugs.
What made a bug hard
I eventually added a self-critique pass that rated every planted bug from 1 to 5 before a task shipped. My early designs averaged about 2.2. Nearly every bug was either a straight policy-to-code diff or something a traceback would point to.
What I moved toward was code that runs cleanly and returns plausible wrong answers. On one task the broken module had missing imports and hardcoded values, so agents skipped repairing it and rewrote it from the prompt. On one model, about 80% of runs scored above 0.9. I replaced it with a module that ran and contained eight quiet mistakes. Here is the flavor, with the names and the domain changed:
# Illustrative only. Assumes non-empty input and positive counts.
def segment_error_rate(events, alias_to_canonical, window_end):
# the map is alias -> canonical; this inverts it, so aliases fall through unresolved
lookup = {canon: alias for alias, canon in alias_to_canonical.items()}
rows = [e for e in events if e["day"] < window_end] # policy: end day is inclusive
by_seg = {}
for e in rows:
seg = lookup.get(e["segment"], e["segment"])
errs, total = by_seg.get(seg, (0, 0))
by_seg[seg] = (errs + e["errors"], total + e["count"])
rates = [errs / total for errs, total in by_seg.values()]
return sum(rates) / len(rates) # policy: pooled rate, not a mean of rates
Other patterns I leaned on:
- Stubs that return a plausible default. A validation that always returns PASS never throws.
- Explicit zero and false values in layered overrides, which truthiness checks silently drop.
- A hidden replay carrying more than half the weight. The grader renamed identifiers, changed values and reran the agent's code, so a fix tuned to the visible data failed.
I also kept the broken workflow runnable, so the agent had to find each failure by checking outputs rather than reading a stack trace.
Low scores were often grader bugs
One of my better-calibrated tasks averaged about 0.5. When I audited it, one check asserted an absolute threshold an order of magnitude below the value in the experiment log that the agent was supposed to sync to. Agents that synced it correctly failed the check, so some of the difficulty I was pleased with came from a grader bug. I replaced the constant with a relative tolerance. I also fixed the reference solution, which had hardcoded the value instead of reading it from the log.
Other grader bugs I found:
- A hosted run below 0.4. A
signal.SIGALRMtimeout crashed the edge-case section in the hosted runner, and that section carried more than half the weight. I switched to a thread-based timeout and re-ran the task. It scored 0.98, which meant it was easy and I had to harden it again. - A 0% pass rate caused by permissions. The hidden replay created a root-owned temp directory that the agent's user couldn't enter.
- A log-parsing regex that dropped half the lines, because some names contained spaces.
- A hidden test that wrote JSON's
falsewhere Python neededFalse.
From then on I treated a surprisingly low score as a question. I ran the reference solution through the grader locally, then read the failing transcripts. On two other tasks I found no grader defect: the reference scored 1.0 locally, and agents were genuinely missing the hidden checks on all ten runs.
The no-op run
The first thing to check is to build the task, change nothing and grade it. Two of my tasks scored about 0.75 and 0.6 that way, because the broken baseline already passed most phases. Both also shipped a leftover log from my own hardening workflow. I deleted the logs and put a gate in front of scoring:
import hashlib
# sha256 of each repairable file as shipped (broken), recorded at build time
BROKEN = {"src/auditor.py": "9f2c…", "config/aliases.json": "41be…"}
def digest(path):
return hashlib.sha256(path.read_bytes()).hexdigest()
def attempted_repair(root):
return any(digest(root / p) != h for p, h in BROKEN.items())
score = weighted(results) if attempted_repair(root) else 0.0
On one task, my first version took its baseline snapshot when the grader module was imported, which happens after the agent has finished. The gate compared the edited files with themselves, and the task scored 0.0. The baseline has to be recorded at build time, somewhere the agent can't reach.
A gate only proves that something changed. The real fix was removing checks the broken code already passed. Reviewers found two more of my tasks with no-op scores around a third. In one of them, several hidden sub-checks were crediting the broken baseline.
I aimed for a no-op of zero. A small remainder from the "didn't touch protected files" check was acceptable, because doing nothing breaks nothing. Beyond that, the grader was paying for something other than the repair.
The agent's reach matters too. A reviewer caught one of my graders reading ground truth from a data directory the agent could write to, so I moved those reads under the protected tests directory. An automated audit found that another grader checked only the shape of a weights output, so uniform 1/n weights passed it.
Calibrating with rollouts
Rollouts were the expensive part, so I ran a small baseline first and paid for ten runs only if that looked healthy. By the end, grading was binary: the target was roughly one pass in five on one model, then one to four in ten on a second.
Binary grading changes the arithmetic. A run passes only if every check passes, so if you treat the checks as independent, the pass rate is roughly the product of the per-check rates. With 11 checks and a 0.2 target, you would need every check at about 0.86, which is very hard to engineer. The alternative is one bottleneck check at about 0.25 with the other ten near 0.99, because 0.99^10 × 0.25 ≈ 0.23. I aimed for that shape and tried not to create a second bottleneck.
One task had no passes in five on each of its first two versions. The hard check was a hidden one with no visible counterpart, so agents never saw that behaviour tested. After I added a hint in version 3, two of five passed. Five runs is a small sample, but it fit the rough rule I had written down by then: each step up in hint strength lifted the pass rate by something like 30 to 50%.
Some zeros had nothing to do with the task. In one stretch the agent sent empty tool calls for hundreds of turns. In another, runs failed because the account had run out of credits.
Hard versus obscure
The most useful lesson came from tasks that got harder for the wrong reason. One task scored 0.6 on one version and zero several versions later, when only the prompt had changed. I had removed the explicit output structure, so agents had to guess the JSON shape the grader wanted.
On another task, the grader required version fields on every record, while the prompt described them as top-level keys. One run in five passed. An automated audit called that contract "inferable-but-undisclosed."
After that, I treated a mistake repeated across runs as a reason to audit the prompt. In one task, agents consistently computed a summary statistic from derived values when the grader expected the raw ones. In another, about 30% of runs read an aggregation rule differently. I fixed both by stating the rule.
By the end, a task had to meet four conditions:
- The reference solution scored 1.0 on three fresh runs.
- The no-op scored zero, or close to it for a reason I could name.
- The pass rate landed in band.
- Someone had read the failing transcripts and could say the agent failed on the reasoning the task was meant to test.
Anything a check enforced had to be stated in the prompt or derivable from the repo. The difficulty was supposed to come from tracing it.