AI Agents · Reinforcement Learning · Code Generation

RHO vs Harness-1: Two Ways to Make the Agent Harness Do the Work

RHO tunes an existing harness from unlabeled trajectories; Harness-1 designs the harness so RL only learns semantic decisions. Same premise, different problems, each with its own numbers.

RHO vs Harness-1: Two Ways to Make the Agent Harness Do the Work

The shared premise, and the fork

Both papers start from the same observation: an LLM agent’s capability is set as much by its harness, the prompts, skills, tools, memory and control flow wrapped around the model, as by the model itself. A survey of the territory, Code as Agent Harness, catalogs how far this has gone: agents whose instructions, tool scripts and even control flow are themselves code artifacts that can be inspected, versioned and edited. Once the harness is a visible artifact, two research programs become possible, and RHO and Harness-1 are the cleanest published examples of each.

RHO treats the harness as an optimization target. You already have an agent running in production; you want it to get better over time without collecting labels. Harness-1 treats the harness as a trainability decision. You are designing an agent to be trained with RL; you want the gradient descent to spend its capacity on judgment, not bookkeeping. One tunes the harness after the fact, the other designs it up front. Choosing between them is choosing which problem you actually have.

Key numbers

DimensionRHOHarness-1Same harness?
What changesSkills, tools and instructions around a fixed Codex agentWhat the 20B policy is asked to do; bookkeeping moves to the environmentNo: one optimizes, one re-architects
Training signalSelf-validation + self-consistency over re-rolled past tasks, ranked by pairwise self-preferenceRL with curated-recall reward over eight retrieval benchmarksNo
Labels neededNoneGold supporting evidence for the recall metric during trainingNo
Headline resultSWE-Bench Pro 0.59 → 0.78 in one round (+19 points)0.730 average curated recall, +11.4 points over the next strongest open search subagentBenchmarks are not comparable
Other benchmarksTerminal-Bench 2: 0.71 → 0.76 (+5); GAIA-2: 0.29 → 0.37 (+8)Largest gains on held-out transfer benchmarksn/a
Versus strongest alternativeMeta-Harness 0.62 at matched single-round budget (label-using); reaches 0.80 only at 10 rounds and ~3x computeBeats the next open subagent by +11.4 points; “competitive with frontier searchers” is a soft, unpinned claimYes within each paper
Optimization cost103 agent calls per round vs Meta-Harness’s 41 single-roundHand-engineered harness: candidate pools, importance tags, dedup, budget-aware renderingNo
Load-bearing ablationDrop self-consistency: 0.56, below the 0.59 vanilla baseline; drop self-validation: 0.70; no diagnosis: 0.60Paper lacks a harness-vs-policy ablation; gains may partly be the well-tuned scaffoldYes within RHO

Read the table for what the rows share and what they do not. Every RHO number comes from SWE-style agent benchmarks with a fixed Codex backbone; every Harness-1 number comes from retrieval benchmarks with a trained 20B model. The 0.78 and the 0.730 measure different things on different tasks, and this page never puts them on the same axis.

RHO: tune the harness you already have

RHO’s scenario is the common one: an agent is deployed, it accumulates trajectories, and nobody has a labeled validation set that matches the next distribution of tasks. Its bet is that the agent’s own judgment over its rollouts is a usable proxy for quality. The pipeline selects a small difficulty-diverse coreset of past tasks with a determinantal point process, re-solves each task several times in parallel, and extracts two label-free diagnostic signals. Self-validation looks inside one trajectory for the agent’s own checks that failed; self-consistency looks across the parallel trajectories for disagreement. Those diagnoses drive a best-of-N proposal of candidate harnesses, and the candidate the agent prefers in pairwise comparison ships.

The result that matters: one round lifts SWE-Bench Pro from 0.59 to 0.78, and the harness it learns is concrete, debuggable tooling, like a check_build_and_lint script that finds the Go toolchain sitting outside the default path. This is not abstract prompt polish.

Two caveats from the paper’s own ablations deserve weight. First, self-consistency is doing most of the work: drop it and SWE-Bench Pro falls to 0.56, below the untouched 0.59 baseline, so a degraded version of the method is actively harmful. Second, the 0.78 is the favorable case. Terminal-Bench 2 gains only +5 and GAIA-2 only +8 on the same method, and self-preference is the agent grading itself, which can reward edits the model likes over edits that generalize; the chosen candidate does not always match the true top scorer. RHO also spends inference to stay label-free: 103 agent calls per round versus 41 for single-round Meta-Harness.

Harness-1: design the harness so RL only learns judgment

Harness-1 starts from a different frustration: RL is a wasteful tool for bookkeeping. A typical search agent trains one policy over a growing transcript to do two jobs at once, making hard semantic decisions (which query, which document) and remembering boring recoverable state (which constraints are open, which of fifty passages is worth keeping). The paper’s harness moves all of the second job into the environment: a candidate pool, an importance-tagged curated set, evidence links, verification records, compressed observations and budget-aware context rendering. The policy keeps only what to search, what to keep, what to verify and when to stop.

Because the reward can now target the thing you care about, the paper trains on curated recall, the fraction of gold supporting evidence that lands in the agent’s curated set, rather than only final-answer accuracy. That choice fits a retrieval subagent whose output feeds a downstream reader. The reported results: 0.730 average curated recall across eight retrieval benchmarks in web, finance, patents and multi-hop QA, +11.4 points over the next strongest open search subagent, with the largest gains on held-out transfer benchmarks, which is the evidence that the learned behavior is a generalizable search skill rather than benchmark memorization.

The caveats are symmetric to RHO’s. Curated recall is a retrieval-side metric: a high-recall subagent can still feed a weak reader, so end-to-end gains are not guaranteed. “Competitive with much larger frontier searchers” is soft, the abstract pins no frontier model and no margin. And the harness itself is hand-engineered, so without a harness-vs-policy ablation the gains may partly be a well-tuned scaffold rather than the RL alone. Building and maintaining that stateful harness is real engineering overhead a plain transcript policy avoids.

When to use which

  • You have a deployed agent, past trajectories, and no labels: RHO. Its entire point is label-free improvement from data you already have, and the +19 on SWE-Bench Pro shows one round is enough to matter.
  • You are training a search or retrieval agent and can design scaffolding up front: externalize bookkeeping Harness-1-style first. Every parameter the policy spends on dedup is a parameter not spent on judgment, and the transfer results are the payoff.
  • Your agent keeps failing on recoverable state, wrong-context errors, lost constraints: that is Harness-1’s diagnosis, and RHO will not fix it, because retrospective tuning edits the harness text, not where state lives.
  • Your agent fails in ways that are hard to diagnose without labels: RHO’s self-consistency signal is exactly for that, but watch the ablation warning, run best-of-N, and never ship the single proposal.
  • You can do both: they compose. A Harness-1-style state-externalizing harness is still a harness made of inspectable artifacts, which means a later RHO-style loop can tune it from production trajectories.

Limits and open questions

There is no controlled experiment putting both approaches on one task, because they target different tasks: RHO is measured on SWE and general-agent benchmarks with a frozen frontier backbone, Harness-1 on retrieval with a trained 20B model. Whether state-externalizing scaffolding would survive RHO-style retrospective editing, and whether RHO’s self-preference signal would give useful diagnoses over a trained 20B retrieval policy, are open questions both papers leave untouched. Treat the two as answers to two different problems that happen to share one premise: the harness is the leverage point.

FAQ

What is the difference between RHO and Harness-1?

RHO optimizes an existing harness from past unlabeled trajectories using self-consistency and pairwise self-preference, leaving the model fixed. Harness-1 designs the harness up front so that RL training only optimizes semantic decisions, with bookkeeping maintained by the environment. One tunes after deployment, the other shapes the architecture before training.

Does RHO need ground-truth labels?

No. It replaces the labeled validation set with two signals the agent computes on itself: self-validation (failed checks inside one trajectory) and self-consistency (disagreement across parallel rollouts of the same task). The SWE-Bench Pro gain from 0.59 to 0.78 happens with no external grading, at a cost of 103 agent calls per round.

What does “state-externalizing harness” mean in Harness-1?

It means routine, recoverable state management (candidate pools, dedup, verification records, budget-aware context rendering) is moved out of the policy and into environment-side code. The 20B policy only decides what to search, what to keep, what to verify and when to stop, which is what RL then trains.

Can RHO and Harness-1 be combined?

Conceptually yes, and nothing in either paper rules it out. A Harness-1-style harness consists of inspectable artifacts, so a later RHO-style loop could tune it from production trajectories. Neither paper tests the combination, so this is an inference, not a measured result.

Which harness approach works without a strong base model?

RHO explicitly assumes a competent base model: its diagnosis relies on the agent’s rollouts disagreeing for informative reasons, and the paper does not test the boundary where a weaker model disagrees noiseily. Harness-1 trains its own 20B policy inside the harness, so its results do not depend on a frontier backbone, but they do depend on the hand-engineered scaffolding and the curated-recall training signal.