SkillOpt vs Ctx2Skill vs OpenSkill: 6 Ways Agents Learn Skills
Six 2026 methods teach LLM agents skills without retraining: SkillOpt edits one skill document for +23.5 points, Ctx2Skill mines skills by self-play, OpenSkill learns unsupervised from the open web, plus three more.
One question, six answers
Every serious agent stack now has a “skills” layer: reusable procedures and domain knowledge that the model did not have when it shipped. The open question of 2026 is where those skills come from. Six recent papers give six different answers, and they disagree on almost everything except one point: the model itself usually stays frozen.
The disagreement is about everything else. What gets updated (a single skill document, a whole skill library, or the model’s own LoRA weights), where the learning signal comes from (scored rollouts, self-play, failure attribution, the open web), and what the whole exercise costs. SkillOpt treats a skill document as the only trainable object and applies optimization discipline to it. Ctx2Skill manufactures its own feedback through self-play when no verifier exists. OpenSkill goes further and assumes no task supervision of any kind, learning from the open web. LatentSkill leaves the prompt entirely and compiles skills into weights. SkillAdaptor edits a skill library from localized failures, and SkillsVote asks the governance question: once many agents write skills into a shared store, how do you stop the store from poisoning itself?
Read the table below as six headline claims, not as a leaderboard. Every number comes from that paper’s own benchmark suite, so cross-paper rankings are meaningless; what transfers is the mechanism and the cost structure.
Key numbers
| Method | What gets updated | Headline result | Where it was measured | Same harness as the others? |
|---|---|---|---|---|
| SkillOpt | One natural-language skill document, bounded add/delete/replace edits | +23.5 average points (GPT-5.5, direct chat); best or tied on all 52 (model, benchmark, harness) cells | 6 benchmarks; e.g. SearchQA 77.7 to 87.3, SpreadsheetBench 41.8 to 80.7 | No, own suite |
| Ctx2Skill | A skill set mined from one long context | GPT-4.1 solving rate 11.1 to 16.5; GPT-5.1 21.2 to 25.8 | CL-bench: 500 contexts, 1,899 tasks | No, own benchmark |
| OpenSkill | Skills plus self-written virtual tests | 43.6% SkillsBench pass, +8.9 over the best baseline, 44.5% human reference | SkillsBench, 11 domains, Claude Opus 4.6 | No, own benchmark |
| LatentSkill | A LoRA adapter compiled from skill text | +21.4 points ALFWorld seen (74.3% vs 52.9 in-context) with 64.1% fewer prefill tokens | ALFWorld + Search-QA, Qwen3-8B | No, own benchmark |
| SkillAdaptor | Text edits to a retrievable skill library | +2.3 WebShop score, +1.5 PinchBench, +1.8 Claw-Eval over frozen backbones | WebShop, PinchBench, Claw-Eval | No, own benchmark |
| SkillsVote | Admission-gated lifecycle of a shared skill library | Up to +7.9 pp on Terminal-Bench 2.0 (offline), +2.6 pp on SWE-Bench Pro (online) | GPT-5.2 coding agents | No, own benchmark |
The spread from +2 to +23 is a property of the benchmarks, not a quality ranking. SkillOpt and LatentSkill report large jumps on task suites where skills have real leverage; SkillAdaptor deliberately measures itself as a low-cost reliability patch and reports the small numbers honestly. Ctx2Skill’s +5 points is arguably the most informative cell in the table, because it comes with absolute levels attached: even the best assisted model solves only 25.8% of CL-bench.
Where the learning signal comes from
The mechanism that decides everything else is the feedback loop. Each paper manufactures a different one.
Scored trajectories, treated like training data. SkillOpt’s bet is that prompt-tinkering becomes reliable once it obeys optimization hygiene. An optimizer model reads rollouts and proposes bounded edits to one skill document; a textual learning rate (an edit budget of roughly 4 units per step) caps how far any version moves, a held-out selection split acts as a validation gate, and a rejected-edit buffer stops the optimizer from re-proposing the same bad idea. The results reward the discipline: removing the slow, epoch-wise meta update costs 22.5 points on SpreadsheetBench (77.5 down to 55.0), and the final skills need only 1 to 4 accepted edits to reach 379 to 1,995 tokens. The catch is the premise: none of this works without an automatic way to score trajectories.
Self-play when no verifier exists. Ctx2Skill targets exactly the case SkillOpt cannot touch: a dense technical document with no ground truth at all. A Challenger agent writes probing tasks from the context, a Reasoner solves them under an evolving skill set, and a Judge returns binary pass or fail; Proposer and Generator agents convert failures into skill updates for each side. Pure adversarial self-play collapses, so a Cross-time Replay step picks the iteration whose skill set balances easy and hard cases, rather than keeping the last one. The loop runs 5 iterations of 5 tasks per context, and the resulting skills lift every backbone tested, though the absolute rates stay low.
The open web as supervision. OpenSkill assumes the hardest deployment setting: no answers, no curated skills, no grader. Its agent retrieves task knowledge and, separately, verification anchors from documentation, repos, and forums; then it writes virtual tests from those anchors and refines its skills against them with a gap-versus-bug classifier. The self-built verifier reaches 80.5% recall but only 56.9% precision against hidden ground truth, which means it over-rejects and burns iterations, but errs in the safer direction. On SkillsBench this zero-supervision loop gets within a point of the human reference (43.6% vs 44.5%), while on SocialMaze and ScienceWorld the margins shrink to 1 to 2 points, so the open-web advantage is task-dependent, not universal.
Failure attribution at the step level. SkillAdaptor keeps a frozen backbone and asks, after a failed trajectory, which skill deserves the blame. A Localizer finds the first actionable fault step (the earliest point where a different action would have changed the outcome), a Linker spreads credit across candidate skills, a Modifier rewrites or mints a skill, and a qualification gate re-executes the task under both skill sets, accepting the edit only at delta of zero or better. That gate is what prevents library rot, and it roughly doubles execution cost for accepted updates. The reported gains are single digits across WebShop, PinchBench, and Claw-Eval, consistently positive and consistently small.
Governance instead of learning. SkillsVote barely touches how skills are written; it controls what survives. The framework profiles a million-scale open skill corpus for environment requirements and verifiability, recommends skills before a run, and after a run decomposes trajectories into skill-linked subtasks with credit attribution, admitting only successful discoveries. The payoff shows up exactly where naive write-back hurts most: shared, long-lived libraries used by coding agents, with up to 7.9 percentage points on Terminal-Bench 2.0 and a more conservative 2.6 points online on SWE-Bench Pro.
Into the weights. LatentSkill is the outlier: it argues that a prompt-stuffed skill library is really parameters in disguise. A hypernetwork pretrained on roughly 171K skill documents compiles any skill text into a LoRA adapter in one forward pass, so adding a skill is inference, not training. The compiled skills cut ALFWorld prefill tokens by 64.1% while beating the in-context baseline by 21.4 points on the seen split, and weight-space composition works, but only when skills are decomposed into aligned components first; naively adding whole adapters fails.
When to use which
The honest answer is that these methods sit on different points of one trade-off triangle: supervision available, weight access, and skill reuse volume.
- You can score trajectories automatically and the same task family recurs: SkillOpt. A sub-2,000-token text file worth +23.5 points is the cheapest deployment story in this comparison.
- The input is one long, dense document with no labels and no verifier: Ctx2Skill, because it bootstraps the missing feedback signal instead of requiring it.
- The agent runs in the field against tasks nobody has graded: OpenSkill, accepting the retrieval cost and the noisy self-verifier. If you can supply supervision, do not pay this cost.
- The backbone is a closed API model, failures leave legible intermediate signals, and tasks are cheap: SkillAdaptor as a reliability patch. Skip it where fine-tuning is allowed and failures are dense-reward.
- Many agents share one skill store over months: SkillsVote, because without admission gates the shared memory degrades faster than individual learning can compensate.
- You own the weights, the skills are procedural, and prompt tokens are the bottleneck: LatentSkill. For knowledge-retrieval skills, plain RAG is still the better tool.
Limits and open questions
Three caveats apply to all six. First, the evaluation is self-referential: each paper benchmarks on its own suite, so the table above can compare mechanisms but never ranks methods. Second, absolute capability is still thin. A best-assisted 25.8% on CL-bench, a self-grader with 56.9% precision, and single-digit frozen-backbone deltas are all reminders that skill learning extends what a model can reliably do; it does not replace training. Third, cost accounting is uneven across the papers: OpenSkill pays open-world retrieval latency, SkillAdaptor roughly doubles execution for accepted edits, SkillsVote profiles a million-skill corpus, and Ctx2Skill spends about 25 agent tasks per document. None of these prices the amortization question (how many reuses justify the mining cost) in full.
The open research questions are shared too. Do mined skills transfer across contexts or must they be re-mined per document? Can a self-built verifier ever be trusted when its errors correlate with the skill’s own blind spots? And does the weight-space route scale beyond one backbone and two benchmarks, which is all LatentSkill demonstrates so far?
FAQ
Which agent skill-learning method has the largest reported gains?
SkillOpt reports the largest average gain: +23.5 points for GPT-5.5 in direct chat across six benchmarks, with +24.8 under the Codex harness and +19.1 under the Claude Code harness, best or tied on all 52 evaluated cells. LatentSkill is close on its own turf with +21.4 points on the ALFWorld seen split. These numbers are not comparable across papers, because each method was measured on a different benchmark suite with different backbones.
Can these agent skill methods work on closed models like GPT-5.5 without fine-tuning?
Yes, most of them are designed exactly for that. SkillOpt, Ctx2Skill, OpenSkill, SkillAdaptor, and SkillsVote all keep model weights frozen and put the learning into text artifacts outside the model. LatentSkill is the exception: it needs weight access to mount the LoRA adapters its hypernetwork generates, so it cannot run on a closed API model at all.
When should I choose SkillOpt over Ctx2Skill or OpenSkill?
Choose SkillOpt when you have an automatic scoring signal (exact match, executable checks, verifiers) and a held-out split, because its held-out gate and textual learning rate assume scored trajectories. Choose Ctx2Skill when the input is a long document with no labels, since it generates its own binary feedback through self-play. Choose OpenSkill when there is no task supervision at all and the agent must scavenge both skills and grading signals from the open web.
How does LatentSkill differ from the prompt-based skill methods?
LatentSkill compiles skill text into a LoRA adapter through a hypernetwork, so the skill lives in the weights and costs almost no context tokens at run time, saving 64.1% of prefill tokens on ALFWorld. The prompt-based methods keep skills as retrievable text, which is portable and inspectable but pays tokens on every step. LatentSkill’s trade-off is that it requires weight access and its evidence covers a single backbone so far.
Do I need a verifier to use any of these agent skill frameworks?
Only SkillOpt strictly requires a reliable scoring function, since its edit acceptance depends on a held-out validation gate. Ctx2Skill and OpenSkill were built for the no-verifier case: Ctx2Skill uses a self-play Judge with binary feedback, and OpenSkill constructs virtual tests from open-web verification anchors, reaching 80.5% recall against hidden ground truth. SkillAdaptor needs failures with observable intermediate signals rather than a formal verifier, and SkillsVote needs tasks it can synthesize for verification.