AIarXiv
Heuristic editor, no API keyVerdict: NotableRetrieval-Augmented Skill Optimization via Cross-Harness Adaptation
An agent skill is a reusable, actionable natural-language artifact that guides an agent to perform a task effectively under a given harness.
VerdictWorth a reader's time today.
Abstract
An agent skill is a reusable, actionable natural-language artifact that guides an agent to perform a task effectively under a given harness. Recent studies have explored the optimization of agent skills, contributing to a growing collection of publicly available skills spanning diverse tasks, domains, and harnesses. Despite millions of publicly shared skills, existing skill optimization methods largely overlook this accumulated knowledge, instead relying solely on expensive agent rollouts to iteratively refine skills for a target task. To address this, we propose Retrieval-Augmented Skill Optimization (RASO), a framework that leverages an external skill corpus as prior knowledge throughout skill optimization. RASO retrieves relevant knowledge from existing skills and adapts it to the target task and harness via Cross-Harness Adaptation, accounting for mismatches in both domain and harness. RASO comprises two complementary stages: Retrieval-Augmented Skill Initialization (RASI) constructs a knowledge-grounded initial skill without requiring agent rollouts, while Retrieval-Augmented Skill Update (RASU) iteratively refines the skill by retrieving external knowledge guided by execution feedback. Across four agent benchmarks and two models, extensive experiments show that RASO consistently outperforms baselines without retrieval-augmented skill initialization and updating.
The editor's rubric
| Dimension | Level | Weight | What that level means |
|---|---|---|---|
| Leverage | ███░░ 3 | 24% | A method or resource many groups across the field will adopt within a year. |
| Magnitude | ███░░ 3 | 18% | Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem. |
| Evidence | ███░░ 3 | 14% | Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data. |
| Novelty | ███░░ 3 | 16% | A genuinely new approach to an open problem. |
| Trajectory | ██░░░ 2 | 18% | Some room to improve with obvious engineering. |
| Stakes | ███░░ 3 | 10% | Meaningful benefit to many people within a few years. |
Editor’s rationale
Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); gains (outperforms); novelty (alternative to status quo); verification (multiple benchmarks); stakes (global scale, general AI).
How the score was computed
- Merit
- 5.6 / 10
- Adjusted merit
- 4.7 / 10
- Attention
- 76%
- Freshness
- 26%
- Citations0 (reference 15, via semantic-scholar, Oct 3, 2026, 13:49 UTC)
- Influential citations0 (reference 3, via semantic-scholar, Oct 3, 2026, 13:49 UTC)
- Hugging Face upvotes46 (reference 25, via hf-daily, Oct 4, 2026, 13:49 UTC)