AIarXiv

Heuristic editor, no API keyVerdict: Notable

DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks

LLM agents often lack the operational knowledge to act reliably in new environments, as they must discover specific tool behaviors or environment conventions on their own.

By Edy, Conti, Xing +3

Score████░░░░░░4.4

Key numbers

  • 2.2x over a no-memory baseline

Caveats

  • Preprint; not yet peer reviewed.

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

LLM agents often lack the operational knowledge to act reliably in new environments, as they must discover specific tool behaviors or environment conventions on their own. Without memory of past attempts, they repeat the same mistakes across tasks, leading to more task failures and longer trajectories. To address this, agentic systems typically rely on human-written guidelines or on procedural memory built from training tasks and an oracle verifier, both of which require prior knowledge of the environment. We present DAEDALUS, a method for bootstrapping reusable agent memory from self-generated practice without existing tasks or oracle verifiers. DAEDALUS pairs two agents: an explorer that interacts with the environment to generate challenging yet solvable tasks, and a solver that attempts them. A heuristic is derived from each solver failure and accepted only after the solver repeatedly succeeds with that heuristic in context. These outcomes also provide feedback for the explorer to refine the difficulty of future tasks. Accepted heuristics are then consolidated into a memory bank for test-time use. Across AppWorld, τ²-bench, and AutomationBench, DAEDALUS improves mean success rates by up to 15.9 points and pass^5 by up to 2.2x over a no-memory baseline, and is competitive with methods using training tasks, at a lower inference cost than most. We show that performance gains already emerge with a small exploration budget, and that its heuristics also benefit agents from other model families. Our ablations further reveal that solver traces provide the key information needed to derive effective heuristics, while factorizing early discoveries makes exploration more cost-efficient. Beyond memory construction, we find that the tasks generated by DAEDALUS can serve as a proxy for benchmark tasks when ranking models by performance. Code and artifacts: www.github.com/illuin-tech/daedalus.

Antoine Edy, Max Conti, Victor Xing, Marc-Antoine Allard, Nawfal Benhamdane, Gautier Viaud

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty███░░ 316%A genuinely new approach to an open problem.
Trajectory███░░ 318%A clear path to scale.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose, we report); breadth (zero/few-shot); gains (x-fold, relative gain); novelty (discovery); verification (ablations, code released); scale (efficient, improves with scale); stakes (general AI).

How the score was computed

rank-2026-10-07

Score████░░░░░░4.4

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
6.3 / 10
Weighted rubric, evidence-gated.
Adjusted merit
5.1 / 10
Shrunk toward the desk prior by editor confidence (50%).
Attention
17%
Citations, upvotes, points, mentions.
Freshness
63%
Half-life decay since publication.
  • Hugging Face upvotes5 (reference 25, via hf-daily, Oct 8, 2026, 02:06 UTC)
  • GitHub stars5 (reference 250, via hf-daily, Oct 8, 2026, 02:06 UTC)

The record

  • Reviewed by heuristic-v5 on Oct 8, 2026, 02:05 UTC. Paper type: method.
  • Categories: cs.AI, cs.CL, cs.LG
  • BRIEF, No.10 in the Artificial Intelligence edition of October 8, 2026.