AIHugging Face

Heuristic editor, no API keyVerdict: Notable

GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis

Working agents need to read diverse files, coordinate tools, and produce deliverables.

By Su, Wang, Zhu +9arXiv

Score██████░░░░5.6

VerdictWorth a reader's time today.

Read the originalPDF

Abstract

Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Existing pipelines either generate files with models, which lack realism and diversity, or build tasks on real files without task-specific verifiers, leaving result quality unchecked. We introduce GraphForge, an evidence-graph based framework that grounds both the task and its verification in real files. Starting from occupation-grounded seeds for controlled diversity, GraphForge assembles a workspace of real files for each seed and builds an evidence graph over their relations. Since the task statement and rubrics are both derived from this graph, task requirements are backed by the workspace files and each criterion is anchored to the files needed to verify it. An initial rollout further tests executability, and a revision agent repairs the task and rubrics against the original files before trajectories are collected. Fine-tuning Qwen3.6-27B on 2,169 GraphForge trajectories brings GDPVal to 1445.7 (+65.7) under OpenHands, and Workspace-Bench-Lite and SpreadsheetBench II to 63.7 (+7.7) and 24.0 (+13.7) under Claude Code. Rejection fine-tuning on the SFT model's own rollouts, with candidates selected by the evidence-anchored rubrics, yields further improvements on all three benchmarks, suggesting that the rubrics provide a useful selection signal. The data and models are available.

Qisheng Su, Hanchen Wang, Guanru Zhu, Huicheng Jiang, Qiuyinzhe Zhang, Kou Shi, Zhen Fang, Ziao Zhang, Qingnan Ren, Zehui Chen, Tao Gui, Feng Zhao

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude██░░░ 218%Solid incremental gain on a meaningful problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty███░░ 316%A genuinely new approach to an open problem.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); novelty (alternative to status quo); verification (multiple benchmarks, code released); stakes (general AI).

How the score was computed

rank-2026-09-29

Score██████░░░░5.6

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.6 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.6 / 10
Shrunk toward the desk prior by editor confidence (40%).
Attention
90%
Citations, upvotes, points, mentions.
Freshness
32%
Half-life decay since publication.
  • Citations0 (reference 15, via semantic-scholar, Oct 2, 2026, 13:49 UTC)
  • Influential citations0 (reference 3, via semantic-scholar, Oct 2, 2026, 13:49 UTC)
  • Hugging Face upvotes132 (reference 25, via hf-daily, Oct 4, 2026, 02:05 UTC)

The record

  • Reviewed by heuristic-v2 on Oct 2, 2026, 02:05 UTC. Paper type: method.
  • TOP, No.7 in the Artificial Intelligence edition of October 4, 2026.