AIHugging Face

Heuristic editor, no API keyVerdict: Routine

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards.

By Dong, Zhao, Yue +7arXiv

Score█████░░░░░5.0

VerdictCompetent work. Briefs at most.

Read the originalPDF

Abstract

Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and adherence to task requirements. This leaves the policy without a learning signal that favors clean, targeted implementations over those containing unnecessary or out-of-scope changes. We introduce GAGAR, a framework for quality-aware credit redistribution in code agent RL. Built on dynamic sampling that retains groups containing both passing and failing trajectories, GAGAR places all trajectories from each group in a shared workspace, where an SFT-trained agentic grader jointly inspects them and ranks the test-passing candidates. Based on this ranking, we downweight lower-ranked trajectories and proportionally rescale the advantages of all test-passing trajectories to restore their original sum. This sum-preserving redistribution retains the relative weights established by quality-based downweighting while shifting credit toward higher-quality implementations. We evaluate GAGAR at industrial scale using pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters). Controlled code-only Flash experiments show improved code agent performance, reduced trajectory-length growth, and more stable training. We further apply GAGAR in large-scale mixed-task RL with both Flash and Pro. Our results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude██░░░ 218%Solid incremental gain on a meaningful problem.
Evidence██░░░ 214%Limited: single setting, weak baselines, or an observational association presented as causal.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); scale (scalable); stakes (general AI).

How the score was computed

rank-2026-09-29

Score█████░░░░░5.0

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
3.7 / 10
Weighted rubric, evidence-gated.
Adjusted merit
3.9 / 10
Shrunk toward the desk prior by editor confidence (36%).
Attention
89%
Citations, upvotes, points, mentions.
Freshness
29%
Half-life decay since publication.
  • Citations0 (reference 15, via semantic-scholar, Sep 29, 2026, 23:53 UTC)
  • Influential citations0 (reference 3, via semantic-scholar, Sep 29, 2026, 23:53 UTC)
  • Hugging Face upvotes113 (reference 25, via hf-daily, Oct 1, 2026, 02:16 UTC)

The record

  • Reviewed by heuristic-v2 on Sep 29, 2026, 23:53 UTC. Paper type: method.
  • BRIEF, No.7 in the Artificial Intelligence edition of September 29, 2026.