AIarXiv

Heuristic editor, no API keyVerdict: Notable

Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills.

By Li, Jiang, Gao +6

Score██████░░░░5.9

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.

Xin Li, Hao Jiang, Xin Gao, Annan Wang, Yuchen Xie, Jinghao Guo, Xingwei Qu, Yichi Zhang, Chau Yuen

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty███░░ 316%A genuinely new approach to an open problem.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose, we report); breadth (many tasks); gains (outperforms); novelty (alternative to status quo); verification (multiple benchmarks, code released); stakes (general AI).

How the score was computed

rank-2026-09-29

Score██████░░░░5.9

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.9 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.8 / 10
Shrunk toward the desk prior by editor confidence (44%).
Attention
90%
Citations, upvotes, points, mentions.
Freshness
50%
Half-life decay since publication.
  • Citations0 (reference 15, via semantic-scholar, Oct 1, 2026, 02:16 UTC)
  • Influential citations0 (reference 3, via semantic-scholar, Oct 1, 2026, 02:16 UTC)
  • Hugging Face upvotes143 (reference 25, via hf-daily, Oct 1, 2026, 02:16 UTC)
  • GitHub stars11 (reference 250, via hf-daily, Oct 1, 2026, 02:16 UTC)

The record