AIHugging Face

Heuristic editor, no API keyVerdict: Routine

TokenRouter: Efficient Serving System for Token-Level LLM Routing

Large language model (LLM) routing distributes inference work across different models, advancing the cost-quality Pareto frontier of LLM serving.

By Fu, Liu, Wang +4arXiv

Score█████░░░░░5.4

Key numbers

  • 64.15x higher decoding throughput than

Caveats

  • Preprint; not yet peer reviewed.

VerdictCompetent work. Briefs at most.

Read the originalPDFCode

Abstract

Large language model (LLM) routing distributes inference work across different models, advancing the cost-quality Pareto frontier of LLM serving. While coarse-grained routing at the session or query level has been widely adopted in production systems, recent algorithmic work shows that fine-grained token-level routing can yield substantial efficiency and quality gains. However, efficiently serving token-level routed inference poses significant challenges to existing systems. Built on single-LLM assumptions, current systems suffer from severe step desynchronization and frequent batch admission delays under token-level routing, and they also impose high implementation complexity on developers. To address these challenges, we design TokenRouter, an efficient and developer-friendly serving system for token-level routed LLM inference. TokenRouter follows the principle of request-centric programming, model-centric execution: developers describe routing logic from the perspective of a single request, while the runtime launches a subserver for each LLM and dispatches requests asynchronously. Each subserver employs a delayed-batching scheduler, whose optimal hyperparameters are derived from a mathematical throughput model of the system. Across diverse routing algorithms, workloads, and model pairs, TokenRouter achieves 2.01-64.15x higher decoding throughput than existing systems, substantially advancing the serving efficiency of token-level LLM routing. Our code is available at https://github.com/thu-nics/TokenRouter.

Tianyu Fu, Tengxuan Liu, Ruoxi Wang, Yixin Dong, Yi Ge, Yichen You, Yu Wang

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence██░░░ 214%Limited: single setting, weak baselines, or an observational association presented as causal.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); gains (x-fold); verification (code released); scale (efficient); stakes (general AI).

How the score was computed

rank-2026-10-07

Score█████░░░░░5.4

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
4.0 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.0 / 10
Shrunk toward the desk prior by editor confidence (38%).
Attention
88%
Citations, upvotes, points, mentions.
Freshness
65%
Half-life decay since publication.
  • Hugging Face upvotes91 (reference 25, via hf-daily, Oct 9, 2026, 13:49 UTC)
  • GitHub stars4 (reference 250, via hf-daily, Oct 9, 2026, 13:49 UTC)

The record

  • Reviewed by heuristic-v5 on Oct 9, 2026, 03:46 UTC. Paper type: method.
  • Categories: cs.CL
  • TOP, No.5 in the Artificial Intelligence edition of October 9, 2026.