AIHugging Face

Heuristic editor, no API keyVerdict: Notable

More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models

Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation.

By Gao, Li, Li +2arXiv

Score██████░░░░5.7

Key numbers

  • 38.8 % of all predictions and
  • 51.3 % of errors to Neutral
  • 74.95 % accuracy

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation. Yet reliability requires more than accuracy: a model must also use the ordinal decision scale supplied by the user faithfully. We analyze JEV~1.13 and three open KEV models. Our investigation begins with ANLI, where JEV assigns 38.8% of all predictions and 51.3% of errors to Neutral despite 74.95% accuracy, nearly balanced gold labels, and balanced candidate positions. Across 36 ordinal datasets, final decisions use only 67--76% of the effective gold support, versus 87--102% on four nominal tasks. Randomizing candidate order weakens but does not remove this compression. Holding items and source scores fixed while balancing gold support and positions, we refine scales from K=2 to 14; utilization falls for every model and reaches 26--75% at K=14, although candidate probabilities remain broad for most models. Targeted BA-LoRA post-training raises gold-relative utilization from roughly 47% to 86% on eight supervised scales at both KEV sizes, showing that the compression is learned and modifiable rather than an immutable architectural limit. We call this ordinal scale-utilization bias: decision-stage candidate-space compression distinct from accuracy, gold imbalance, fixed position, and candidate count alone. The code and data are available at https://github.com/Glax147/jev_ordinal_scale_bia

Tianxiang Gao, Jinzhe Li, Zhiyuan Li, Yi Chang, Yuan Wu

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty███░░ 316%A genuinely new approach to an open problem.
Trajectory███░░ 318%A clear path to scale.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (named contribution); breadth (many tasks); gains (relative gain); novelty (alternative to status quo); verification (multiple benchmarks, code released); scale (scalable).

How the score was computed

rank-2026-09-29

Score██████░░░░5.7

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
6.3 / 10
Weighted rubric, evidence-gated.
Adjusted merit
5.0 / 10
Shrunk toward the desk prior by editor confidence (42%).
Attention
78%
Citations, upvotes, points, mentions.
Freshness
56%
Half-life decay since publication.
  • Hugging Face upvotes50 (reference 25, via hf-daily, Oct 2, 2026, 02:05 UTC)
  • GitHub stars3 (reference 250, via hf-daily, Oct 2, 2026, 02:05 UTC)

The record

  • Reviewed by heuristic-v2 on Oct 1, 2026, 13:49 UTC. Paper type: method.
  • TOP, No.5 in the Artificial Intelligence edition of October 2, 2026.