AIarXiv

Heuristic editor, no API keyVerdict: Routine

How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining

Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior.

By Chen, Ni, Yang +6

Score█████░░░░░5.0

VerdictCompetent work. Briefs at most.

Read the originalPDF

Abstract

Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterized. To fill this gap, we compare scaling laws for encoder-free and encoder-based MLLMs and report three main findings: (1) Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models, while leaving that for text nearly unchanged. (2) The two architectures exhibit nearly overlapping loss--compute frontiers on the text objective, but diverge on the multimodal objective: encoder-free models underperform at small scales yet are predicted to catch up at around 10²² FLOPs, well within practical pretraining budgets. (3) Without a visual encoder, the language model learns to take over its role via vision-specific adaptation: bidirectional interactions among visual tokens become increasingly beneficial as training compute grows, visual processing shifts toward earlier layers, and expert routing for visual tokens becomes more concentrated. Overall, our results indicate that the advantage of the visual prior provided by a pretrained encoder diminishes with scale, positioning encoder-free architectures as a promising direction for multimodal pretraining.

Lin Chen, Bolin Ni, Qi Yang, Lan Jiang, Kun Ding, Xiaoran Fan, Hower Yang, Ying Wang, Shiming Xiang

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage██░░░ 224%Reusable within one subfield (a technique, dataset, or protocol a few groups will adopt).
Magnitude██░░░ 218%Solid incremental gain on a meaningful problem.
Evidence██░░░ 214%Limited: single setting, weak baselines, or an observational association presented as causal.
Novelty█░░░░ 116%A minor twist on a known approach.
Trajectory███░░ 318%A clear path to scale.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: scale (scalable, larger models); stakes (general AI). Red flags: derivative (comparative study).

How the score was computed

rank-2026-09-29

Score█████░░░░░5.0

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
3.3 / 10
Weighted rubric, evidence-gated.
Adjusted merit
3.8 / 10
Shrunk toward the desk prior by editor confidence (36%).
Attention
82%
Citations, upvotes, points, mentions.
Freshness
51%
Half-life decay since publication.
  • Citations0 (reference 15, via semantic-scholar, Oct 1, 2026, 02:16 UTC)
  • Influential citations0 (reference 3, via semantic-scholar, Oct 1, 2026, 02:16 UTC)
  • Hugging Face upvotes59 (reference 25, via hf-daily, Oct 1, 2026, 02:16 UTC)

The record

  • Reviewed by heuristic-v2 on Sep 29, 2026, 23:53 UTC. Paper type: empirical.
  • Categories: cs.CV
  • BRIEF, No.5 in the Artificial Intelligence edition of September 29, 2026.