AIHugging Face

Heuristic editor, no API keyVerdict: Routine

Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue

Real-time full-duplex speech models can listen while speaking, enabling natural interaction without rigid turn boundaries.

By Ma, Feng, Jin +7arXiv

Score█████░░░░░5.0

VerdictCompetent work. Briefs at most.

Read the originalPDF

Abstract

Real-time full-duplex speech models can listen while speaking, enabling natural interaction without rigid turn boundaries. Existing benchmarks evaluate turn-taking, interruption handling and multi-round dialogue, but largely centre on a designated user rather than an assistant participating in a shared conversation among several people. We introduce Duplex-MPE to evaluate when such an assistant should answer, remain silent or stop speaking. The benchmark contains 2,000 scenarios with three or four human speakers and one assistant, each paired across explicit and implicit addressing of the same request. Models receive continuous conversation audio without transcripts or supplied turn boundaries. Four scores measure fresh response initiation, answer accuracy, silence preservation and stopping when a human resolves a request. We evaluate five open-weight speech systems: MiniCPM-o 4.5, Moshi, FLM-Audio, Voila and Freeze-Omni. MiniCPM-o 4.5 leads on three scored capabilities, while frequent speech from other systems can coexist with inaccurate answers or failures to remain silent. A transcript-based Gemini 3.1 Pro reference responds 64.3 percentage points more often to explicit than implicit requests; paired tests detect no significant response-rate difference for the speech systems.

Chengqian Ma, Wenhao Feng, Weixuan Jin, Gaole Dai, Tianyu Xie, Yuexiao Ma, Zhaolu Kang, Xiangyu Zhao, Xiawu Zheng, Fei Chao

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude██░░░ 218%Solid incremental gain on a meaningful problem.
Evidence██░░░ 214%Limited: single setting, weak baselines, or an observational association presented as causal.
Novelty███░░ 316%A genuinely new approach to an open problem.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); novelty (alternative to status quo).

How the score was computed

rank-2026-09-29

Score█████░░░░░5.0

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
4.0 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.0 / 10
Shrunk toward the desk prior by editor confidence (34%).
Attention
86%
Citations, upvotes, points, mentions.
Freshness
25%
Half-life decay since publication.
  • Citations0 (reference 15, via semantic-scholar, Sep 29, 2026, 23:53 UTC)
  • Influential citations0 (reference 3, via semantic-scholar, Sep 29, 2026, 23:53 UTC)
  • Hugging Face upvotes85 (reference 25, via hf-daily, Oct 1, 2026, 02:16 UTC)

The record

  • Reviewed by heuristic-v2 on Sep 29, 2026, 23:53 UTC. Paper type: method.
  • BRIEF, No.8 in the Artificial Intelligence edition of September 29, 2026.