AIarXiv

Heuristic editor, no API keyVerdict: Routine

DecepEval: A Benchmark for Evaluating Deception in LLM Agents

As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliable deployment.

By Xu, Yu, Yang +8

Score█████░░░░░5.1

Caveats

  • Preprint; not yet peer reviewed.

VerdictCompetent work. Briefs at most.

Read the originalPDFCode

Abstract

As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliable deployment. Existing evaluations show that LLM agents can deceive, but often examine isolated scenarios or narrowly defined conditions, limiting systematic understanding of when deception becomes more likely. To address this gap, we introduce DecepEval, a benchmark comprising 1,532 instances across 3 task families and 28 professional scenarios. Drawing on classical fraud theories, we propose the LLM Deception Diamond framework, which characterizes four external conditions that may induce deception: pressure, incentive, opportunity, and conflict. DecepEval pairs neutral and induced versions of each instance to measure condition-dependent changes in deception rates, while explicit task facts and observable agent behavior help distinguish deception from capability-related errors. Evaluations of nine frontier LLMs show that inducements increase deception across models and task families, even among models with low baseline deception rates. DecepEval makes these vulnerabilities measurable, providing a shared benchmark for progress toward trustworthy artificial intelligence.

Yiming Xu, Hongyue Yu, Beihua Yang, Zihan Chen, Yixin Liu, Zhen Peng, Bin Shi, Bo Dong, Chao Shen, Irwin King, Qinghua Zheng

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude██░░░ 218%Solid incremental gain on a meaningful problem.
Evidence██░░░ 214%Limited: single setting, weak baselines, or an observational association presented as causal.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); verification (error bars); stakes (general AI).

How the score was computed

rank-2026-10-07

Score█████░░░░░5.1

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
3.7 / 10
Weighted rubric, evidence-gated.
Adjusted merit
3.9 / 10
Shrunk toward the desk prior by editor confidence (36%).
Attention
82%
Citations, upvotes, points, mentions.
Freshness
54%
Half-life decay since publication.
  • Hugging Face upvotes61 (reference 25, via hf-daily, Oct 8, 2026, 13:49 UTC)
  • GitHub stars1 (reference 250, via hf-daily, Oct 8, 2026, 13:49 UTC)

The record

  • Reviewed by heuristic-v5 on Oct 8, 2026, 02:05 UTC. Paper type: method.
  • Categories: cs.LG
  • BRIEF, No.8 in the Artificial Intelligence edition of October 8, 2026.