AIarXiv
Heuristic editor, no API keyVerdict: NotableCUAWright: A Minimal Unified Interface for Digital Agents
The prevailing approach to computer-use agents couples a model with a domain-specific harness: a browser or desktop environment equipped with human engineered tools that are fixed before task execution.
Key numbers
- 33.2% relative improvement in partial
- 37.5% compared with the published
- 4.7% and 44.0% in success
VerdictWorth a reader's time today.
Abstract
The prevailing approach to computer-use agents couples a model with a domain-specific harness: a browser or desktop environment equipped with human engineered tools that are fixed before task execution. As models' coding capabilities improve, the GUI native and static harness prevents them from direct programmatic operation on system state, as well as flexible construction of tools. To this end, we introduce CUAWright, a minimal terminal harness of roughly 3K lines of code that uses bash commands as its sole action interface, and a file system as its evolvable space for dynamically creating tools and managing the context. We conduct comprehensive experiments across a wide range of digital tasks, and demonstrate that by giving the agent a minimal, programmable interface, it achieves substantially stronger results compared to their GUI or hybrid CLI interface across a wide range of tasks. On OSWorld 2.0, CUAWright delivers a 33.2% relative improvement in partial reward while reducing estimated cost by 37.5% compared with the published GPT-5.5 baseline. On Online-Mind2Web and the long horizon Odysseys benchmark, CUAWright substantially outperforms GUI native harness by 4.7% and 44.0% in success rate, respectively. Furthermore, we found the gains extend to CAD applications that require accurate visual understanding and CLI interaction: on CADGenBench and BenchCAD, our unified harness yields 8.1%-41.6% relative improvements over other CLI based harnesses with GPT-5.5. Together, these results suggest digital environments are far more programmable than their GUI interfaces imply, and a minimal terminal-focused harness is the key for better performance and efficiency.
The editor's rubric
| Dimension | Level | Weight | What that level means |
|---|---|---|---|
| Leverage | ████░ 4 | 24% | A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing). |
| Magnitude | ████░ 4 | 18% | A qualitative jump: a capability or regime that did not exist before. |
| Evidence | ███░░ 3 | 14% | Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data. |
| Novelty | ██░░░ 2 | 16% | A new combination of known ideas. |
| Trajectory | ███░░ 3 | 18% | A clear path to scale. |
| Stakes | ██░░░ 2 | 10% | Benefits a professional community (practitioners, clinicians, engineers). |
Editor’s rationale
Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (wide range, programmable); gains (relative gain, absolute improvement, outperforms); scale (efficient); stakes (general AI).
How the score was computed
- Merit
- 6.3 / 10
- Adjusted merit
- 5.1 / 10
- Attention
- 61%
- Freshness
- 42%
- Hugging Face upvotes1 (reference 25, via hf-daily, Oct 6, 2026, 01:15 UTC)
- GitHub stars6k (reference 250, via hf-daily, Oct 6, 2026, 01:15 UTC)