ClimateOpenAlex
Heuristic editor, no API keyVerdict: NotableEvaluation of Zero-Shot Annotation Models for Waste Classification Training Data in Textiles, Scrap, and Packaging Waste
Generating accurately annotated training data remains a major bottleneck for deploying computer-vision-based sorting systems in heterogeneous waste streams.
Key numbers
- 45.7% for packaging
- 43.0% for six-class textiles
- 21.0% for scrap
VerdictWorth a reader's time today.
Abstract
Generating accurately annotated training data remains a major bottleneck for deploying computer-vision-based sorting systems in heterogeneous waste streams. This work evaluates whether general-purpose zero-shot object detectors can reduce annotation effort for lightweight packaging waste, post-shredder scrap, and post-consumer textiles. YOLO-World, Grounding DINO, and OWL-ViT models were compared using mAP50, inference latency, and the Perfect Image Ratio (PIR), representing images requiring no annotation correction. Prompt performance and ablation were investigated, and the resulting annotations were further evaluated through downstream YOLOv8n training and direct annotation-time trials against manual and semi-automatic workflows. In the initial detector comparison, the best configurations achieved PIR values of 45.7% for packaging, 43.0% for six-class textiles, and 21.0% for scrap. Prompt reduction improved performance and reduced inference latency, while multi-class tasks required ensemble-aware prompt ablation. PIR-selected zero-shot annotations yielded downstream training performance close to equivalent manually annotated subsets. Zero-shot pre-annotation reduced human annotation and correction time by 47.0%, 56.9%, and 64.9% for packaging, scrap, and textiles, respectively. Zero-shot detection therefore cannot replace human verification, but can substantially reduce manual annotation effort and support efficient creation of waste-specific object-detection datasets.
The editor's rubric
| Dimension | Level | Weight | What that level means |
|---|---|---|---|
| Leverage | ███░░ 3 | 16% | A method or resource many groups across the field will adopt within a year. |
| Magnitude | ███░░ 3 | 20% | Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem. |
| Evidence | ███░░ 3 | 20% | Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data. |
| Novelty | ██░░░ 2 | 10% | A new combination of known ideas. |
| Trajectory | ███░░ 3 | 14% | A clear path to scale. |
| Stakes | ██░░░ 2 | 20% | Benefits a professional community (practitioners, clinicians, engineers). |
Editor’s rationale
Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: breadth (general-purpose, zero/few-shot); gains (relative gain); verification (ablations); scale (efficient).
How the score was computed
- Merit
- 5.4 / 10
- Adjusted merit
- 4.6 / 10
- Attention
- 0%
- Freshness
- 84%
- Citations0 (reference 15, via openalex, Sep 29, 2026, 23:38 UTC)
- Field-weighted citation impact0 (reference 3, via openalex, Sep 29, 2026, 23:38 UTC)