PhysicsarXiv
Heuristic editor, no API keyVerdict: NotableFLAT: Smoothing the Rugged Landscape for Learnable, Sample-Efficient Traffic Calibration
Calibrating microscopic traffic models for digital twins is an expensive black-box optimization problem: tuning car-following and lane-changing parameters requires a full simulation run, affording only a tight budget…
Key numbers
- 4.4 times more simulations to match
- 81% on average
VerdictWorth a reader's time today.
Abstract
Calibrating microscopic traffic models for digital twins is an expensive black-box optimization problem: tuning car-following and lane-changing parameters requires a full simulation run, affording only a tight budget per recalibration window. Matching raw trajectories yields a rugged objective that sparse surrogates cannot learn, reducing sequential acquisition to near-random probing. We present FLAT, which couples what to optimize with where to sample next. An eight-dimensional behavioral fingerprint smooths the parameter-error landscape, making the objective learnable from a few dozen samples; annealed lower-confidence-bound (LCB) acquisition then spends each remaining run where it most reduces error. The surrogate, interchangeable among a Gaussian process (GP), random forest (RF), or multi-layer-perceptron (MLP) ensemble, plugs into the same LCB loop. Across six heterogeneous real-world scenes, FLAT-GP achieves the lowest scene-averaged behavioral error, winning 6/6 scenes against SPSA, GA, and CMA-ES and 5/6 against TPE under the matched budget. Some baselines need up to 4.4 times more simulations to match. Ablations show objective choice shifts final behavioral error by 81% on average, removing sequential LCB raises the six-scene mean by 20%, and surrogate choice shifts it by at most 4.2%, confirming gains trace to objective geometry and sequential allocation rather than surrogate capacity.
The editor's rubric
| Dimension | Level | Weight | What that level means |
|---|---|---|---|
| Leverage | ███░░ 3 | 18% | A method or resource many groups across the field will adopt within a year. |
| Magnitude | ████░ 4 | 20% | A qualitative jump: a capability or regime that did not exist before. |
| Evidence | ███░░ 3 | 22% | Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data. |
| Novelty | ███░░ 3 | 22% | A genuinely new approach to an open problem. |
| Trajectory | ██░░░ 2 | 10% | Some room to improve with obvious engineering. |
| Stakes | ██░░░ 2 | 8% | Benefits a professional community (practitioners, clinicians, engineers). |
Editor’s rationale
Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); gains (x-fold, relative gain, won); novelty (alternative to status quo); verification (ablations, code released); scale (efficient).
How the score was computed
- Merit
- 6.0 / 10
- Adjusted merit
- 4.9 / 10
- Attention
- 0%
- Freshness
- 74%