How the benchmark was run
Methodology builds more trust than logos. Here is exactly what we compared, under what rules, and what we deliberately don’t publish.
The setup
A multi-week parallel benchmark on a real UK multi-drop operation. The published figures cover the steady-state window, after an initial calibration phase in which delivery turnaround times were still being aligned between the two systems. On each operating day the same live order set — stores, volumes, delivery windows — was planned twice: once by the incumbent, an industry-standard automated planning system running in production, and once by VroomDesk. Both plans were produced for the same depot, the same fleet and the same constraints.
Ground rules
- Identical inputs. Same orders, same stores, same volumes, same delivery windows — no curated subsets.
- Identical constraints. Same fleet and vehicle types, same driver duty-hour limits, same depot departure rules.
- Baseline from the incumbent’s own outputs. The comparison uses the figures the incumbent system itself reported — not our re-simulation of its plans.
- Every day in the window counts. All operating days in the steady-state window are included — no cherry-picking. The earlier calibration phase, while turnaround times were still being aligned, is excluded and disclosed here.
Results (indexed, incumbent = 100)
| Metric | Automated planner | VroomDesk | Change |
|---|---|---|---|
| Distance | 100 | — | −X.X% |
| Routes | 100 | — | −X.X% |
| Driver hours | 100 | — | −X.X% |
Totals over the steady-state window, expressed as an index against the incumbent’s reported totals.
What we don’t publish — and why
Absolute mileages, route counts, exact dates and the identity of the operation are withheld under confidentiality. Publishing them would identify the client. The percentages and the rules above are the part we can show; the full day-by-day breakdown is available under NDA in a pilot conversation.
One honest caveat
These results come from one network. Savings depend on the shape of your operation — drop density, window tightness, fleet mix. That is exactly what a 2–4 week pilot measures: your historical routes, replanned 1:1 against what you ran.
Want the same comparison on your own data?
Book a pilot