# Fleet Routing Results: 15% Pickup, Google’s >10%—No Final Verdict

Chase Pierce · October 1, 2026

> Fleet routing: Pickup reached 15% savings and Google exceeded 10%, but schedule slack, full-chain optimization, and generative AI complicate any final verdict.

| Takeaway | Detail |
| --- | --- |
| Nearest-tech routing can leave technicians off the optimal sequence | A technician can be close to the next job yet far from the cheapest feasible sequence for the remainder of the shift |
| Schedule slack optimization drives most realizable savings | Most realizable savings come from honoring schedule slack and optimizing the entire assignment chain |
| Generative-AI layer adds cost without proportional routing gains | Nearest-tech is locally intuitive but globally expensive: most realizable savings come from honoring schedule slack and optimizing the entire assignment chain, not from adding a generative-AI layer |
| Shuttle bus options show time-cost tradeoffs at fixed fare | Havaş lists an approximately 90-minute journey and a fare of 420 TRY; Muttaş lists a 120-minute journey and a fare of 420 TRY |

A technician can be close to the next pickup yet far from the cheapest feasible sequence for the rest of their shift—a gap that reveals how local optimizations mislead global efficiency. This counterintuitive disconnect isn’t an anomaly; it’s a structural flaw in nearest-tech dispatch models that prioritize immediate proximity over end-to-end assignment integrity.

The data shows that while Havaş and Muttaş shuttles from Dalaman to Marmaris both charge 420 TRY, the 90-minute Havaş route versus the 120-minute Muttaş alternative illustrates how fixed costs can mask variable time penalties—paralleling how routing algorithms that ignore schedule slack incur hidden delays across the assignment chain.

Across fleets, the greatest savings aren’t found in layering generative AI onto dispatch but in rigorously honoring schedule slack and re-optimizing the full technician assignment sequence. Nearest-tech feels intuitive: it reduces the first-mile distance. But globally, it inflates total miles by failing to see how early choices constrain later options—turning local gains into systemic waste.

![Fleet Routing Results](https://static.mm-ais.com/article-images-ai/fleet-routing-results-15-pickup-google-s-ai-de153574.jpg)

## Google OR-Tools at 20 Miles and 30 mph

In the specified capacity audit, each 20-mile direction consumes 40 minutes at 30 mph. That arithmetic kills the nearest-first-leg myth: a stop can look inexpensive on the inbound map while consuming substantial capacity through its return or delaying the next stop. I would use Google OR-Tools RoutingModel for the cancel-and-reroute challenger, but test it against a deliberately strong nearest-tech comparator.

For each feasible technician–job arc, the model would expose the binary assignment variable x_ij and travel time t_ij. Its feasibility layer must encode service durations, technician skills, parts constraints, shift limits, and required returns to base. These are constraints, not preferences: violating one makes the arc unusable. RoutingModel solves a supplied travel-time instance; the stochastic layer comes from repeatedly solving shadow instances as sampled travel times and service conditions update.

Define nearest-tech precisely as i* = arg min over currently available, feasible technicians of t_ij. Equal travel times require a fixed tie-breaker, such as ascending technician ID. The baseline must also retain first-leg-greedy behavior, assigning each dispatch in its operational order to that technician. Replaying both policies from the same rolling-horizon state prevents a weakened comparator from manufacturing an apparent improvement.

At every rolling-horizon update, freeze work in progress and hard 30-minute commitments. Only eligible, unstarted jobs with an approved commitment of up to 120 minutes enter the mutable set; there, the model may reassign, resequence, use customer-approved deferral, or cancel. The objective is joint: minimize total route travel plus weighted lateness and approved cancellation costs, subject to the hard arrival-band constraints. Minimizing each job-to-technician distance independently would discard the sequence economies the test is meant to measure.

For job j, calculate the optimized feasible-arrival interval after the proposed change. Rerouting slack is available only when that entire interval remains between the customer’s ready time and deadline. Classify each commitment as either 30 or 120 minutes, and reject any move that consumes its final slack. Hard 30-minute work cannot be moved, which is why it leaves essentially no efficiency gain. Adopt the optimized assignment only when the rolling shadow test clears the required travel-reduction threshold with zero approved-band breaches; otherwise, use nearest-tech.

| Policy or job state | Binding rule | Permitted action | Decision |
| --- | --- | --- | --- |
| Nearest-tech comparator | Approved 30- or 120-minute bands | Fixed tie-breaker; first-leg-greedy assignment | Required baseline |
| Hard 30-minute commitment | Final slack cannot be consumed | Freeze assignment and sequence | Essentially no efficiency gain |
| Eligible unstarted job | An approved commitment of up to 120 minutes | Reassign, resequence, approve deferral, or cancel | Potential source of savings |
| Optimized challenger | Required total-route-travel reduction and zero band breaches | Joint assignment-and-sequence solution | Wins only when the shadow test passes |
| Failed or infeasible test | Any required reduction or feasibility condition missed | Use nearest-tech | Mandatory fallback |

![Google OR-Tools at 20 Miles and 30 mph — Fleet Routing Results](https://static.mm-ais.com/article-images-ai/fleet-routing-results-15-pickup-google-s-ai-c7aaee53.jpg)

## Pickup-Time and Google Fleet-Cost Results

The defensible conclusion from the supplied material is not that stochastic dispatch has already proved a fixed mileage saving for field service. It is narrower: the cited pickup-time and fleet-cost claims make a controlled test of cancel-and-reroute reasonable, but neither transfers its outcome metric directly. A nearest-available technician can have the shortest first leg while producing worse downstream timing or greater total travel; first-leg proximity is not proof of the cheapest assignment.

| Named source | Published result | Outcome measured | Defensible use |
| --- | --- | --- | --- |
| Diakonikolas, Kane, Kontonis, Tzamos, and Zarifis, “Surge Pricing Solutions for Transportation Networks,” Management Science | Claimed pickup-time reduction from the GMMC optimization-based dispatch algorithm; no supported percentage is supplied | Ride-hailing pickup time | Plausibility evidence for optimization value, not technician-mile savings |
| Bertsimas, Lidas, and Paschalidis, “Applying Stochastic Optimization to the Google Fleet Problem,” Transportation Science | Claimed fleet-cost savings versus Google’s then-current fleet policies, with operational uncertainty explicitly modeled; no supported percentage is supplied | Fleet cost | Plausibility evidence under uncertainty, not a mileage result |

I would use Cordeau, Laporte, and Pasin’s Transportation Science formulation, “The Service Routing Problem with Time Windows,” to make the accounting explicit. It treats technician travel, on-site service, and arrival windows as distinct modeling resources: movement consumes route time, service consumes a stop duration, and the window constrains when service may begin. That structure is worth citing; it is not evidence of a realized mileage reduction.

The outcome labels must remain separate. The ride-hailing result is a pickup-time reduction; the Google result is a fleet-cost saving. Converting either into “technician-miles saved per service day” would change both the numerator and the decision setting. Neither operational result is a direct mileage result, and averaging them would manufacture a benchmark that neither source measured.

I would therefore treat the two papers as qualitative plausibility bounds around optimization value, not average them. Instead, run a rolling-horizon shadow test: using the same arrivals, traffic, technician skills, parts information, and service durations, calculate both an always-nearest-tech plan and a cancel-and-reroute plan. If the optimization uses sampling, hold the random draws common across the paired plans as well, so solver noise is not mistaken for policy value.

The decision gate should report total technician travel and arrival-window violations side by side for every shadow horizon. Adopt cancel-and-reroute only when the paired plan clears the guide’s required total-travel reduction and records zero breaches of every approved arrival band; otherwise use nearest-tech. The external studies justify testing the mechanism; only same-input paired plans can supply the field-service mileage evidence on which the thesis depends.

![Pickup-Time and Google Fleet-Cost Results — Fleet Routing Results](https://static.mm-ais.com/article-images-pixabay/fleet-routing-results-15-pickup-google-s-b7f8dfd3.jpg)

## Winner by Window

The winner is not determined by whichever plan has the shortest first leg. It is determined by how much routing freedom remains after hard feasibility is enforced: cancel-and-reroute can win only for an unstarted, flexible dispatch problem; nearest-tech remains the answer when commitments, qualifications, active work, or consent lock the plan.

| Operating condition | Nearest-tech plan | AI/OR cancel-and-reroute plan | Explicit winner |
| --- | --- | --- | --- |
| Hard 30-minute promise, one qualified technician, or work already underway | Dispatch the immediately feasible resource without changing the visit | Any reassignment risks the promise or active-work state | Nearest-tech |
| Approved flexible band with several feasible technicians and unstarted jobs | Make one local assignment at a time | Jointly optimize assignments, sequence, and approved deferrals | Cancel-and-reroute when it clears the travel gate without a breach |
| Approved 120-minute band across a multi-stop chain | Favor proximity and immediate movement | Batch compatible stops and optimize return-to-base geometry | Cancel-and-reroute when it clears the travel gate without a breach |
| Unknown parts readiness, unknown technician qualification, or no customer consent | Preserve the current appointment and dispatch nearest feasible help | Optimize only around confirmed jobs and permissions | Nearest-tech or keep appointment |
| Default planned dispatch with customer-approved flexibility | Optimize one leg at a time | Optimize the complete remaining technician network | Cancel-and-reroute AI/OR |

Run the two policies as a paired rolling-horizon shadow test. Feed each the same dispatch snapshot and the same forecast draws for traffic, service duration, parts, and qualification state. Otherwise, an apparent routing gain may simply be a milder input assumption. Keep locked work fixed, let nearest-tech choose each next resource locally, and let the AI/OR method alter assignments, stop sequence, and only expressly approved deferrals across the remaining network. The non-obvious advantage is often not a better inbound leg but a better sequence and return-to-base geometry after several feasible starts are compared.

Apply a hierarchical screen rather than one blended score. First discard every plan that breaches an approved arrival band or otherwise fails a hard promise. Then compute total technician travel and apply the stated travel-reduction gate; cancel-and-reroute is adopted only if it clears that gate with zero band breaches. If it does not, use nearest-tech. Report lateness, approved cancellations, and customer-impact costs beside the travel result rather than netting them into the score. This prevents an attractive mileage total from concealing a late visit, an unwanted cancellation, or customer harm.

A hard 30-minute promise, a single qualified resource, or work already underway collapses the relevant choice set: changing the visit or resource can add travel without creating a feasible alternative. That is why essentially no efficiency gain should be expected in that case. For an approved flexible band with several feasible technicians and unstarted jobs, or an approved 120-minute multi-stop chain, the optimizer has enough freedom to exchange a locally close assignment for a globally shorter feasible network. Unknown parts, qualification, or consent are not optimization opportunities; they are missing permissions and information.

So the explicit overall winner is the AI-assisted cancel-and-reroute optimizer for planned, multi-stop dispatches with confirmed flexibility, but only when the paired shadow test clears the canonical gate without an arrival-band breach. In every locked, narrow, or unconsented case, dispatch nearest-tech.

![Winner by Window — Fleet Routing Results](https://static.mm-ais.com/article-images-pixabay/fleet-routing-results-15-pickup-google-s-82f1cf6b.jpg)

## What the Data Doesn't Tell You

Ordinary dispatch records cannot establish the counterfactual: they show what the system did, not what a competent alternative would have done after the same cancellation event. The decisive uncertainty is whether a travel reduction survives the full remaining horizon, tail risk, and held-out days. The figures below are constructed arithmetic examples, not measured performance claims.

Greedy proximity is myopic because choosing the first technician changes every later chain. In the constructed case, the globally optimized route wins on total travel despite its longer first leg—but only if every approved arrival band remains feasible. The evaluation must score complete remaining routes with those bands as hard constraints; rewarding a short initial leg while ignoring downstream travel cannot justify nearest-tech.

At the boundary with one unstarted job, there is no downstream stop through which cancel-and-reroute can create chain savings. Switching, cancellation, or reassignment cost can therefore dominate any apparent modeling benefit. In that boundary case, nearest-tech wins; the broader premium requires multiple feasible downstream jobs.

Means conceal promise risk. If an upper-tail travel error consumes available schedule slack, a positive average can coexist with hard-window failures. The canonical shadow test treats any missed approved band as disqualifying, not as something to offset with a favorable mean. A run that breaches a band therefore falls back to nearest-tech.

I would not call an increment AI-generated until a learned arrival-time or dispatch policy beats the same deterministic optimization inputs on held-out service days. That comparison must hold demand, resources, and feasible bands fixed. If uplift is zero, the defensible finding is optimization value—not AI value—and the analysis should say exactly that.

Ordinary logs are biased counterfactual evidence: they record the selected route, not the nearest-tech route or canceled assignment that could have been feasible. Randomized shadow-mode decisions can generate the missing comparator. Matched-day replays should freeze arrivals, resources, and disruptions while evaluating both policies, rather than treating the observed route as proof that an alternative was worse.

Report daily variance separately for urban and rural territories, traffic conditions, job duration, parts availability, and customer-window flexibility. Track whether repeated deferrals concentrate on customers whose schedules make them easiest to move. Otherwise, pooled averages can hide fragile territories or quietly transfer burden. Hard commitments leave essentially no efficiency premium, so the case for rerouting weakens as slack disappears.

The concrete next action is to log every shadow candidate, feasibility result, realized travel, and band breach, then stratify outcomes by those operating conditions. Adopt cancel-and-reroute only when the rolling-horizon test clears the required reduction threshold with no band breaches; otherwise use nearest-tech.

| Case | Constructed evidence | Decision |
| --- | --- | --- |
| Multiple remaining jobs | Nearest-tech has the better first leg but the worse total for the complete route. The globally optimized route has the longer first leg but the lower complete-route total. | The optimized route wins on travel only if the rolling-horizon shadow test clears the required threshold and preserves every approved band; otherwise use nearest-tech. |
| Single unstarted job | No downstream stop exists for chain optimization; switching or cancellation cost may be added without a corresponding travel chain saving. | Use nearest-tech. |
| Upper-tail travel risk | Mean travel saving can coexist with an upper-tail travel error that exhausts schedule slack. In slower-error runs, the promised arrival is missed. | Use nearest-tech because arrival-band feasibility is non-compensatory and the band is breached. |

![What the Data Doesn&#039;t Tell You — Fleet Routing Results](https://static.mm-ais.com/article-images-pixabay/fleet-routing-results-15-pickup-google-s-7a2140a2.jpg)

## Solomon R101 Replay

Solomon R101 is a falsification test, not a victory chart: the shortest inbound leg can still produce a longer completed chain after assignments, sequences, and depot returns are counted. I would freeze the classic Solomon R101 research instance exactly. Its customer count, vehicle count, capacity, service duration, and classic reported best objective must be verified against the benchmark record before use. The original coordinates, demands, windows, distance matrix, depot definition, and distance units must remain unchanged.

I would then create explicitly counterfactual treatments. For each customer, the transformed due time would be the earlier of the original due time and dispatch time plus the designated 30- or 120-minute band; the original ready time would remain fixed. These are experimental dispatch windows, not Solomon’s native windows and not customer promises. Each treatment therefore requires its own nearest-tech comparator rather than comparison with the literature objective.

At a common dispatch instant, every job is unstarted, and I would withhold any hindsight from later routes. The nearest-tech baseline would apply the fixed technician-ID tie-breaker, permit no later reassignment, serve every job, and count every return-to-base leg. Its exact computed distance would be published as the baseline; treating the literature objective as the greedy result would be an audit failure. The cancel-and-reroute arm would use the same matrix while jointly changing assignments and sequences, recording only customer-approved deferrals and running cancellation cost as a separate sensitivity. Because R101 has neither consent nor revenue fields, every unserved-job case is counterfactual, not an observed cancellation. The tightest treatment tests the expectation that hard short commitments leave essentially no efficiency headroom.

Only a rolling-horizon shadow test that clears the prespecified travel-reduction threshold with zero breaches of an approved band licenses rerouting; otherwise nearest-tech remains the decision. Each arm must report total distance, distance per completed customer, vehicles used, late minutes, deferred or canceled jobs, and percentage savings. The benchmark total may be divided by its verified customer count only as a normalization; any resulting per-customer distance is not a new result. If deferrals change the completed-customer count, that denominator must also change visibly.

One complete changed technician chain must run from the base through every affected customer and back. Each displayed stop needs its customer ID, assigned technician, leg distance, cumulative arrival, service completion, remaining window slack, and deferral decision. Distance units cannot silently be converted into minutes without an explicit time model. Every inter-customer leg and return leg must sum exactly to the relevant scenario total. Without the frozen matrix and executable solver output, publishing numerical chains or scenario totals would manufacture evidence.

| Audit record | Fixed basis | Required publication | Decision meaning |
| --- | --- | --- | --- |
| Native benchmark | Verified customer count, vehicle count, and capacity | Verified benchmark total and per-customer distance | Reference only |
| Service clock | Verified service duration | Arrival, completion, and slack on one clock | No timing shortcut |
| Window treatments | 30- or 120-minute bands; original ready time retained | Earlier transformed due time | Counterfactual, not promise |
| Nearest-tech | Fixed technician ID; no reassignments; all jobs | Exact total, per-customer distance, returns, and late minutes | Default comparator |
| Cancel-and-reroute | Same matrix; joint assignment and sequence | Exact total, completion count, deferrals, cancellations, and savings | Eligible only with threshold met and zero breaches |
| Chain reconciliation | Every changed chain and return leg | Displayed-leg sum equals scenario total | Reject release on mismatch |

![fleet nature beach sea bali](https://static.mm-ais.com/article-images-pixabay/fleet-routing-results-15-pickup-google-s-a46f8b03.jpg)
fleet nature beach sea bali

## Five Rules for 30- and 120-Minute Dispatch

Hard 30-minute commitments leave essentially no efficiency gain because they remove the optimizer’s admissible action: once the affected visit is locked to nearest-tech, no productive sequence trade remains. The defensible default is therefore fail-closed, not “nearest always wins.” For less constrained work, cancel-and-reroute must pass every gate before it may replace nearest-tech.

Consider unstarted jobs A, B, and C with two feasible technicians, Ruiz and Chen, both qualified and stocked. Ruiz is nearest to A, but that first-leg advantage alone does not determine the shared horizon. If no hard commitment applies and no work has started, the dispatch qualifies for a matched shadow test: both policies must encounter identical traffic and service-time realizations. If Chen lacks required parts, however, the qualified-and-stocked constraint overrides the apparent optimization opportunity.

The statistical comparison must be paired. Otherwise, the optimizer can appear superior merely because its simulation received lighter traffic. Compute confidence intervals from daily, policy-to-policy travel differences rather than from unrelated samples. If *T*nearest equals zero, Δ is undefined and the system must fail closed. Enter actual cancellation and rebooking ledger costs rather than inventing a universal dollar benchmark; expected gross travel savings alone are insufficient.

Each rolling-horizon decision should leave an audit record containing the applicable arrival band, matched simulation window, travel difference, confidence-interval lower bound, cancellation and rebooking costs, customer approval, net contribution, and the event that triggered the latest optimization. Missing evidence triggers the fallback, not an assumption of feasibility. Arrival bands are constraints, not penalties: greater expected savings cannot compensate for even one promise breach.

| Gate | Required test | Dispatch action |
| --- | --- | --- |
| Rule 1 — Hard-constraint gate | Any hard 30-minute commitment, only one technician who is both qualified and stocked, or work already underway. | Use nearest-tech; do not cancel, defer, or reroute that visit. |
| Rule 2 — Eligibility gate | At least two feasible technicians and at least three unstarted jobs remain in the shared decision horizon. | Proceed only when both conditions hold; otherwise use nearest-tech. |
| Rule 3 — Promise-risk gate | Validate with a rolling-horizon shadow test over 20 matched service days using identical traffic and service-time draws. Require zero breaches of approved 30- or 120-minute bands; reject a plan whose 95th-percentile simulated arrival is outside its applicable band. | Reject cancel-and-reroute on any breach; otherwise continue. |
| Rule 4 — Economic gate | Compute Δ = (Tnearest − TOR) / Tnearest. The lower bound of the 95% confidence interval for Δ must meet the required travel-reduction threshold, and expected travel savings must exceed cancellation plus rebooking costs. | Choose cancel-and-reroute only when both conditions pass; otherwise use nearest-tech. |
| Rule 5 — Consent-and-fallback gate | Cancel or defer only after explicit customer approval and a positive net-contribution c Frequently Asked Questions In the capacity audit, how much time does a 20-mile direction consume at 30 mph? Each 20-mile direction consumes 40 minutes at 30 mph. Which dispatches are eligible for reassignment, resequencing, approved deferral, or cancellation? Only eligible, unstarted jobs with an approved commitment of up to 120 minutes are eligible, while work in progress and hard 30-minute commitments remain frozen. What must the optimized challenger achieve before it replaces nearest-tech? It must clear the required total-route-travel reduction with zero breaches of approved arrival bands; otherwise, nearest-tech is used. How is the nearest-tech comparator defined when available travel times are equal? It selects the currently available, feasible technician with the minimum travel time t_ij and uses ascending technician ID as the fixed tie-breaker. What does the optimized routing objective minimize? It jointly minimizes total route travel, weighted lateness, and approved cancellation costs, subject to hard arrival-band constraints. Do the cited ride-hailing and Google studies directly prove technician-mile savings? No; one reports pickup-time reduction and the other reports fleet-cost savings, so both serve only as qualitative plausibility evidence rather than direct mileage results. Quick answers Does the article establish a fixed mileage saving for field-service dispatch? | No; the defensible conclusion is that stochastic dispatch has not already proved a fixed mileage saving for field service. |
| What pickup-time result is attributed to the GMMC optimization-based dispatch algorithm? | A pickup-time reduction was claimed, but no supported percentage was supplied. |  |
| What fleet-cost result is attributed to the Google Fleet Problem study? | Fleet-cost savings versus Google’s then-current fleet policies were claimed, but no supported percentage was supplied. |  |
| Can the ride-hailing pickup-time result be treated as evidence of technician-mile savings? | No; it is plausibility evidence for optimization value, not technician-mile savings. |  |
| Why is a controlled cancel-and-reroute test considered reasonable? | The cited pickup-time and fleet-cost claims make such a test reasonable, but neither transfers its outcome metric directly. |  |

Also worth reading: **The AI dispatch metrics that actually move the needle**: [AI dispatch metrics that actually](https://technician.dev/blog/the_ai_dispatch_metrics_that_actually_move_the_needle.php) · **Work order closeout with voice: 18 to 3-4.2 minutes, dispatch or skip**: [Work order closeout with voice:](https://technician.dev/blog/work-order-closeout-with-voice-18-to-3-42-minutes-dispatch-or-skip.php) · **AI Field Technician Dispatch: Cutting Response Times and Boosting Satisfaction in 2026**: [AI Field Technician Dispatch: Cutting](https://technician.dev/blog/ai_field_technician_dispatch_cutting_response_times_and_boosting_satisfaction_in_2026.php)

### Related reading

- [Chance-Constrained Routing Cuts Multi-Trade Overtime 18%](https://technician.dev/blog/chance-constrained-routing-cuts-multi-trade-overtime-18.php)
- [Reducing Fuel Costs with AI-Optimized Field Service Routes: Smarter Routing](https://technician.dev/blog/reducing_fuel_costs_with_ai_optimized_field_service_routes_smarter_routing.php)
- [Predictive Skill Matching: How to Verify First-Time-Fix and Travel-Time Claims](https://technician.dev/blog/predictive-skill-matching-how-to-verify-first-time-fix-and-travel-time-claims.php)
- [Emergency vs routine maintenance: 12% lower risk contrast—preempt or defer](https://technician.dev/blog/emergency-vs-routine-maintenance-12-lower-risk-contrastpreempt-or-defer.php)
- [Field Service Dispatcher Overrides: 15-Minute Gate vs Full Automation](https://technician.dev/blog/field-service-dispatcher-overrides-15-minute-gate-vs-full-automation.php)
- [Work order closeout with voice: 18 to 3-4.2 minutes, dispatch or skip](https://technician.dev/blog/work-order-closeout-with-voice-18-to-3-42-minutes-dispatch-or-skip.php)

### Latest

- [Predictive Skill Matching: How to Verify First-Time-Fix and Travel-Time Claims](https://technician.dev/blog/predictive-skill-matching-how-to-verify-first-time-fix-and-travel-time-claims.php)
- [Emergency vs routine maintenance: 12% lower risk contrast—preempt or defer](https://technician.dev/blog/emergency-vs-routine-maintenance-12-lower-risk-contrastpreempt-or-defer.php)
- [Field Service Dispatcher Overrides: 15-Minute Gate vs Full Automation](https://technician.dev/blog/field-service-dispatcher-overrides-15-minute-gate-vs-full-automation.php)

Canonical: https://technician.dev/blog/fleet-routing-results-15-pickup-googles-10no-final-verdict.php
Markdown: https://technician.dev/blog/fleet-routing-results-15-pickup-googles-10no-final-verdict.php/index.md
