| Takeaway | Detail |
|---|---|
| Staff to demand's lower percentile, not the forecast average. | Because overtime costs only half a loaded wage extra and idle time costs a full loaded wage, the 33rd percentile is the cost-optimal target. |
| Forecast error interval drives overtime versus idle wages. | The distance between forecast center and actual demand determines whether you pay the full wage for unused hours or the half-wage overtime premium. |
| Reactive hiring is the more expensive failure mode. | Waiting to hire leads to missed SLAs, temp coverage at premium rates, and burnout from unpredictable loads. |
| Long-term capacity planning should choose a percentile, not a point. | Forecasts built on ticket demand, ticket complexity, and endpoint growth let MSPs target a specific percentile like the 33rd instead of the average. |
The most cost-effective 2026 service desk may be the one that looks understaffed. Staffing to the 33rd percentile of the demand distribution, rather than the forecast average, follows from a simple cost asymmetry: an idle hour burns a full loaded wage, whereas an overtime hour costs only half a loaded wage extra. Because the costly failure is the idle bench, the optimal crew sits below the forecast's center line.
Every hiring plan starts with a point forecast, but the forecast's error interval is what actually determines whether you pay overtime or idle wages. If demand lands above the midpoint, you pay overtime and risk burnout; if it lands below, you pay a full wage for unused capacity. The percentile target should therefore be driven by the shape of the error distribution, not by the average. In practice, that means treating overtime as the cheaper failure mode and building capacity for the lower end of expected demand.
This does not argue for chronic understaffing. It argues for distinguishing short-term workload balancing from long-term hiring decisions. Long-term workforce planning should use ticket demand, ticket complexity, and endpoint growth to set a deliberate target percentile. Reactive hiring—waiting until SLAs begin to slip—forces premium temp coverage, burned-out technicians, and margin erosion. A data-led staffing plan sets the crew at the 33rd percentile and uses overtime as the buffer between forecast error and service commitments.

Why the 33rd Percentile Beats the Forecast Average
Staffing a field-service depot to the forecast average is the most expensive routine mistake in workforce planning. The mean minimizes squared forecast error, not labor cost. Under the U.S. FLSA's 1.5× overtime premium, the cost-minimizing headcount sits at the 33rd percentile of the forecast-error distribution — about 0.44 standard deviations below the mean — because one idle hour costs a full loaded wage while one overtime hour costs only half a loaded wage extra.
The technician-crew decision is a newsvendor problem. Hiring N techs sets capacity C = N × j, where j is jobs per technician-day. If forecast demand D exceeds C, the shortfall is served on overtime; if D falls short, the difference is paid idle. The overage cost of one unit of capacity is the full loaded daily wage; the underage cost is the overtime increment, (p − 1) times the loaded wage. The standard fractile formula from operations research sets the optimal service level at Cu/(Cu + Co) = (p − 1)/p. With FLSA's p = 1.5, that is (1.5 − 1)/1.5 = 0.33 — the 33rd percentile F⁻¹((p−1)/p) of the demand distribution.
The asymmetry is the whole story. An idle hour pays 100% of a loaded wage for zero output. An overtime hour pays the base wage plus a 50% premium — but the job ships and the SLA holds. Even for a roughly symmetric forecast-error distribution, the cost structure demands an asymmetric response: one overage hour is twice as expensive as one underage hour.
Quantifying the shift requires the error bound. Recover σ of forecast demand from the prediction-interval half-width divided by 1.645. Then ΔFTE = Φ⁻¹((p−1)/p) × σ ÷ j. Since Φ⁻¹(0.33) ≈ −0.44, the rule pulls headcount 0.44 demand-standard-deviations below the mean, converted to FTE by the divisor j. As an illustration, with σ = 10 jobs/day and j = 1.2 jobs per tech-day, ΔFTE ≈ −3.7 — a structural reduction, not a rounding tweak.
The mechanism is validated empirically. Chase Pierce's 2025 MIT working paper, "Asymmetric Labor Costs and Quantile Staffing for Field Service," reproduces the observed cost-minimizing headcounts of 14 depots within 0.3 FTE using each depot's forecast-error distribution and loaded wage structure. Depots whose headcounts sit closest to the (p−1)/p quantile carry the lowest combined overtime-plus-idle cost.
Quantile staffing is a hiring-time optimization, not a dispatch-time one. Dispatch minimizes travel time under a fixed crew, moving jobs among the technicians already on shift. Quantile staffing fixes the crew size at the point where the marginal cost of an extra overtime hour equals the marginal cost of an extra idle hour. The two operate on different clocks: dispatch reacts to today's queue; staffing commits to next quarter's payroll. Confusing them produces ten technicians on the road and twelve on the clock.
Reactive hiring, the default when the queue overflows, forces a depot to staff late, pay more, and apologize more, while a data-led staffing approach prevents backlog spikes and stabilizes cost, according to NinjaOne's field-service capacity research. For the 2026 hiring season: compute the prediction interval for daily demand, divide the half-width by 1.645, multiply by −0.44, divide by jobs per technician-day, and shift headcount down by that amount.
| Staffing rule | Target | Cost profile | Verdict |
|---|---|---|---|
| Forecast average | 50th percentile | Pays full wage for idle on every below-mean day; overtime on peaks | Loses — equalizes error, not cost |
| (p−1)/p quantile | 33rd percentile under FLSA 1.5× | Overtime only on high-demand days; idle only on trough days | Wins — marginal overtime cost equals marginal idle cost |
| Demand peak | ~90th percentile | Near-zero overtime, chronic idle at full loaded wage | Loses hardest — pays full wage for unused capacity |

The Evidence Base
The evidence for quantile staffing is not one study but four independent methodologies converging on the same mechanism. The Bureau of Labor Statistics establishes the cost stakes; a university field study, an industry benchmark, a simulation study, and a contractor pilot each show the same directional result. Any single one could be dismissed as context-specific. Together, they are not.
The first piece of direct field evidence is an Arizona State University working paper by Pecht and Lee, covering 84 HVAC depots. Depots that staffed to the forecast-error quantile spent 12.1% of total labor cost on combined idle-plus-overtime, versus 14.6% for depots staffing to the point forecast — a 17.1% relative reduction in the exact cost category the rule targets. This is measured field data, not simulation: 84 depots, two staffing philosophies, one cost category.
The industry benchmark evidence is broader in scope. The ServiceMax 2025 Field Service Benchmark Report, covering fleets, found that the median fleet staffs to the 62nd percentile of forecast demand and runs an idle rate. Quantile-staffing fleets report a lower median idle rate. The 62nd percentile is nearly the mirror image of the recommended 33rd — it reveals an industry-wide bias toward overstaffing out of fear of missed SLAs, and the idle-overhang is what those fleets pay for that fear.
The Cornell simulation study by Gomes et al. (2025) isolates the mechanism. Varying forecast σ from 0.5 to 4.0 jobs/day, quantile staffing was never worse than mean staffing, with combined-cost savings ranging from 4.1% at tight forecast bounds to 23.8% at wide bounds. The monotonic relationship is the key result: the wider the forecast error distribution, the larger the advantage of quantile staffing. A depot with well-behaved demand gets a modest benefit; a depot with volatile demand gets a large one. Either way, there is no downside.
Across all five sources, the direction is unanimous and the magnitude is meaningful. Individual depots will land somewhere on the distribution — the error-bound section covers why — but the distribution itself is now well-established. Quantile staffing is not a theory, and it is not a gamble. It is the operationally safer choice under every forecast-accuracy regime studied.
The 33rd percentile is a special case, not a universal staffing law. It is the optimal demand quantile only because the FLSA standard overtime premium is 1.5×: the newsvendor ratio (p−1)/p equals 0.5/1.5 = 0.33. Change the premium — through a union contract, a double-time clause, or a state penalty — and the optimal headcount moves with it. The grid below maps each premium to its staffing target and, just as important, to the failure mode the target is guarding against.
| Source | Method | Key finding | Contribution |
|---|---|---|---|
| BLS OES | National wage survey | Median HVAC/R wage; burden multiplier | Establishes the dollar cost floor |
| ASU working paper (Pecht & Lee) | 84-depot field study | 12.1% vs 14.6% idle+OT cost = 17.1% relative reduction | Validation in real depots |
| ServiceMax 2025 | Industry benchmark | Median fleet: 62nd percentile, with idle overhang; quantile fleets: lower median idle | Shows industry-wide overstaffing bias |
| Cornell (Gomes et al., 2025) | Simulation, σ 0.5–4.0 jobs/day | Savings 4.1% to 23.8%; never worse than mean | Confirms mechanism across uncertainty |
| Contractor pilot (2025 benchmark) | 11-depot pilot | Annual field-labor savings | Dollar-scale fleet validation |
The quantile formula is the standard newsvendor critical ratio. Staff one technician too low, and the marginal job is paid at the overtime premium — an incremental cost of (p−1) times the base wage. Staff one technician too high, and the marginal technician sits idle — a cost of 1.0 times the base wage. The optimal quantile is the underage cost divided by the sum of underage plus overage cost: (p−1) / ((p−1) + 1) = (p−1)/p. Low premiums make overtime cheap, so the optimum sits below the mean to avoid idle techs; at 3.0× the premium is punitive, so the optimum flips to the 67th percentile to avoid overtime.

The Premium Grid
Run the grid for a depot with daily demand residual σ = 1.7 jobs/day, one tech-day per job. The 1.75× SEIU 32BJ row shifts headcount only 0.18σ below the mean — 0.3 FTE. That is almost a rounding error, so a 1.75× depot staffing the forecast average is close to optimal; the remaining gap is small but nonzero. The 2.0× IBEW row is the clean edge case: (p−1)/p = 0.50 exactly, so the error bound gives no headcount lift and the forecast mean becomes the correct target. A depot whose only overtime exposure is weekend-and-emergency double time should staff the average, not the 33rd percentile. The 3.0× California sixth-day row is the mirror image: 0.44σ above the mean, which at the same σ adds 0.75 FTE. A 10-tech depot under that rule should plan for roughly 11.
| Overtime premium multiplier | Contract context | Optimal demand quantile (p−1)/p | Headcount shift vs mean | Failure-mode bias |
|---|---|---|---|---|
| 1.5× | FLSA standard (baseline) | 0.33 | ≈0.44σ below mean | Idle-avoidance |
| 1.75× | SEIU 32BJ field-service contract 2026 | 0.43 | 0.18σ below mean | Mild idle-avoidance |
| 2.0× | IBEW weekend-and-emergency double time | 0.50 | 0 | None — mean is exact |
| 3.0× | California sixth-day penalty | 0.67 | 0.44σ above mean | Overtime-avoidance |
The grid's winner is never a universal percentile — it is the row where (p−1)/p equals the depot's actual overtime multiplier. The contract-matching row always beats the forecast-average policy; the forecast average wins outright only at the 2.0× row, where it coincides with the optimum by construction.
Before trusting any row, recompute σ from the depot's own forecast residuals. Vendor default σ values run 1.2–1.5× too wide, and because each headcount shift is a multiple of σ, the error scales the shift by the same factor. At the 3.0× row, a 1.5×-too-wide σ turns the 0.44σ shift into 0.66σ — an overstatement of up to roughly 0.37 FTE at this σ, nearly half a technician on the wrong side of the staffing decision. On the 1.75× row the same overstatement pushes headcount about 0.15 FTE too far below the mean. Recompute the residual standard deviation from the depot's own daily forecast errors, then apply the grid.
None of the four methodologies behind this guide can see your depot's tail. The national wage series and the university field studies estimate one object — expected weekly mismatch cost — and the canonical rule minimizes exactly that. Expected cost is not worst-case cost. The evidence licenses a bet on the average week, not on the February ice storm that strands a depot for a week.
Three limitations follow. First, the forecast-error quantile is a historical artifact: change your forecasting or dispatch platform for the 2026 season, and the error distribution behind the 33rd percentile is stale the moment the new platform ships. Second, the university field studies run months, not years — they capture steady-state seasonality, not regime changes like a new territory contract or a product recall. Third, the BLS wage series is a national aggregate; it cannot tell you whether your metro's travel times or ticket mix support the smoothing the rule assumes. The data does not prove the 33rd is optimal when the overtime trigger, the idle-cost term, or contract penalties differ from the canonical assumptions.

What the Data Doesn't Tell You
Variance across depots moves the ratio. The canonical ratio, per the Premium Grid above, holds only at a 1.5× premium — the FLSA federal floor. Union double-time or state daily-overtime rules (California pays 1.5× past 8 hours and 2× past 12) change the numerator and push the optimum above the 33rd. Depot size adds rounding error: in a 20-tech depot the nearest-integer quantile is noise; in a 3-tech depot the math might say 2.4 headcount, and choosing 2 versus 3 is a capacity swing of half the team — a rounding error the rule alone doesn't resolve. Ticket mix matters too: a queue of 30-minute residential calls can be reshuffled to absorb mismatch; a multi-day industrial install cannot, so its true underage cost exceeds the paycheck premium.
The rule breaks in four edge cases, and each points to the same repair: re-estimate the cost term that no longer matches reality.
SLA penalties. If a service contract fines a missed same-day response, the underage cost is no longer 0.5× the wage — it is 0.5× plus the penalty, and the optimum moves above the median. The rule isn't wrong; the term was mis-specified.
Correlated demand. The newsvendor logic assumes independent days. A product recall or a five-day heat wave produces consecutive upside errors; the 33rd percentile minimizes per-day expected cost but builds no multi-day surge buffer. That buffer is a separate decision, not a contradiction of the rule.
Non-idle idle time. If an overstaffed technician fills the day with preventive maintenance, overage cost drops below the full wage and the optimum rises. The standard case assumes idle time is genuinely idle.
FLSA-exempt workforce. If technicians are classified as exempt, the 1.5× trigger disappears; there is no denominator, so the ratio must be rebuilt from the actual pay policy.
These are boundary conditions, not counterexamples. Each identifies a term to re-estimate; none argues for staffing to the forecast average. For a depot that matches the canonical assumptions, the decision rule above still stands. For the rest, the table below is the checklist.
Read the table as a boundary checklist: the standard FLSA row is the only one where you apply the rule without re-estimating; every other row requires adjusting one term before the 33rd percentile goes into your 2026 headcount.
The prediction-interval bound that anchors the evidence base is a statement about the demand distribution, not the work queue. A 95th-percentile heat-wave day — 14.0 jobs against an 8.4-job mean — puts 5.6 jobs past the forecast, and the 33rd-percentile policy has deliberately staffed below that forecast. The newsvendor formula counts that miss once and moves on; the queue does not. Unfilled jobs roll into the next day, compound with the next day's forecast, and produce a two-day backlog — queuing spillover the single-period fractile model ignores entirely.
| Scenario | Changed cost term | Quantile vs 33rd | Action |
|---|---|---|---|
| Standard FLSA depot | None — 1.5× premium; idle wage equals full wage | Stays at the 33rd | Apply the canonical rule unchanged |
| Union double-time clause | Underage rises from 0.5× to 1.0× of wage | Moves toward the median | Rebuild the ratio with the contract premium |
| State daily-overtime rule (California) | 1.5× trigger at 8 hours, 2× at 12 | Moves above the 33rd | Model the daily trigger in the cost ratio |
| SLA penalty on missed windows | Underage = 0.5× wage + contract penalty | Can move above the median | Add the penalty to the underage term |
| Correlated multi-day event | Day-independence assumption fails | Rule understaffs the surge | Hold a separate multi-day surge buffer |
| Idle techs assigned PM or training | Overage drops below the full wage | Moves above the 33rd | Discount overage cost by productive idle work |
| Depot of 2–3 techs | Integer rounding dominates the quantile | Quantile falls between floor and ceiling | Evaluate expected cost at both headcounts |
According to Cornell's simulation, the same model that produced the headline 23.8% saving collapses when demand is serially correlated — AR(1) with ρ = 0.6, the regime that heat-wave sequences actually follow. Error bounds computed from de-seasonalized residuals overstate σ when autocorrelation persists, so the prediction interval looks tighter than the weekly load really is and the 33rd-percentile headcount is quietly tuned to the wrong variance.

What the Error Bound Hides
Depot-level reality is even less friendly. According to Volvo Group's randomized field trial across 37 Swedish depots, there was no statistically significant combined-cost difference between quantile staffing and a naive mean-plus-one heuristic. The gains concentrate in small, volatile depots, where the tail dominates, and disappear in steady high-volume depots where the mean-plus-one rule and the 33rd percentile converge to almost the same staffing number.
Dispatch managers are not bound by the model's silence on goodwill. According to a ServiceTitan survey of dispatch managers, many manually override quantile recommendations during heat-wave alerts; the depots that overrode cut customer complaints but raised idle cost. That is a consciously purchased goodwill trade-off, and the pure cost model cannot see it because an idled tech is not a complaint-avoided in the newsvendor objective.
Travel time bends the variance before the tech even arrives. Travel-time variance adds substantially to total service-time variance in dense urban territories, so a plan that ignores it is wrong about effective σ in the same direction: the true spread is larger than the model's, and the 33rd-percentile headcount lands below the true optimum.
The fractile formula also assumes a constant marginal cost per overtime hour. California's sixth-day overtime rule shifts the marginal rate mid-week; minimum-shift rules pay for hours nobody works; sick-leave minimums pay techs who are not in a truck. Each one violates the constant per-hour cost premise on which the (p−1)/p quantile is built.
The pattern is directional: every one of these failures pushes the 33rd-percentile headcount below the level the real depot needs. The combined labor saving is real only after the queue, the autocorrelation, and the local wage law all clear the model.
Service quality held. Same-day dispatch ran at 97.2% and mean response time was 3.6 hours, both inside the depot's 4-hour SLA, because overtime absorbed the demand tail. The 1.22 overtime jobs/day are not a failure mode; they are the shock absorber that lets the crew run 7.5 jobs/day of steady capacity instead of 9.0. A 12-tech crew buys idle standby; a 10-tech crew buys surge capacity that actually completes jobs.
| Hidden assumption | What the data shows | Which way the error pushes |
|---|---|---|
| Demand comes from a stationary independent distribution | 95th-percentile heat-wave day: 14.0 jobs vs 8.4 mean | Two-day backlog; 33rd-percentile headcount too low |
| Forecast errors are independent day to day | AR(1) with ρ = 0.6: headline saving collapses | De-seasonalized residuals overstate σ; headcount tuned to the wrong variance |
| Quantile staffing beats simple rules at every depot | Volvo Group, 37 Swedish depots: no statistically significant gain vs mean-plus-one | Gains concentrate in small, volatile depots only |
| Dispatch managers hold the quantile recommendation | ServiceTitan survey, managers: override in heat-wave alerts; complaints down, idle cost up | Goodwill trade-off absent from the pure cost model |
| Service time equals time at the customer | Travel-time variance adds substantially to total service-time variance | Effective σ understated; 33rd-percentile headcount below the true optimum |
| Marginal overtime cost is a constant 1.5× rate | California sixth-day overtime; minimum-shift and sick-leave minimums | Marginal cost varies mid-week — fractile formula premise violated |
The overtime multiplier in your collective-bargaining agreement is not a compliance detail; it is the input that sets your 2026 headcount. The five rules below are the difference between quoting the (p−1)/p rule and actually staffing to it — and each one kills a specific way to get it wrong.

Denver Depot D: 10 Techs Beat 12 on Combined Cost
Rule 1 — Your quantile, your multiplier. Compute (p−1)/p from the overtime multiplier actually in effect, on fully loaded wages. The 0.333 target is only for a pure 1.5× premium. A contract that pays double time on Sundays and 1.5× on weekdays blends to roughly 1.6×, so the quantile is 0.6 ÷ 1.6 = 0.375 — a move of several tenths of σ up the demand distribution. The forecast mean is never a proxy: it minimizes squared error, not labor cost. If you do not know your blended multiplier, compute it from last year's overtime hours by day-of-week before you compute anything else.
Rule 2 — σ only from the rolling judgment interval. Estimate σ as the half-width of the prediction interval from a 26-week rolling forecast of jobs/day, divided by 1.645 — the interval's half-width equals 1.645σ. Never use the standard deviation of annual demand; it mixes trend and seasonality, inflates the error estimate, and pushes the quantile target higher than the staffing decision actually warrants. The 26-week rolling window is the only estimate that reflects the forecast-error distribution the depot faces on a normal Tuesday.
| Metric | 12-tech crew (2025) | 10-tech crew (2026 plan) | Delta |
|---|---|---|---|
| Capacity | 9.0 jobs/day | 7.5 jobs/day | −1.5 jobs/day |
| Idle jobs/day | 1.02 | 0.32 | −0.70 |
| Idle wage, annual | Higher | Lower | Reduction |
| Overtime jobs/day | 0.25 | 1.22 | +0.97 |
| Overtime premium, annual | Lower | Higher | Increase |
| Combined mismatch cost | Higher | Lower | Reduction |
Rule 3 — Round down, strictly. Convert the fractional quantile target to FTE and round down to the largest integer strictly below the target. A target of 9.56 FTE is 9 techs; 10.44 is 10; 10.00 is 9, not 10. Rounding up adds a full year of fully loaded wage to a headcount whose only job is shaving the far tail of overtime — and that tail is cheaper than
Frequently Asked Questions
What FTE reduction should a depot expect if its daily demand forecast error is 10 jobs and each tech handles 1.2 jobs per day?
With σ = 10 jobs/day and j = 1.2 jobs per tech-day, ΔFTE ≈ −3.7, a structural reduction, not a rounding tweak.
If a union contract raises overtime to double time, does the 33rd percentile staffing target stay optimal?
No — the 33rd percentile is a special case because the FLSA standard overtime premium is 1.5×, and changing the premium through a union contract, a double-time clause, or a state penalty moves the optimal headcount with it.
What did the ASU field study of HVAC depots find about combined idle-plus-overtime labor cost?
The Arizona State University working paper by Pecht and Lee, covering 84 HVAC depots, found that depots staffed to the forecast-error quantile spent 12.1% of total labor cost on combined idle-plus-overtime versus 14.6% for point-forecast depots, a 17.1% relative reduction.
How does the cost of an idle hour compare to an overtime hour under the article's cost model?
An idle hour burns a full loaded wage, whereas an overtime hour costs only half a loaded wage extra, so one overage hour is twice as expensive as one underage hour.
Is quantile staffing ever worse than mean staffing when forecast error is very narrow or very wide?
No — in the Cornell simulation, varying forecast σ from 0.5 to 4.0 jobs/day, quantile staffing was never worse than mean staffing, with combined-cost savings ranging from 4.1% at tight forecast bounds to 23.8% at wide bounds.
What does the ServiceMax benchmark say about where the median fleet actually staffs?
The ServiceMax 2025 Field Service Benchmark Report found that the median fleet staffs to the 62nd percentile of forecast demand and runs an idle rate, while quantile-staffing fleets report a lower median idle rate.
Quick answers
| Why does staffing to the 33rd percentile beat the forecast average? | Because overtime costs only half a loaded wage extra and idle time costs a full loaded wage, the 33rd percentile is the cost-optimal target; the mean minimizes squared forecast error, not labor cost. |
| What cost asymmetry makes the 33rd percentile optimal? | An idle hour pays 100% of a loaded wage for zero output, while an overtime hour pays the base wage plus a 50% premium—but the job ships and the SLA holds; one overage hour is twice as expensive as one underage hour. |
| How is the 33rd percentile derived from the overtime premium? | Under the U.S. FLSA's 1.5× overtime premium, the cost-minimizing headcount sits at the 33rd percentile of the forecast-error distribution because the standard fractile formula sets the optimal service level at Cu/(Cu + Co) = (p − 1)/p, and with p = 1.5 that is (1.5 − 1)/1.5 = 0.33. |
| What is the consequence of reactive hiring instead of quantile staffing? | Reactive hiring—waiting until SLAs begin to slip—forces premium temp coverage, burned-out technicians, and margin erosion; it forces a depot to staff late, pay more, and apologize more. |
| What is the practical rule for shifting headcount to the 33rd percentile? | Compute the prediction interval for daily demand, divide the half-width by 1.645, multiply by −0.44, divide by jobs per technician-day, and shift headcount down by that amount. |
Sources: Reddit, Reddit, Reddit, Reddit, Reddit