Back to blog
Operations

SLA Breach Runbooks: Prioritizing Delivery Exceptions Before They Escalate

Aditya Singh

April 12, 2026

Prevent SLA failures with structured runbooks, clear alerts, and prioritized exception handling.

Key takeaways
  • Separate hard breach from soft risk. A recoverable late delivery needs immediate intervention; an unavoidable breach needs notification and rescheduling. Conflating them wastes escalation capacity.
  • Intervene in the 60–90 minute window. That is the last point a resequence, replacement driver, or window extension still works. Alert at the moment of breach and every option is gone.
  • Prioritise by impact, not queue order. FIFO fails under pressure. A three-tier P1/P2/P3 system weighted by penalty risk and category criticality protects the deliveries that matter most.
  • Notify customers in parallel. Send the proactive update at the same moment the dispatcher is alerted — before the customer notices — to cut inbound complaints and preserve goodwill.
  • Runbooks improve through review. Weekly post-incident analysis turns recurring zone or driver patterns into structural fixes rather than repeated firefighting.

Every delivery operation breaches SLAs. The difference between operations that contain the damage and operations that spiral into customer churn is not the frequency of breaches—it is the speed and quality of the response. A well-designed runbook does not prevent every breach, but it ensures that when a risk emerges, the right people know, the right actions are taken in the right order, and the customer is informed before they have to ask.

Building that system requires three things: a precise taxonomy of breach types, an exception prioritization framework that reflects business impact rather than queue order, and an escalation chain with clear ownership at each step. This guide covers all three.

Classifying SLA Breach Types

Hard Breach vs Soft Risk

A hard breach is a delivery that cannot meet its committed window under any realistic intervention—the driver is too far away, the time remaining is insufficient, or a blocking event (vehicle breakdown, access issue) has made completion impossible. The only response is customer notification, rescheduling, and a root-cause log. A soft risk is a delivery that is trending toward a breach but is still recoverable: the driver is running 20 minutes behind and could catch up with a route resequence, a replacement driver, or a customer window extension. Treating soft risks as hard breaches wastes escalation resources. Treating hard breaches as soft risks creates false confidence and delayed customer notification.

Visibility Gaps as a Breach Category

A third category that most runbooks miss: visibility breaches. A driver who has not pinged location in 12 minutes may be fine or may be stranded. An order with no scan event in 40 minutes may be on track or misrouted. Visibility gaps are not confirmed SLA risks, but they are unresolved uncertainties that require a check-in action. Include them in your runbook with a defined response window—contact driver within 5 minutes of the gap, escalate to supervisor if no response in 3 minutes.

  • Hard breach: delivery cannot meet committed window — notify customer, reschedule, log root cause
  • Soft risk: delivery trending late but recoverable — intervene immediately with resequence or replacement
  • Visibility gap: no location or scan event within threshold — check in with driver, escalate if no response
  • Partial breach: multi-item order with some items confirmed, others delayed — notify on partial and set expectation for remainder

"The biggest improvement we made to our breach management was separating 'soft risk' from 'hard breach' in our alerting. Before that, everything looked like a crisis. After, dispatchers could focus on the things they could actually fix."

— Dispatch Operations Lead, Regional E-Commerce Operator

Building Early-Warning Detection Thresholds

The 90-Minute Rule

The practical window for meaningful intervention is 60–90 minutes before a committed delivery time. At that point, a dispatcher can still resequence stops, dispatch a replacement driver from a nearby hub, or call the customer to extend the window with goodwill intact. At 15 minutes before breach, the only realistic option is customer notification. Your alerting thresholds must surface soft risks at 90 minutes or earlier—not at the moment of breach.

60–90 min

Intervention window

Surface soft risks at least an hour before the committed window. That is the difference between a resequence that saves the delivery and a notification that only manages the damage.

Configuring Risk Flags

Set ETA drift thresholds per delivery tier. For same-day deliveries with tight 2-hour windows, flag any driver whose real-time ETA for an upcoming stop is more than 15 minutes beyond the committed time. For next-day standard deliveries, flag at 30 minutes drift. For time-critical deliveries (medical, regulated, high-value), flag at 10 minutes. The threshold should reflect the consequence of a breach in that tier, not a uniform rule applied across all orders.

  1. Define breach consequence tiers: critical, standard, low-priority
  2. Set ETA drift alert thresholds per tier (10 / 15 / 30 minutes)
  3. Configure automatic flag to dispatcher dashboard at threshold breach
  4. Set a secondary alert to supervisor if dispatcher has not acted within 10 minutes
  5. Log all flags and responses for weekly review

Exception Prioritization: Impact Over Queue Order

Why First-In-First-Out Fails Under Pressure

When exceptions queue up during a peak period or weather event, dispatchers default to FIFO—handling the oldest flag first. This is operationally intuitive but economically wrong. A 20-minute delay on a regulated pharmaceutical delivery or a high-value B2B contract has a materially higher cost than the same delay on a standard consumer order. Your prioritization framework must weight by business impact, not arrival time in the exception queue.

A Practical Prioritization Matrix

Assign each order a priority score at dispatch based on three factors: contract penalty risk (does this customer have a breach penalty clause?), delivery category criticality (regulated, medical, perishable, standard?), and downstream impact (does this delivery unlock a next step for the customer that has its own deadline?). When exceptions queue, surface the highest-score items first. The matrix does not need to be complex—even a three-tier system (P1 / P2 / P3) consistently applied beats FIFO.

TierExamplesETA drift flagEscalation
P1Regulated, medical, penalty-clause contracts10 minAccount owner looped at supervisor step
P2High-value B2B, same-day, high-LTV repeat15 minSupervisor if dispatcher idle 10 min
P3Standard next-day, flexible windows30 minDispatcher handles, logged for review
Exception priority tiers and response
  • P1: Regulated deliveries, penalty-clause contracts, medical/perishable items
  • P2: High-value B2B accounts, same-day consumer commitments, repeat high-LTV customers
  • P3: Standard next-day consumer, low-value orders, flexible window deliveries
  • Override rule: any visibility gap on a P1 order escalates automatically to supervisor
  • Review priority assignments weekly—tier creep (everything becoming P1) is a common failure mode

Escalation Chain and Customer Notification

Define Ownership at Each Escalation Step

An escalation chain without named owners is a responsibility diffusion mechanism. Every step in your runbook must have a specific role—not a team, a role—who is accountable for that action within a defined time window. Dispatcher acts within 10 minutes of flag. If unresolved, supervisor acts within 5 minutes. If unresolved, operations manager is notified. For P1 orders, the customer account owner is looped in at the supervisor escalation step, not after the breach is confirmed.

Proactive Customer Notification Reduces Complaint Volume

The notification trigger belongs inside the runbook, not outside it. When a soft risk is flagged and the 90-minute intervention window is open, the customer notification should go out simultaneously with the dispatcher alert—not after resolution is attempted and failed. Customers who receive a proactive update before they notice a delay have a dramatically lower escalation rate than customers who reach out to ask where their delivery is. Build the notification step into the runbook as a parallel action, not a sequential one.

Post-Incident Analysis That Actually Improves Runbooks

A runbook that is never updated is a policy document, not an operational tool. Weekly review of breach incidents—how many, what tier, what root cause, how long from detection to resolution—is the mechanism that keeps runbooks current. Look for patterns: if the same zone appears in breach logs three weeks in a row, the issue is structural (capacity, routing, hub distance) not execution. If the same driver appears repeatedly, it is a coaching or assignment issue. Post-incident analysis is how you convert reactive breach management into proactive prevention.

Frequently asked questions

What is the difference between an SLA hard breach and a soft risk?

A hard breach is a delivery that cannot meet its committed window under any realistic intervention—notification and rescheduling are the only options. A soft risk is a delivery trending toward a breach that is still recoverable with immediate action: resequencing, a replacement driver, or a customer window extension. Runbooks must treat these differently or waste escalation resources on unrecoverable situations.

How early should an SLA risk alert be triggered?

Alert thresholds should surface soft risks 60–90 minutes before the committed delivery window. That is the practical intervention window. For tight same-day 2-hour windows, flag at 15 minutes ETA drift. For next-day standard, flag at 30 minutes. For critical or regulated deliveries, flag at 10 minutes. Alerting at the moment of breach eliminates all intervention options.

How should delivery exceptions be prioritized when multiple flags appear simultaneously?

Use a three-tier priority system based on business impact: P1 for regulated, penalty-clause, and medical deliveries; P2 for high-value accounts and same-day consumer commitments; P3 for standard next-day orders. Surface P1 exceptions first regardless of queue order. Review tier assignments weekly to prevent everything drifting into P1.

When should the customer be notified about a potential SLA breach?

Notify customers proactively when a soft risk is flagged and the intervention window is open—do not wait for confirmation that the breach is unavoidable. Proactive notification before the customer notices a delay reduces inbound complaint volume by 30–50% and preserves goodwill. Build the notification trigger as a parallel action in the runbook, not a sequential step after resolution is attempted.

How does Geofleet help manage SLA breaches and exceptions?

Geofleet's dispatch platform provides configurable ETA drift alerts per delivery tier, automatic exception flagging to the dispatcher dashboard, priority queue ordering by contract and category, and built-in customer notification triggers. The AI copilot surfaces at-risk deliveries before the breach window closes and suggests specific interventions—resequence, reassign, or notify—based on current driver positions and remaining windows.

Explore the Geofleet command center

See how AI agents can optimize your operations. Start free or book a walkthrough with our team.