Back to blog
AI & Automation

How Much Does AI-Powered Dispatch Cost When You Pay the Model Provider Directly?

Amit Saini

August 19, 2026

When you bring your own model provider, AI dispatch cost stops being a mystery inside a software subscription and becomes a line on a bill you can read. Here is how to estimate it, and how to lower it.

Key takeaways
  • AI dispatch cost is calls multiplied by tokens multiplied by rate. Once you know how many AI calls your operation makes a day, roughly how long each is, and what your provider charges, the estimate is arithmetic.
  • Most calls are short and cheap; a few are long and matter. Address cleanup and note classification are hundreds of tokens each. Copilot and agent reasoning runs into thousands per call because tool results are included.
  • Under illustrative assumptions, a 500-order operation lands around a cent per order. The exact figure depends on your provider rates, which you should check, but the order of magnitude is small relative to a failed delivery.
  • Model tiering is the biggest lever. Routing short tasks to a small model and reserving the large model for reasoning typically removes most of the cost without changing outcomes.
  • BYO-LLM makes the number visible. With your own provider account the spend appears on your dashboard by day and by model, so you can tune it instead of guessing.

AI dispatch cost is one of the first questions an operations lead asks and one of the last a vendor answers clearly. When the model usage is buried in a per-seat subscription, the honest answer is that nobody outside the vendor knows. When you bring your own model provider, as Geofleet now allows, the cost becomes a line on your provider bill, and you can estimate it before you turn anything on.

This guide walks through which Geofleet AI features consume tokens and roughly how, a worked example for a 500-order-a-day operation with every assumption labelled, the levers that reduce the bill, and how to keep it visible over time. We do not quote real provider list prices as facts, because they change; we use placeholder rates and tell you where to check the current ones.

Which Geofleet AI features consume tokens

A token is roughly three quarters of a word. Every AI call sends a prompt (input tokens) and receives a response (output tokens), and providers price the two separately. What drives cost is not the number of features but the shape of the calls each one makes.

FeatureTriggerTypical prompt shapeApprox. tokens per callModel tier
Address cleanupEach imported order with an ambiguous addressRaw address in, normalised address out300 to 500Small
Note classificationEach order with a free-text noteNote in, one or two labels out200 to 400Small
Dispatch copilotEach dispatcher questionQuestion plus several tool results3,000 to 8,000Mid or large
AI agentsEach exception or SLA event handledMulti-step reasoning with tool calls5,000 to 15,000Large
MCP-driven assistantsEach cross-system promptPrompt plus results from multiple servers5,000 to 12,000Mid or large
Report narrationEach generated summaryNumbers in, short prose out1,500 to 3,000Mid
Cost drivers by Geofleet AI feature (approximate, illustrative)

Two features that people assume are AI-heavy are not: auto allocation and route planning are optimisation engines, not language models, and do not consume tokens. Our auto allocation explainer covers how that works.

The three variables in the estimate

AI dispatch cost per month equals the sum over features of calls per month, multiplied by tokens per call, multiplied by the rate per token for the model used. Each variable comes from a different place.

  • Calls per month: from your own operation. Orders per day drive the short calls; dispatcher headcount and exception rate drive the long ones.
  • Tokens per call: from the table above, refined after two weeks of real usage on your provider dashboard.
  • Rate per token: from your provider price list or your negotiated agreement. Check the current published rates for each model tier; they change.
Token-flow diagram showing Geofleet AI features feeding into a token meter split between a small model tier and a large model tier, flowing to the tenant’s own provider bill, with cost levers listed alongside
Where the tokens come from, which tier they land on, and the levers that shrink the flow before it reaches your provider bill.

Worked example: a 500-order-a-day operation

Everything below is illustrative. The volumes are plausible for a mid-size regional operation and the rates are placeholders, not any provider’s price list. Swap in your own numbers.

Assumptions

  • 500 orders a day, 26 operating days a month.
  • Address cleanup on every order at 400 tokens; note classification on every order at 300 tokens. Both on the small tier.
  • 60 copilot questions a day at 6,000 tokens each; 40 agent runs a day at 8,000 tokens; 30 MCP assistant prompts a day at 10,000 tokens. All on the large tier.
  • Placeholder blended rates: $X per million tokens on the small tier and $Y per million on the large tier. For a rough feel we use $0.50 for X and $5.00 for Y, purely as round illustrative figures. Check current published rates.

Arithmetic

  • Small tier: 500 times (400 + 300) = 350,000 tokens a day, about 9.1 million a month. At $0.50 per million that is about $4.55 a month.
  • Large tier: (60 times 6,000) + (40 times 8,000) + (30 times 10,000) = 980,000 tokens a day, about 25.5 million a month. At $5.00 per million that is about $127 a month.
  • Total: roughly $132 a month, or about one cent per order, under these assumptions.
~1 cent per order

Illustrative result

Under the placeholder assumptions above, a 500-order-a-day operation spends on the order of $130 a month on model usage. Your figure will differ with real rates and real usage, but the order of magnitude is what matters: it is small next to the cost of a single failed delivery.

The example also shows where the money is. The small tier handles 1,000 calls a day for a few dollars a month. The large tier handles 130 calls a day and accounts for almost all of the bill. Every lever below aims at that large-tier line.

For scale, our real cost of failed deliveries guide puts the price of one failed attempt at many multiples of that per-order AI figure. If the copilot and agents prevent a handful of failures a week, the AI line pays for itself several times over.

Levers that cut AI dispatch cost

Model tiering

The largest lever by far. If the example above ran every call on the large tier, the small-tier work alone would cost ten times more. Route address cleanup and classification to a small, fast model and keep the large model for copilot, agents, and MCP reasoning. With bring-your-own-LLM you choose the model or deployment for each, and on Azure OpenAI that maps naturally to two deployments, as our Azure guide describes.

Prompt length

Long calls are long because tool results are included. Returning only the fields a question needs, capping list sizes, and summarising route events instead of dumping them can halve the input tokens on copilot and agent calls. This is also good security practice, as our guardrails guide explains under PII minimisation.

Caching

Many prompts share a large, stable prefix: the system instructions, tool definitions, and tenant configuration. Several providers offer prompt caching that charges less for repeated prefixes. Where your provider supports it, structure prompts so the stable part comes first, and the saving applies automatically to every copilot and agent call.

Batching

Address cleanup and note classification do not need to run one order at a time during import. Batching twenty orders into a single call reduces per-call overhead, and where your provider offers a discounted batch mode for non-urgent work, overnight cleanup of tomorrow’s orders can use it.

LeverApplies toTypical effectEffort
Model tieringShort structured callsRemoves most of the small-task costConfiguration only
Prompt lengthCopilot, agents, MCPCuts input tokens on the most expensive callsPrompt and tool response design
CachingAny call with a stable prefixLower price on repeated instructions and tool definitionsProvider setting plus prompt ordering
BatchingImport-time cleanup and classificationFewer calls; possible discounted batch rateScheduling change
Rate limits and capsEverythingBounds the worst caseProvider and Geofleet settings
Cost levers and where they apply

"We assumed AI would be the expensive line. After the first month on our own key it was smaller than our SMS bill, and moving classification to a small model made it a rounding error."

— Operations director, regional grocery delivery

How BYO-LLM makes the cost visible

When Geofleet runs through your own provider account, configured under Settings, Developer, Integrations, in the LLM model providers section, every call is billed to you at your rate and shows on your provider dashboard. You can see tokens by day and by model, set a monthly cap, and receive alerts at thresholds you choose. There is no vendor markup on model usage and no blended margin hiding the heaviest users.

  • Create a dedicated project or deployment for Geofleet so its usage is separated from any other AI work in your organisation.
  • Set a spend cap for the first month at two to three times your estimate, then tighten it once real usage is known.
  • Review tokens per feature after two weeks and update the estimate. Real usage is the only number that matters.
  • If you have committed spend with a provider, this usage draws it down instead of adding a second bill.

Put the estimate next to the ROI

Run the arithmetic above with your own volumes, then compare it against the savings from fewer failed deliveries, faster dispatch, and reduced manual reporting. Our ROI calculator gives you the other side of the equation.

A quick estimate template

  1. Count orders per day and operating days per month.
  2. Estimate daily copilot questions (dispatcher headcount times questions each), agent runs (exception rate times orders), and MCP prompts.
  3. Multiply each by the tokens per call from the cost-driver table, split by tier.
  4. Look up the current published rate for each tier on your provider, or use your negotiated rate.
  5. Multiply, sum, and divide by monthly orders for a per-order figure.
  6. Enable BYO-LLM, run for two weeks, and replace the estimates with dashboard numbers.

That is the whole method. The point of paying the provider directly is not only that it is usually cheaper. It is that you can do this arithmetic at all, and keep doing it as the operation grows.

Frequently asked questions

How much does AI dispatch software cost per order?

When you pay the model provider directly, the AI component is a function of calls, tokens per call, and your provider rate. Under illustrative assumptions for a 500-order-a-day operation, it works out to roughly a cent per order. Your figure depends on current provider rates and your actual usage.

Which AI features in dispatch software use the most tokens?

Copilot questions, AI agent runs, and MCP-driven assistant prompts, because each includes tool results and multi-step reasoning, typically several thousand tokens per call. Address cleanup and note classification are a few hundred tokens each and cheap on a small model.

Does route optimisation or auto allocation consume LLM tokens?

No. Those are optimisation engines rather than language models and do not make model calls. Token cost comes from the copilot, agents, MCP assistants, address cleanup, note classification, and report narration.

What is the best way to reduce LLM API cost in logistics?

Model tiering: route short, structured tasks to a small model and keep the large model for reasoning. After that, trim prompt length by returning only needed fields, use prompt caching where your provider supports it, and batch import-time cleanup.

Does Geofleet charge a markup on AI usage with bring-your-own-LLM?

No. With your own provider and key configured under LLM model providers, the provider bills you directly at your rate. Geofleet does not add a margin on model usage, and the spend appears on your provider dashboard.

How do I control AI spend before I know real usage?

Create a dedicated project or deployment for Geofleet at your provider, set a monthly cap at two to three times your estimate, add alert thresholds, and review tokens per feature after two weeks. Then tighten the cap to match real usage.

Explore the Geofleet command center

See how AI agents can optimize your operations. Start free or book a walkthrough with our team.