Median monthly spend:$3.54K
$ / 1M tokens:$1.11
Efficiency cost gap:4.15×
Efficient $ / 1M:$0.72
Frontier model %:59%
Cache hit rate:78.74%
Fast / priority share:2.05%
Prompt caching savings:74.1%
Usage units:110.1T
Providers:11
Model families:16

What does efficient AI spending look like?

Data from 1500+ businesses up till July 14th 2026

The price of inefficient AI

Same work. Unoptimized teams pay several times more per token.

4.15×

Efficient teams use caching to reduce costs

Higher cache reuse is one reason efficient teams pay less.

Where the money goes

Inefficient48% frontier
Average68% frontier
Efficient51% frontier

What separates them

Same tokens, very different habits

What models you use and how you use them contribute heavily to how efficient you are with your tokens.
BEHAVIOR / CAL

Inefficient

Pays too much

Average

Some waste

Efficient

Spends carefully

Model choice
Rely on expensive, frontier models.
Lean on balanced models for most workloads.
Minimizes frontier model usage, relying on cheap efficient models as daily workhorses.
Cache use
Rarely caches tokens
Reuses cached input often, but less consistently than efficient teams
Reuses cached input very consistently
Instructions
Long and packed with unnecessary context
Efficient on some teams, overly long on others
Short, focused, and free of unnecessary context
Older models
Older, overpriced options are still widely used
Some remain in use while teams upgrade
Mostly replaced with newer, better-value options

Find your statistics

See where you stand

Connect your data to understand your AI spend, trends, and savings opportunities

The middle of the pack

Meet the average company

Average company profile

Median monthly
$3.54K
$ / 1M tokens
$1.11
Frontier model %
59%
Cache hit rate
78.74%

Top models by spend

100%all model spend
  • Claude Opus 4.819%
  • Claude Opus 4.615%
  • Claude Sonnet 4.613%
  • Claude Opus 4.712%
  • GPT-5.511%
  • Other models30%

How much usage runs in fast mode?

2.05%

of tokens are fast or priority.

Anthropic · Fast mode

0.65%

OpenAI · Priority tier

3.67%

Unit economics

Your model choice can really impact costs.

Monthly cost calculator

Model rate keys
Task inputs

Draft a 150-word cold outreach email to this prospect referencing their recent press release. Map their news to our Enterprise Tier value propositions.

Input tokens / task6,000
Cached input tokens / task30,000
Output tokens / task500
Tasks / month20,000
Monthly estimate
Sonnet 520,000 tasks
$460

Prompt caching

(Re)order the prompts.

Putting in constant reusable context first can lead to big savings.

Prompt order

Adding things like ids, names, or numbers to the prefix leads to cache invalidation and missed savings.

01System prompt · changing prefixDynamicCurrent Ticket to EvaluateTKT-88492 · Senior Graphic Designer · Primary 4K monitor flickers black every 10 seconds while using Adobe Premiere.
02System prompt · follows dynamic dataReusableTriage Rulebook & Context2,000+ words of routing rules, edge cases, SLAs, and priority matrices.
Cached input0
Input tokens3,000
Output tokens80
Cost / 10,000 tasks$68.00
Same result · { "department": "Hardware", "priority": "High" }

Token governance

See where every token lands

Get control over AI spend without slowing adoption. Manage everything in one place, broken down by team, model, and workload, with spend limits, alerts, and more.

Spend control

GC-02 / ALERTS

Elizabeth Smith's API Key

Total spend in June: $9,994.35

Set spend limit

Amount

$1,000

Frequency

Monthly

Notify me when spend reaches

50%80%90%100%

Elizabeth Smith will always receive alerts. Choose who else to notify.

Teammates
CancelUpdate

Stories

Field notes on AI spend

Benchmarks, playbooks, and hard-won lessons on how the fastest teams measure, route, and govern their LLM usage.

Tips & tricks

How to maximize value per token