Your company
March 2026
Mid-market tier. Spend rose modestly, led by Engineering on frontier models.
Most tokens go to Engineering.
Efficiency improved as more input was served from cache.
Sample brief
Same work. Unoptimized teams pay several times more per token.
4.15×
higher cost per 1M tokens for unoptimized teams versus efficient ones.
Higher cache reuse is one reason efficient teams pay less.
What separates them
Inefficient
Pays too much
Average
Some waste
Efficient
Spends carefully
Find your statistics
Connect your data to understand your AI spend, trends, and savings opportunities
The middle of the pack
2.05%
of tokens are fast or priority.
Anthropic · Fast mode
0.65%
OpenAI · Priority tier
3.67%
Unit economics
Draft a 150-word cold outreach email to this prospect referencing their recent press release. Map their news to our Enterprise Tier value propositions.
Prompt caching
Adding things like ids, names, or numbers to the prefix leads to cache invalidation and missed savings.
Token governance
Get control over AI spend without slowing adoption. Manage everything in one place, broken down by team, model, and workload, with spend limits, alerts, and more.
GC-02 / ALERTS
Elizabeth Smith's API Key
Total spend in June: $9,994.35
Set spend limit
Amount
Frequency
Notify me when spend reaches
Elizabeth Smith will always receive alerts. Choose who else to notify.
Stories
Benchmarks, playbooks, and hard-won lessons on how the fastest teams measure, route, and govern their LLM usage.
A big AI bill doesn't mean you're using too much AI. It means you're buying it wrong. How to separate wasteful spend from value with smarter defaults, the right model tiers, and a culture of efficiency.
OpenAI used to write to the token cache for free. With the 5.6 models that changed — and the new cache-write fee quietly reshapes what efficient usage looks like.
Businesses using AI
50.4%
Half of businesses on Ramp now pay for AI. The strongest predictor of who adopts isn't industry or city — it's who funded the company.
Tips & tricks