You're Buying AI Wrong
A big AI bill doesn't mean you're using too much AI. It means you're buying it wrong. How to separate wasteful spend from value with smarter defaults, the right model tiers, and a culture of efficiency.
A big AI bill doesn't mean you're using too much AI. It means you're buying it wrong. How to separate wasteful spend from value with smarter defaults, the right model tiers, and a culture of efficiency.
A big AI bill does not mean you are using too much AI. It means you are buying it wrong.
Most companies I talk to have two contradictory problems: they overspend on the AI they use, but use far less AI than they should. Whether a bill feels big or small has less to do with the aggregate amount spent and more to do with your ability to justify the value it provided. You want to maximize dollars spent on productive AI and minimize dollars spent on wasteful AI. Everything else follows from that.
The good news is you can get spending under control with three things:
Start with a simple question: how much does $100,000 buy in AI? For the same spend you can buy about 5 billion tokens of the smartest model on the market — or roughly 42x more of an open-weight model.
That tells you what $100K buys, but not what it buys you. How much work will these tokens produce? Tokens measure usage, nothing more. They are inherently non-fungible and therefore the wrong tool for measuring business value. Ignore tokens and look at something more legible instead: the work your organization does with AI. A PR reviewed. An invoice coded. An outbound email sent. These are atomic units you can price, compare, and measure ROI from.
Your AI usage today is a mix of tasks that succeed and tasks that fail — but providers charge you by usage, not by outcome. To get the best out of that arrangement, maximize the success rate of your tasks while running them on the cheapest form of AI you can.
Cost = tasks attempted x cost per attempt
Value = successful tasks x value per task
If you have not seriously audited your AI usage, you are almost guaranteed to be on the wrong side of this equation.
Every AI request is shaped by three choices: which model you use, how hard it thinks, and how fast it responds. Providers make those decisions for you and set them as defaults because it simplifies the experience — but those defaults are optimized for broad adoption, not your economics.
Configure your defaults well and the same inputs can produce the same outputs for a tenth of the price.
The moment a new frontier lands, last generation's frontier becomes second-best at half the price. Most organizations passively ride these waves, but upgrading every workflow to the newest model "just because" is the most invisible way to burn money in software today.
Keep your defaults boring: start routine work on the cheapest model that passes your quality bar — even if it was yesterday's frontier — and move up only when the task demands it. Newer frontier models should expand what you try, not re-run the same work at a higher price. Using extra intelligence when a cheaper model produces the same outcome is just margin donated to your provider.
Modern models let you set how hard they think, and providers default to the highest setting. Each step up the scale roughly doubles the tokens burned — but the performance difference only shows up on extra-hard problems. On everything else, you are paying double for the model to double-check work it already got right.
Providers sell the same model at several speeds. "Fast mode" and "flex / batch mode" are simply deciding how quickly you want the same model, with the same intelligence, to respond. For many workflows a 2x premium to get the same result a few seconds sooner is a poor trade.
One simple test: is a human waiting on the answer? If yes, pay for speed. If your agent is responding to another agent, or running an automation nobody will look at until later, flex or batch mode is the right pick.
For the right tasks, the frontier is worth it. A more advanced model might surface an architectural decision that saves months, catch a security hole before it ships, or find an approach your team simply doesn't have. That single insight can create more value than all of your day-to-day usage combined.
Difficulty and value run in opposite directions. Easy tasks are common; hard tasks are rare — and almost all of your upside lives in the rare ones.
So sort everything into two buckets and stop deliberating:
| Bucket | Model | Reasoning | Spend |
|---|---|---|---|
| Easy work | Cheapest that clears your bar | Low–medium | Strict |
| Hard work | Best on the market | Max effort | Unlimited |
Smart defaults prevent waste from people who never touch their settings. But once someone starts making explicit choices, they are almost always incentivized to spend more. The engineer picking the model never sees the bill — they are judged on outcomes, not efficiency. The biggest model at the highest setting is the safe career move every time. Nobody has ever been fired for buying the frontier.
The companies that manage AI cost best build a culture where efficiency is treated as an engineering achievement, not a constraint. Teams celebrate getting the same outcome with a smaller model, less reasoning, or a slower tier just as much as they celebrate shipping the feature.
In practice, a handful of workflows drive the majority of your bill. You don't need to optimize everything — you need to find the ten lines that solve 80% of the problem.
Treat AI spend like a budget problem and you'll make people afraid to use AI. The bill shrinks, but so does the upside. Treat it as an operating-model problem instead: routine work done as cheaply as possible, expensive models reserved for the moments where more intelligence actually changes the outcome.
So yes — you are probably spending too much on AI. You are also probably using too little of it. Both are symptoms of the same problem: you haven't separated waste from value.
Start there.