August 8, 2026 · Business
Rippling was burning 40 percent of its engineering budget on AI tokens. It fixed that without cutting anyone off.
The instinct when an AI bill spirals is to cap usage. Rippling's own numbers show why that would have been the wrong move.
In March, Rippling's CFO Adam Swiecicki walked into an executive meeting with a number that stopped the room: the company was on pace to spend 40 percent of its entire R&D headcount budget on AI tokens, growing 80 percent month over month. Chief Product Officer Matt MacInnis put it plainly: "We were incredulous." One engineer alone was running up $50,000 a month in model usage. Roughly 10 to 15 percent of employees accounted for 60 percent of the bill. TechCrunch reported the full story this week, and Rippling put out its own announcement on August 7.
My position: the interesting part is not that AI got expensive. It is what Rippling did instead of the obvious response, and what that says about how a business should handle a runaway AI bill.
The obvious move, when 10 to 15 percent of your people drive 60 percent of a cost line, is to cap them. Set a token budget per seat, throttle the outliers, done. That instinct is not dumb. A skewed spend distribution like that is the classic signature of waste, and MIT's Project NANDA found in its widely cited 2025 study of enterprise AI deployments that 95 percent of generative AI pilots produce no measurable return on the company's books. If most AI spend is waste, a hard cap is the rational response.
Rippling didn't cap usage. It built an AI Spend Console: a dashboard tying every dollar of model spend to the employee, team, and business system that produced it (GitHub pull requests, Salesforce deal activity), paired with an internal gateway that routes each prompt to the cheapest model capable of handling it instead of defaulting everyone to the priciest frontier option. By July, token volume was flat versus the March peak of 605 billion tokens (July ran about 600 billion), but the bill for that volume came in at 37 percent of what April cost. Overall spend as a share of the R&D budget dropped from 40 percent to 15 percent. Check the math: 15 divided by 40 is 37.5 percent, matching the reported 37 percent same-volume cost comparison almost exactly.
That is the tell. If the problem had been waste, cutting total volume would have been the lever that mattered, and usage would have dropped with the cost. It didn't, it stayed essentially flat. The entire savings came from changing which model handled each request, not how many requests got made. The heavy users weren't necessarily wasteful; many were defaulting to the priciest model for tasks a cheaper one could handle just as well. A blunt cap would have throttled real output alongside the waste, hitting hardest the employees who, per MIT's framing of the "GenAI divide," most resemble the minority that actually extracts value by going deep.
MacInnis's diagnosis of why this problem exists is the sharper point, and it is the free-market case in one sentence: "inference providers, like Anthropic and OpenAI, have absolutely no incentives to help you control your spend." That's not a knock on those companies, it's just their business model: a frontier lab sells tokens, and has no reason to tell you a cheaper model would do. The fix for that misaligned incentive never needed to come from the vendor, or from a regulator. It came from a third party building better tooling, and from competition among models cheap enough to route to in the first place. DeepSeek's V4 Flash, released in public beta on July 31, prices output at $0.28 per million tokens, with cache-hit input near $0.003 per million on repeated context. When the routing option is that cheap, defaulting every task to the priciest frontier model gets harder to justify on the merits, not because anyone mandated it.
One caveat worth stating plainly: this works because Rippling could measure output. Engineering has pull requests as a legible proxy for value produced. TechCrunch notes the company is still working out the equivalent signal for sales, support, or marketing, where productivity doesn't reduce to a commit count. A lot of the functions burning AI budget right now lack that clean signal, and importing this playbook wholesale into one of those roles will just produce a dashboard nobody trusts.
For a small or midsize business without Rippling's engineering bench, the lesson isn't "build your own router." It's the diagnostic step that came before it: before capping anyone's AI usage, find out whether the cost comes from how much people use or from which model they use for tasks that didn't need it. Those are different problems with different fixes, and confusing them is how you cut your best output along with your waste. That diagnostic, plus routing routine work to smaller models while reserving frontier spend for what needs it, is core to what we do in MojoAI engagements. If your AI bill is climbing faster than you can explain, that conversation is worth having before you reach for a cap. Get in touch.
Sources
Every factual claim above is drawn from these independently published sources, linked inline where first referenced.
Tell us what's on your mind.
You don't need a polished brief to reach out. A two-line email about what's bugging you is plenty; we'll tell you straight if we're the right fit, and what we'd tackle first.
We'll scope the work around your workflow, goals, and timeline before quoting anything, so you know what's included before committing.