Subscribe to the Next Level BizTech podcast, so you don’t miss an episode!
Amazon Music | Apple Podcasts | Listen on Spotify | Watch on YouTube
The conversation explores the challenges and implications of implementing AI tools in business, particularly focusing on budgeting issues that arise after initial success. It highlights a real-world example involving Uber, where the rapid adoption of an AI coding tool led to unexpected budget constraints, prompting a reevaluation of financial planning and resource allocation.
Video Transcript
Transcript is auto-generated.
Josh Lupresto (00:00)
So picture this: a company rolls out this awesome AI coding tool to, you know, 5,000 engineers back in December. Everybody loves it, productivity’s up, engineers are thrilled, everybody’s high-fiving, huge win. Then only four months later, four months, the CTO looks up and says, Uh-oh, my entire annual AI budget for the year is gone. Gone. And he says out loud, on record,
I’m back to the drawing board because the budget I thought I would need is blown away already. Four months into the year. not a made-up company. That is one that you know very well. Company is called Uber. So that actually happened this spring. and look, they’re not alone, right? We’ve got another one for you. There’s a company that spent five hundred million dollars in a single month because they turned on AI, no usage caps. Crazy. Five hundred million in a month. if you can’t tell.
Here’s why I’m fired up to talk about this. This is a brand new category of runaway spend. Just landed on your laps, landed in every one of your customers’ businesses. And right now, almost nobody is really governing it. And if you’ve been in this channel for any length of time, which a lot of you have, you’ve seen a lot of this and you know exactly what this means. That is not a problem. That is a doorway, and we love doorways. We jump in.
Welcome back, everybody. Next level biz tech. I’m your host, Josh Lupresto SVP of Sales Engineering at Telarus Today’s episode is titled The AI Bill Shock and the Runaway Cost that nobody, almost nobody, is governing. So, the biggest story in enterprise tech recently, it’s not about a new model. We’ve talked about that on some of the previous episodes. it’s just the bill that everybody got for all of the old ones. So
Let me let me let me set the stage here, set the table a little bit on what’s going on, because these numbers are insane. So for about two years, you know, kind of the unofficial enterprise AI strategy was super simple. Send every request to the biggest, most powerful model, and don’t ask too many questions, right? And and I know I’m talking about enterprise. I realize enterprise isn’t where everybody focuses, but I think it’s really important to understand what enterprise does because whatever they do.
We also see the SMB, the mid market, everybody follows suit. So it’s a great indicator of things to get in front of. So you think about, you know, they’re sending all these requests off. in the beginning, I think it was cheap enough to ignore. Nobody’s watching the meter, you know, the lovely power meter outside your house that just spins and spins and spins. two things happened at the same time. Usage exploded, and then the pricing model changed underneath everybody.
So we went from this kind of flat, predictable subscription to consumption-based pay per token billing. great for the model providers, right? But here’s the part that catches everybody off guard. So even the smart technical teams, per token prices have actually been falling. But the amount of tokens being consumed has been rising way, way faster than the price is dropping. So unit price goes down.
But the invoice triples. Why is that? Well, the models are getting better, and that’s because of agents. This is a key thing to understand. So when you type one question into a chat bot, that’s just a singular call, singular API request, right? But an AI agent, the thing that everybody’s deploying right now, doesn’t make one call. It takes the in the the inbound, it plans.
It pulls a bunch of context, it maybe calls some tools, it checks its own work, it retries if it didn’t meet the criteria, you know, it’s following those instructions, and then it just loops. So you think about just a single agentic task, that can fire anywhere from five to 30 model calls. And in some cases, that can burn through a thousand times the tokens in just a simple query, one query like that. So the customer thinks they bought a calculator.
And what they actually plugged in was a meter running at you know highway speed 24 hours a day. And look, this isn’t because you know, a a couple reckless startups. Some good data for you here. We love some numbers. 73% of enterprises reported that their AI costs came in over projection this year. 70%, 73%. So the discipline that companies built over the last decade to control project.
Cloud spending, what they call it? They called it FinOps. And so it’s the share of those professionals now responsible for AI. That jumped from about a third last year to basically all of them, 98%. Let’s look at GitHub, right? The place where everybody sends their code, it’s the code repository. So their own coding tool moved to usage-based billing. And then some of the power users saw those costs jump up to 50, 10 to 50 times.
What does this sound? You know, this sound familiar, right? We saw VMware licensing, we saw different changes, we saw all of those things. Business model changes, prices go up, customers have to figure out a path. So back to this, you know, costs going up 10 to 50 times. We got a name for it now, because we don’t introduce enough you know, acronyms and names. But here’s one for you, you’re gonna get used to. It’s called token maxing. Engineers you know these different orgs competing to burn the most tokens.
Well what’s the point here? The point of this is the technology worked great. The governance did not exist. It just moved too fast. And that gap, that gap between powerful tool and zero oversight of what it costs, that gap is exactly where a great advisor lives. let’s let’s think about part two here. why is this your home turf? So so let me connect this.
Directly to you because I want you to feel how this lands right in your wheelhouse. Think about, let’s get in the time machine here. think about where this channel started. Started in telecom, right? And what was one of the very first ways that a good advisor proved the value to a customer? You went in, you pulled the carrier bills, you found circuits they were paying for, maybe some things that they weren’t using, and you had these lines. Nobody’s audited these things in five years, right? You’re a hero. So it was all these overages that nobody was watching.
Now, you did an expense audit, maybe you right-sized it, you renegotiated it, you switched carriers. you saved them real money before you ever even really sold them anything new. All right, so let’s fast forward a little bit from that. Cloud comes along. All right, everybody, we gotta get off-prem. We gotta go to cloud. hyperscalers, everything’s private cloud. Let’s go. So that exact same story plays out. People move to AWS, they move to Azure.
And you know, first they said it was for the devs and for you know it wasn’t for production, and then all of a sudden it was. And the bill was consumption based, a little bit unpredictable, maybe spiraled for some, but what did it create? It created this whole new discipline. We called it out earlier, called FinOps. So it you know, it grew up to get it under control. And I think advisors were you were right there.
Helping these customers tag spend, you know, let’s right size the instances, maybe we can get you some reserved things like that. Let’s just let’s stop the bleed and help your team manage it because nobody’s been here before. Okay. So here’s the here’s the thing I want you to hear. AI spend is the third wave, I think, of this exact same movie, same plot, new technology.
And I think you’ve you’ve seen these before. You already know the ending, right? It is consumption based. Check. It’s unpredictable and spikes without warning. Check. nobody really centrally owns it. Check. Now the bill arrives and no one can tell you which team or which workflow drove it. Check. Kind of get the story here. That is the telecom bill in 1999. That’s the cloud bill.
In 2015. And that is the AI bill right now in 2026. Is it not? So you’ve done this before, twice, actually, if not more. the customer does not need you to be the AI researcher. they need you to be the person that walks in and says, Hey, let me find out where this is actually going and let’s get it under control.
Which is just the most natural thing for anybody to say. I love that that is where this channel has thrived and where it’s originated at. And think about that. Anything that you’ve said right there is not about replacing anything that you’ve already sell sold them. Your connectivity, your security, your cloud, your UCAS, all of that stuff. The SD WAN, it it all stays. This is a brand new line item that you get to add to the conversation. It’s purely additive. I and I was talking to a a a partner about this last night.
It is just a new reason why to get back in front of every customer you already have and every prospect you are chasing. Just like you have in all those original conversations that we started doing cloud and contact center and security, the customers were all going through it for the first time ever. So when you talk about these relevant points, you gauge their interest because they’re not seeing anybody yet within their own orgs that knows how to tackle these or even externally, right? So a great spot for you to come in.
All right, part three, the fix and what do you get to sell? So let’s get some let’s get some concrete here. Now, when you find this runaway AI bill, what’s the fix and what does that really open up for you to sell? two big levers here. They’re gonna sound pretty familiar because it’s kind of some of the same moves that we’ve already made, right? That you’ve made in years past and some of these other shifts. Lever one, right sizing.
In telecom, you didn’t, you know, you didn’t put every site on the biggest circuit. you from an AI perspective, you don’t need to send every task to the biggest, most expensive model. what a what a ton of companies are doing is routine ticket triage, document extraction, simple classification, right? Not the Ferrari, you don’t you don’t need the Ferrari model. it can often run on a smaller or cheaper, fine-tuned model.
For a fraction of the cost. If you want to help them get that built, get that spun up, get that managed, right? You can do that. and then look, you know, you’ve saved the frontier model for re the the really hard stuff. Okay, that was our first lever. Think about what’s our second lever here. This is putting a smart layer in front of everything that automatically sends each request to the right model that can do the right job. So with
You know, think about automatic failover goes to a backup provider, first one goes down or jacks up prices, no single vendor lock-in. and I think that’s kind of important right now. not to get too political, but as you think about if you’re helping guide these businesses, one model provider might feel some sort of way, one model provider might feel some other sort of way. If they are beholden to the beliefs of that model provider.
And then all of a sudden they decide in regulation or they decide that this is how we want to bias and weight the model really could hurt the customer’s business. And so you really don’t want that lock-in. This is a bad time to have that lock-in. Now, here’s the here’s the number that I think makes the whole business case for you. If you look at real enterprise data, companies that routed everything to frontier models were paying on average 18 to about 40-ish.
cents per million tokens. So these companies over here on the other side that are running a tiered routed setup, about two and a half bucks. So that’s almost eight times cheaper for the same work. Eight times. That is a number that you can walk into any CFO’s office and just put it right there on the table. So here’s I I guess here’s the other part that should get you excited because right sizing and and and routing
don’t just save money. They do a lot more than that. I think it gets the conversation, it it moves the conversation along. But they open up a whole shelf of brand new things for you to sell. It feels like the shelf is getting pretty big these days. I just picture, you know, we’re we’re opening up the coat, that guy in in, you know, New York that’s got all the watches, you go, my gosh, how much stuff do we really have here? So always look to things like the Tolera solution map and and and others, right? Is that constantly expands.
This is a whole new category of suppliers, or it’s an expansive product set from some existing suppliers that we have. Maybe now they’re gonna need AI gateways sitting in front of the models to help enforce some of those spending limits. FinOps for AI platforms, you know, that gives finance a real dashboard and kind of like this, you know, this per team chargeback. Maybe they want to look at different language models, smaller language providers, language model providers.
another one, we’re talking about this yesterday with some new tooling that came out with one of our suppliers. Browser-based security and monitoring. You know, we’ve always looked at security as kind of this data loss prevention. And data loss prevention means like, okay, I’m gonna you know look at my people’s emails and make sure they don’t send any PDFs that shouldn’t go to domains or shouldn’t go to.ru, you know, things like that or.cn. That’s great to make sure the data doesn’t leave the network, but what about?
You know, we’ve we’ve we’ve kind of from browser security and application monitoring, we don’t want a big brother, right? There’s kind of that weird feeling. But you’ve got to catch these things earlier and you’ve gotta have browser-based security, application level monitoring in the browser to ensure what your people are sending where or what your customers people are sending where. that’s a miss if you don’t capture it there. think about you know, you got a customer.
instead of paying for the token forever, they’re paying for steady, heavy, sensitive workloads, run that model maybe in their own, maybe some hardware they want to do, right? I don’t know, maybe they’re even going that route. But you’ve just got to think about how are they gonna run that on their own hardware and how do we get some data sovereignty as a bonus of that? So i you think about for a second, let’s talk about that. the sovereignty point probably deserves a little a little thread on its own. There’s a little compliance landmine.
Hiding in all of this. We’re all running so fast, sometimes we don’t even see it. if you’re paying attention to some of these latest models, some of them are coming out of China, you know, whether it’s Deep Seek or others, they’ve captured because of some cost savings, they’ve captured a real chunk of US enterprise usage. I think it was something north of like 30% token volumes at any point, because they are dramatically, dramatically cheaper.
So here’s the catch, maybe that your customer doesn’t know. A direct API call to one of those providers could route their data through servers in some other jurisdiction, some other country. For a regulated customer, you know, you think healthcare, finance, government, that cheap model, it could kind of blow up their data residency, compliance, all those things. And you being the person that catches it.
before it becomes a headline, that’s just trust that you can’t buy. you know, these these relationships are defined. Not you know, not always until we go through the trenches and we go do some hard things, then they can lean on us for some of the easy things. So the fixed the fix is not just cost cutting. it’s cost cutting, it’s resilience, it’s compliance, and a whole new set of solutions.
In your bag. I think that’s the good stuff, right? It just constantly opens up more. So let’s go to part four. We always try to weave in lately. We’ve been trying to make sure that we weave in questions to these new styles, some of these new formats. So let’s make let’s make something that you can use in your very next customer call. I think you find this opportunity the same way you find every good opportunity, right? I’ve I’ve since the beginning of this podcast, I go back to my what’s my favorite thing.
It’s questions. Ask the right questions. Go get the book that I have no affiliate code on, unfortunately, called Power Questions. If you want a real great weekend refresher, it’s Power Questions by Andrew Sobel. Get it on Amazon. teaches you the power in that, in asking those questions. Aptly titled, I know. All right, write these down. Start broad. You’re walking in, you’re asking who owns your AI spend right now.
And can they tell you this week what you spent the last month of tokens on and how much and what? Nine times out of ten, the answer is a big long pause. Everybody looks around, kind of makes weird eyes, maybe doesn’t make eye contact. nobody owns it, nobody can break it down. You just found the gap, right? Then quantify that surprise a little bit. Give them something, ask something like this Has an AI or cloud bill
Come in higher than you expected in the last two quarters. We’ve got quarters of history of this now. If the answer on that is yes, and it usually is, you’ve got a live wound, and and and kind of a reason to act now, not someday, right? Thinking of, you know, we say live wound, I just instantly go to like, the the bad guy shoots the good guy in the movie, and the bad guy just, you know, while he’s trying to get information out of the good guy, presses on that wound. gosh, you can just you know, you
You give them a reason to act and a reason to fix it. So then what do you do? You just simply find the workloads. and so you you ask, hey, we’re just a little bit exploratory here. what are you using AI for today? And which of those are simple, high volume tasks versus you know, some complex, high-stakes ones? That’s you spotting the right size opportunity out loud, right? And just sizing it up accordingly.
the high volume stuff is where I think a lot of the savings is hiding, honestly. So so then think about lock in. I I talked about a little bit earlier as we get a little bit political, or some of these guys do. Ask if your main AI provider doubled its prices tomorrow or went down for a day, what’s your plan? most of these guys have no answer, especially if they’re going all, you know, all the eggs in one basket.
That’s i it’s the same thing, right? That’s your opening for routing, failover, multi-provider resilience. We did it when we were doing backup and data. Hey, what’s your backup data plan? well, you know, it’s it’s it’s I’m gonna grab this flash drive or we’re gonna push these things up to, you know, this data warehouse. When’s the last time you tested it? well, it’s a great question. I mean the reality is like I I I came from a backup world and we saw this constantly. That was what we had our our
Predicated success on, people never really tested it. You really had to push because things just happen different when you test. All right. the compliance check. Here’s another question for you. Do you know where your AI requests are physically being processed? And has anybody checked that against your data residency requirements? For a regulated customer, that question alone,
That’s that’s getting people called in the room and maybe sometimes getting people called out of the room.
So you notice what all five of these have in common. If you caught the last couple episodes, you’re gonna see the trend, you see where I’m going with this. You’re not selling a product. You are helping them see a risk and a cost that they didn’t know that they had. You lead with the bill, you lead with the exposure, and the solution starts to sell itself because you’re the one who found it first. Here’s the honest part. When you start pulling on that thread.
Some of this gets technical fast. Routing architecture, model selection, design, compliance. That’s not the moment to wing it if you’re not comfortable with it. That is the moment to bring in some reinforcements. Let’s talk about these reinforcements here. part five. Just think about don’t go it alone. Now, if that
In that beginning part, you’ve you’ve drummed up these these conversations and these needs. I I I don’t want you to feel like you you you’ve got to be the deepest AI expert in the room to win this deal. You do not. Your job is to see the opportunity, ask the questions, own the relationship. This deep technical build is what our sales engineering team is here for. This is where we thrive. We love this stuff. So when the conversation gets to something like, you know, okay.
How we actually architect a routing layer? Which models do we tier? Where does the workload go for cost and sovereignty? How do we set up guardrails? You know, should I have a FinOps dashboard? That’s your cue. Loop us in. We’ll go as deep as the customer needs, and obviously our suppliers will continue that thread on, go even deeper. you just stay where your value is massive. Quarterback that deal. You are the trusted face, you are the one.
Getting that conversation to the next level.
So you bring the relationship, you’ve got that instinct, you’ve spotted the opportunity. let the engineering horsepower help you build it, right? I think together that is something the customer can absolutely not get from a model vendor’s website, and they can’t get from a spreadsheet. That is how these deals are gonna be won. Okay, we gotta bring it home here. So remember, a company gave five thousand engineers.
This awesome tool. And four months later, four months, the budget was gone. That’s not a story about AI being bad, right? There’s enough of that, I think, being perpetuated. That is a story about a powerful new technology that has just moved faster than anyone’s ability to govern it. And every single time that happened to the beginning, we talked about in telecom, in cloud, and now AI, the advisor.
Who stepped in to bring order to this chaos didn’t just save the customer money. You deepened the relationship. You earned the next three deals and the next three deals and you became indispensable. This is the third wave of that movie that you already know how to win. The technology’s new, but this this play is not, right? We gotta trust the playbook. Find the runaway spend, bring in the questions.
Bring in your sales engineering team and let’s figure out how to turn the customer’s bill shock into this favorite conversation, this new conversation. And remember when. All right. So let’s look ahead just a tiny bit. Let’s watch for the FinOps world to kind of keep standardizing this. I I see a lot of potential supplier expansion here.
there’s now an effort to build open standards around AI costs. The tooling and the language are only gonna get better and make it easier for you to kind of sell into, right? We’re gonna keep tracking that. We’ll keep bringing things like that as those come along to you here on this podcast or any other Toleris events, the hit calls, all those good things.
So this week, what I want you to do, go pick one customer and ask them the first question. Who owns your AI spend? And can they tell you what they spent last month? Just listen for the pause. It’s hard. Let it pause, let it wait, and watch what happens. that’s the show for today. if this gave you a customer to go call, to go text, to go tee up a conversation with, go do it.
I’m your host, Josh Lupresto SVP of Sales Engineering at Telarus and this has been Next Level Biz Tech wrapping up on the AI Bill Shock. We’ll see you next time.