Subscribe to the Next Level BizTech podcast, so you don’t miss an episode!
Amazon Music | Apple Podcasts | Listen on Spotify | Watch on YouTube
The AI race just flipped. For years it was all about building bigger models. Now the center of gravity has moved to serving those models fast, cheap, and everywhere. Josh Lupresto, SVP of Sales Engineering, breaks down the “Inference Flip” in plain English: what training vs. inference really mean, why today’s models “think” before they answer, and how that shift pushes the hard problem—and the money—into real‑time inference.
What you’ll learn
- Training vs. inference, explained simply (write the textbook once, read it billions of times)
- Why letting models “think” increases tokens 10–100x at answer time—and shifts cost to inference
- The memory wall: why moving data, not math, bottlenecks GPUs
- New hardware built for inference: wafer‑scale systems (e.g., Cerebras), LPUs (e.g., Groq), and system‑level designs
- What changes in the real world: costs fall, latency drops, and AI moves to the edge (“physical AI”)
- Where trusted advisors win as customers face more diverse AI deployment choices
Practical takeaways
- Real-time use cases become feasible: voice agents without awkward lag, live translation, multi‑step agents, instant fraud checks
- Edge/onsite AI gets viable for latency, cost, or data‑residency needs
- Reopen “shelved” AI projects—math and performance assumptions may have flipped
Video Transcript
Transcript is auto-generated.
Josh Lupresto (00:00)
I want to start with a number that sounds made up. One company put out an inference chip this year that ran a small AI model about 40 times faster than NVIDIA’s flagship GPU on the same job. And we got another one that claims 25 times faster on the same part of the job. The part that actually matters to your customer. And NVIDIA, the four trillion dollar king of this whole thing, turned around and paid.
20 billion bucks to buy a startup’s technology just to make sure that they weren’t caught flat footed. 20 billion. Here’s the thing. For the last five to 10 years, the AI story has been about one thing. It’s been about building bigger models, bigger, bigger, bigger. And in the last few months, kind of quietly, the whole center of gravity moved. The race stopped being about.
Building the smartest AI, and it just became about running it fast, running it cheap, and running it everywhere. The industry is now coining this term. They’re calling it the inference flip. And it’s the most important shift, I think, in computing that people haven’t heard of yet. So today I want to open the hood. I want to do this a little bit different. it’s a little less sales and tips and things like that, and it’s a little more tech.
I I wanna get feedback on that from you and I wanna see if you like it. So let’s get into it.
Welcome back to Next Level BizTech. I’m your host, Josh Lupresto SVP of Sales Engineering, and today’s title is The AI Inference Flip.
So today we’re gonna go a little deep on the plumbing because the shift I think is gonna reshape what your customers can actually do with AI. And I don’t want you hearing about it secondhand. So let me define the two words here. This whole episode hangs on in plain English training and inference. Training and inference. So just to set the ground here a little bit, training is when you build the model and you take this mountain of data.
An enormous, enormous amount of computing power over weeks or months, and you just teach the AI, right? You’re teaching. That is a giant, one-time, incredibly expensive event. So think of it as think of it as like writing a book, writing a textbook. It takes a huge team, it takes a long time, but you do it once. So now inference, what is inference? That is every time you actually use the model after that.
You ask it a question and it answers. That’s inference. That’s the textbook being read. But the book being read is cheaper than writing it, of course. But here’s the key. You write the textbook once and then you get it read billions of times per day.
So for years, all of the money, all of the attention, all of the giant NVIDIA clusters that everybody had, they were pointed at training. Building the biggest model was the whole game, because building a bigger model was a smarter model inherently, and smarter one. That’s kind of what we were led to believe. But if you look at kind of where the actual work is now, every chatbot reply, every AI agent doing a task, every coding assistant, every
AI voice call, that’s all actually inference. It’s happening constantly at a massive scale and it never stops. And I I think in this you have analysts that are now projecting that that demand for inference computing will outstrip training by a staggering margin. I don’t think we’ve seen this yet. So this is kind of the first time.
So some of the estimates put it, I think, more of a hundred times greater within the next couple years.
So the center, you know, over the last little while, the center of this industry quietly rotated. You know, we spent years obsessed with building the models. Now the race is all of a sudden about serving them billions of times a day as fast and as cheap as possible. That is the inference flip. And I think once you see it, you can’t unsee it. So I’m gonna hopefully this lays a little bit of the foundation for you. All right.
Let’s jump into part two. Why now? and the TLDR is because the models started to think. So, what does that mean? Why did this flip happen right now? Why not two years ago? there’s there’s a specific technical reason, and it’s kind of interesting. The models started to think. And I’m not talking about AGI, I’m not talking about robots, I’m not talking about all of that stuff. Let me explain. So the older AI models worked a lot like
you got familiar with autocomplete, right? You ask and they immediately start typing an answer, word by word, no real deliberating, hey, what what did you want to say here? it was fast, but they would fall on their face when anything genuinely hard would come up. Think like multi step math when I’m trying to help my son with his with his homework or complex logic. Kind of breaks down. Then the industry figured out something that that changed everything.
if you let the model think just a little bit before it answers, generate a whole internal train of thought, try a few approaches, check its own work, and then give you the final answer, it gets dramatically smarter. And the stats on this were kind of nuts. So there’s a there’s a famous example on a hard math test. The old style model, before we got into the thinking, scored around 12%. I know everybody that graduates with C, they call them graduates, but 12% is not good. the same underlying model.
Given the ability to think, now jumped to 74%. Nothing about the training changed, right? You’ve still got that inherently same set of training data there. All that changed was just letting it spend a little more time and and and effort at the moment you ask that question, just to think. So here’s why it matters for chips, and here’s here’s kind of the whole ballgame, I guess. When a model thinks.
It is generating tokens, a lot of tokens. We’re hearing a lot about tokens, token maxing, token this, token that. P CFOs are very concerned. So all that internal reasoning, the scratch work that it does before answering, can be ten to a hundred times more computing per question than the old style, right? So there’s a little trade-offs you gotta understand with some of this. So the models got smarter.
By doing way more work at the exact moment your customer is sitting there waiting for an answer. And that pushed the hard problem and the money out of the training data center and straight into inference, where speed is everything. Okay, so so think about this from the user’s seat. If AI’s gotta think for two minutes before it answers, ain’t got time for that, right? That’s like
Going to Chick-fil-A and then going somewhere else and go, no, no, no Chick-fil-A set the bar. That’s not gonna work for me. And I love Chick-fil-A. if you think about it just as hard and it can hand you the answer in three seconds, it’s kind of magic. So suddenly this most valuable thing in the world is not a chip that can train the model over three months. It is a chip that can think fast and in real time. That is a fundamentally different kind of machine. All right. So let that
Let that sink in. Let’s let’s jump into the third part of this, what we’re gonna call the memory wall and how we now have options here. So let’s talk about why a normal GPU, the graphics style chip card that NVIDIA got famous for isn’t actually ideal for this. And I’m gonna try to my propeller hat is spinning, but I’m gonna try to keep this as human as I can.
So the bottleneck for the inference, now that you understand that, isn’t really raw math horsepower. It’s memory. Specifically, the constant shuffling of the model’s data back and forth between the chip’s memory and to the part that does the computing. Engineers have a technical term for this. It’s called the memory wall. The chip spends a shocking amount of its time just waiting, waiting for data to move. think up think about it like this.
since we’re talking food and we’re talking Chick-fil-A. So think if you’ve got a world-class chef who has to run a warehouse over to that warehouse across the street to grab every single ingredient one at a time. Cooking isn’t the slow part. It’s the running back and forth across the street that is. So a wave of companies looked at that and said, Well, what if we build a chip specifically to kill this memory wall? And so the approaches on this are are are crazy.
You may have heard of these guys. if you haven’t, you will. Company called Cerebrus. So you think about most of the chips and the race, it’s been about get the chips smaller, smaller, smaller. And now these guys are talking about five nanometers, four nanometers, three nanometers, like a crazy, crazy. It makes a hair look big. So most chips then these days are about the size of a postage stamp. So when they they print these out on the press.
You cut a bunch of them out of this giant silicon wafer. Cerebra said, nah. We’re gonna make the entire wafer one single giant chip. Huge. It’s the size of a dinner plate. So why? So that the model’s data can live right there on the chip next to the computing and nobody’s running across the street to do anything. So those guys just went public this year. Great timing. super, super smart CEO. I’d say
Go check out, read a little bit on Cerebrus. They are running a brilliant operation. So just like all the partnerships that we see, all the money exchanging hands, those guys went public this year, like I said, and they’ve got a partnership with OpenAI, $20 billion, right? They signed a deal with AWS, where they’re claiming to answer, you know, doing kind of the the answer generating part of the inference.
Doing that 25 times faster than a standard GPU. And the standard GPU was already fast, right? It blows my mind what these GPUs can do. And these guys just surpassed it. And then there’s Grok. So by the way, watch these names because this is this is where it gets a little bit spicy. Grok, G-R-O-Q, built something called an LPU. We haven’t talked about LPUs yet, using a different memory approach that just makes
The speed extremely consistent and predictable. So no waiting, no jitter. You know, we don’t like jitter on voice calls, things like that. Just a steady fire hose of tokens. So it was so good at inference that at the end of last year, NVIDIA said, you know what? Let’s pay $20 billion to license that technology and bring the core groc team in-house. That is insane. Let that sink in for a second. This
Four trillion dollar behemoth, the most dominant chip company on the planet, looked at this little inference startup and decided it was cheaper to buy the threat than to fight it. lots of good stuff. go back and, you know, we talked about it on one of the previous episodes, The NVIDIA way. There’s a lot of little stories like that. fascinating book. I encourage you to read it if you didn’t go get it yet.
And then the the you know, the story goes on. We could probably shout out a few more names. There’s another startup that came in out of nowhere saying, hey, we could run a small model 40, 50 times faster than Nvidia’s flag chip. So it’s just kind of this, you know, size and inference trade-off. But here’s the pattern out of all of it. The unit of competition used to be the chip, right? Now it’s this entire rack mounted system. It’s the whole integrated system of compute and memory and networking engineered top to bottom to do one thing.
Serve the answers instantly. Nvidia saw it, right? the newest, their newest platform is designed this way. The whole industry is kind of retooling around inference at the same time. So inference has really become the name of the game. All right. Let’s go part four. What actually changes now that you’ve kind of got this foundation built up? So let’s zoom back out because it’s easy to get lost in all of these chip names.
What does the inference flip actually change in the real the real world? changes a couple things. One, the AI gets faster and cheaper to run. And when you have chips that are 10 to 20 to 40 times more efficient at inference, the cost of running AI at scale falls off a cliff. And the speed goes through the roof. Two great things for that to happen. Things that were just too slow or too expensive to be practical, suddenly a year ago become viable.
We talked about this in in the early building of the AI practice here, a a thing called Jevons paradox. In the beginning, you know, well, my gosh, electricity, you know, back we’re going back to the 1800s, right? You had Jevons, who was this 1800s economist, said when the cost of something gets cheap enough, that’s when it becomes really viable, right? This thing that’s great in this beautiful idea, it’s just so expensive. Everybody runs to it on the top of this S curve, but then it
Once the Jevons Paradox rule says it gets down here to a cost perspective, things like electricity and and and and other things and cars, then they become viable. So what’s the second thing here? r real-time AI becomes real. this is a big one, right? when a model can think and kind of deplo think deeply and deploy these answers still in a heartbeat, you unlock stuff that only works if it’s instant.
All this natural voice AI gets better. It starts responding like a human. None of the awkward lags. we’ve seen this a lot in the CX side really improving in the last year or two. And especially now, I think it’s even going to get better. So agents that complete a multi-step task while a customer waits, live translation, real-time fraud detection, the list goes on, right? So the lag was the thing that killed some of these use cases. But the inference flip is really what.
kills the lag if you think about it. third thing. So AI starts leaving the cloud. So when inference gets this efficient, you don’t always need a gigantic data center. You can start running some serious AI and the customers could start running some serious AI on just some smaller hardware in a factory, managed with some help, right? It becomes a little more deployable.
they could run it in a vehicle, they could run it on a device, right? We saw this with what mobility did for 5G and putting some of these you know, fleet management and fleet tracking and just as these technologies became available. So the the industry, of course, why not introduce some new names? the industry’s got a name for this. They’re calling it physical AI, edge inference, and that’s a big part of why capital is pouring, just pouring into this space right now.
It’s good time for those guys to be in this space. Okay, so put those three together, and here’s the headline. The constant strain on what AI can do is shifting from how smart is the model to how fast and how cheap can we run it? And how close to the customer can we put it? The intelligence is almost a solved problem, right? We’re seeing a lot of a lot of gains in that.
So the delivery of that is is the new frontier, I think. It’s a completely different set of winners and losers, and and I mean it’s just one that we have been watching for the last few years. Who’d have thought somebody comes in to do something different than NVIDIA had proven as the leader, right? All right. part five. as we get kind of towards the last couple thoughts here. So what does this mean for us? Why do we care? And again, I would love
some feedbacks and and comments on this. you know, I know we went a little more technical into this episode, but we don’t often we don’t often do that. so we wanted to experiment and go, is this is this some of the depth that you’d like, or did we go too deep? Did Lepresto get a little too nerdy for you? And that doesn’t help you. So anyway, give us some comments, shoot me a note and just kind of let me know your your your thoughts on that, right? Do you want to see more like this or mm
Bring it back up a level. Okay, so so final five, you know, part five here. What does this mean for us? So let’s let’s do the practical translation here. So every one of those three changes turns into something that your customer can now buy and deploy that maybe they couldn’t a year ago, or maybe they don’t think they maybe they tried it. And we talked about it, right? It’s better real-time voice AI in the context center, AI agents that
Don’t make the customer wait. AI running at the edge for the customers who care about latency or or or cost or data staying in the building. The use cases for maybe some of those things was cool demo, but I don’t know that it’s ready for us. Now those guys are crossing into what’s ready now, right? So I think that’s enabling our suppliers to just keep bringing quick iterations, leveraging that technology, and allows you to go to market with options that.
Maybe couldn’t solve a use case months ago, a year ago. And there’s a strategic piece to this, I think. as the inference world gets a little more diverse, GPUs for some things, these big giant wafer scale chips like the Cerebro stuff we talked about for others, edge devices for a different use case, your customers’ uses and you know, they’re gonna have this more diverse set of choices, right? The customers’ uses are just getting more complicated.
They’re not getting simpler. So, which I think is always exactly where the trusted advisor crushes this. You don’t need to know all all of the things that we’re talking about about how to build these chips and you know, why inference and why the memory wall and all those things, but you just you need to understand that the ground moved. And now I think there’s maybe some new options worth putting on the table.
So let me leave you with what I always I I I have to leave you with, which is just a couple questions to take into your very next customer conversation because I think this shift opens the door a little bit and can help you walk through, give you some talking points. All right. First question Where in your business would AI be valuable if it responded instantly instead of with a lag?
So that one kind of surfaces the real-time use cases, you know, the voice agent, the CXI, the live assists, things that maybe they didn’t think were possible before. Second, and and this is my favorite. Did you look at use cases in AI in the last year and shelve it? Because maybe it was too slow or too expensive. you know, back to our Jevons paradox thing, right? When things get cheaper, they become much more scalable. Because here’s the thing on that. The
The ground moved underneath that no. They may have had a founded no before, but the math that killed that project may not be true anymore. is that a dead deal that you get to reopen? Is that a use case that they thought they just couldn’t accomplish that maybe now we can?
And third, what what is holding you back from doing more with AI right now? Is it speed or is it cost? I think their answer tells you you know which door to walk through because this inference flip here that we’re starting to see is hammering on both. So here is your one concrete move, right? To do something different, something new, a new conversation. Pick a customer, maybe.
Who shelved an idea for being a little bit you know, an AI idea for for being too slow or too pricey, and just reopen that. Like I said, the ground moved underneath the original no, and you just have to be the one that tells them. So hopefully the technical foundation for some of this gives you a little bit of color into that conversation if you want to go into it. And of course, shameless plug here. When a customer asks and gets excited and says, my gosh, could we run this real-time voice agent or could we get better at inference?
That’s kind of the moment the technical design matters. And that’s exactly what our sales engineering team is here to help you with. You spot this, find the opening. You’ve obviously manning that relationship. We’ll go deep on the architecture. And obviously our suppliers will take it when it gets to a certain point. So you bring the trust, we’ll just bring that horsepower. That’s how this frontier stuff actually gets deployed. And again, look at it, right? Go back to the last couple episodes. You’re seeing it. You got validation in it.
That’s why these guys are putting the forward-deployed engineers out into the customer’s environment because it takes these kind of architecture discussions. Okay. So final thoughts. Let’s put a bow on this. For five plus years, the whole game was about building a bigger brain. And almost overnight, this game just became how do we let that brain just think fast, run cheap, and live everywhere? From you know, a a dinner s a dinner plate size to chip.
To a device in your customer’s warehouse. That is the inference flip, and that is how intelligence got built. Now, the entire industry is racing to deliver it. That’s the exciting part for us because delivery, deployment, getting technology into the hands of the businesses and the management of that, the deployment of that, the scoping of that, that is the part where humans and advisors shine. that is
The smartest model in the world does nothing sitting in a lab, right? It only does it right when it’s running in your customer’s business, solving a real problem in real time. So, you know, Ms. Cleo Crystal Ball here. As you look ahead, just keep an eye on the edge, keep an eye on physical AI and how fast these inference chips show up inside of products that your customers are already using. This is moving quicker than almost anything that we’ve tracked that we’ve talked about so far. So
the capital’s gotta get deployed and those guys have to produce results. And they will. And we’ll keep watching it from here and we’ll keep bringing you new and and shiny updates as they come out. So that’s gonna wrap us up for today. if you learned something new, again, go share it with somebody that wants to geek out on it. But more importantly, give me some feedback. Did you like this level of depth? Did we go too deep? was it just right? Or do you wanna see more or more of those like this as well in the future?
I’m your host for today, Josh Lupresto SVP of Sales Engineering, and this has been Next Level Biz Tech, the AI and Inference Flip. And we will see you on the next one. Thanks, everybody.
Level Biz Tech.