← All posts

The Model Is Not the Product

Eran Sandler
A photograph bisected by the ocean surface: a small, ordinary lump of ice above the waterline, and beneath it a vast data centre — server racks, pipework and cabling — descending into darkness.

We tend to talk about AI systems as if the model name tells us what we’re using. “We’re using GLM,” “we’re using DeepSeek,” or “this runs Qwen.” But the model is only one part of what determines what you actually experience. The same model can feel faster or slower, behave differently, and be more or less reliable depending on how it’s quantized, served, cached, and scheduled. As open-weight models get better and more providers offer the same ones, I think this distinction is becoming much more important.

The other day, we were researching LunaRoute performance improvements on an 8×B200 cluster serving GLM-5.3-Vision-NVFP4 (yes, we added GLM-5.2 Vision to GLM-5.3!) and reduced interactive latency from 2.7 seconds to 0.4 seconds. For real human use, this kind of shift is a massive improvement, going from a stutter to a real-time interactive feeling.

The surprising part for us was that this didn’t require a change to the weights, model, or hardware. We only changed how batch/background requests competed with interactive requests under heavy load, focusing on responsiveness where humans need it while maximizing tokens per second across the cluster.

Now, this wasn’t some “free lunch”; background tasks slowed down because we prioritized the interactive experience. But here’s the kicker: one single decision about how to allocate resources improved the user experience by nearly seven times without us changing a single parameter of the model itself. We did this on a stack we own from top to bottom.

In the wild, different inference providers are tweaking hundreds of these variables at once. They might all slap the same model name on their API, but under the hood they’re using different checkpoints, quantization tricks, hardware, caching strategies, and scheduling policies. It’s never really “the same model.”

As great open weights models become more common, access to the weights isn’t the bottleneck anymore. More and more of the value is moving into the systems that actually serve them.

The model name is only part of the story.

The serving system is the product.

The Problem with Model Aliases

We’ve been conditioned to think of the “model” as the yardstick. We argue about whether GLM is better than Qwen or if a new DeepSeek checkpoint can beat Opus on a coding test.

That’s fine for picking a “family” of models, but it’s pretty useless when you’re trying to pick a service to run your production app.

Your app doesn’t talk to a bunch of abstract math weights. It talks to a complex system that decides which version of those weights is loaded, how compressed they are, what hardware they run on, and whether your request has to wait in line behind a massive batch. That system decides if you get a snappy response or if you’re stuck waiting because someone else is running a long prompt.

Providers can mess with everything: the chat templates, the sampling settings, how they parse tool calls, you name it. The model name stays the same, but the reality of the service can shift underneath your feet.

Take quantization, for example. Three different providers might all say they’re serving the same weights. One might serve the original high-precision version, while another uses a heavily compressed version to save money. That compression might hurt a complex coding task, make no meaningful difference at all, or sometimes even produce better results. You don’t really know until you evaluate it on your own workload. You’d never know just by looking at the alias.

Then there’s the engine. Whether a provider is using vLLM, SGLang, or their own proprietary secret sauce matters. Even if two people use the same open-source engine, their specific configurations and patches can lead to completely different performance.

And don’t get me started on hardware. A well-managed cluster of older chips will often feel way better than a brand-new GPU cluster that’s being crushed by too much traffic.

As a customer, you feel the whole system — the weights, software, hardware, caching, and scheduling — as one single experience.

Lessons from the SSD Industry

We’ve actually seen this exact movie before in the hardware world.

Back in 2021, Western Digital swapped out the internals of their WD Blue SN550 SSD without changing the model number. At first, it seemed fine. But once the cache filled up, the write speeds tanked from 610 MB/s to 390 MB/s. People were rightfully upset, and WD eventually had to promise they’d change the model number whenever the specs changed that much. (BleepingComputer)

Samsung did something similar with the 970 Evo Plus. They changed the controller and kept the name. The new version was actually better for short bursts, but for big, sustained jobs, the speed dropped from 1.5 GB/s to 800 MB/s. It also ran hotter.

The real-world result was more complicated. In a copy of a roughly 155 GB file, the larger cache offset the lower post-cache speed, and the two versions completed the transfer in approximately the same time, with some tests showing the revision slightly ahead. (TechSpot)

The drive wasn’t strictly “better” or “worse.” It was just different. But buyers had no way of knowing what they were actually getting because the name was the same.

That’s the exact problem we have with inference providers today.

A background update might boost overall speed but ruin latency for some users. A new quantization might save money but break specific coding tasks. A review you read last week might be totally accurate for the version they served then, but completely wrong for what they’re serving now.

A published review can remain perfectly accurate while no longer describing the product being delivered.

The Serving Stack Determines the Experience

We saw this ourselves.

We ran a test on our own stack: at 87 simultaneous requests, we were pumping out 42,700 tokens per second. But when we pushed it to 188 requests, that dropped to 25,100.

We had hit a wall where the overhead of managing all those requests actually slowed everything down. The system had crossed its saturation knee. At that point, adding more requests reduced rather than increased useful throughput.

Where a provider chooses to sit on that curve is a business decision.

Do you want to maximize profits by cramming as many requests onto a GPU as possible, or do you want to keep things snappy for the users?

Both providers can claim they’re running the same model, but your experience will be night and day.

The model does not make that tradeoff. The provider does.

Caching is another huge variable.

If you’re building an agent, you’re sending the same prompts and files over and over. If the provider is smart and caches that state, they don’t have to re-process it every single time.

In our tests, keeping a prefix in GPU memory cut the time to first token from 206ms to 68ms. Even if it was just in system RAM, it was still way faster than a cold start: 0.96s vs. 2.54s.

For a single request, this is a latency optimization. Across dozens of calls in an agent trajectory, it changes the speed and cost profile of the application.

Yet “supports prompt caching” is still not a complete product specification. Providers differ in how much they retain, where it’s stored, how eviction works, whether it survives between sessions, and whether customers receive any of the economic benefit.

A provider might look slow in a simple benchmark but be a superstar for your specific agent’s workflow.

Why This Matters for Agents

If a chatbot gives a slightly off answer, you just move on. But for an AI agent, small errors compound. If an agent has to make 30 decisions in a row, the math gets scary fast.

Let’s say each step is 95% reliable. The chance of finishing all 30 steps correctly is only 21%. But if you bump that reliability up just two points to 97%, the chance of success jumps to 40%.

That tiny 2% improvement nearly doubles your success rate.

In the real world, agents aren’t just rolling dice; errors cascade. That’s why you can’t just look at one response. You have to look at the whole task.

A slightly better quantization or a more reliable tool-call parser can be the difference between an agent that works and one that’s a waste of money.

And it’s not just about performance — it’s about the bill.

A “cheap” token is expensive if the agent fails and has to retry the whole task. For agents, we should be looking at the cost per successful outcome, not just the price per million tokens.

This is where I think the market starts to split.

Selling generic tokens is going to be a race to the bottom. But serving agents is different. If one provider is slightly more expensive per token but much more likely to actually finish the job, they’re the obvious choice.

That doesn’t guarantee attractive margins for every inference provider. Price competition will remain severe, and plenty of providers will struggle to demonstrate meaningful differentiation.

But the buying criterion changes.

Customers stop asking only how much a million tokens cost and start asking how much it costs to get a piece of work completed.

One of these is a commodity market.

The other is a place for real innovation and differentiated systems.

Abstraction Is Not the Problem

Look, I get it. Most developers shouldn’t have to worry about GPU topology or batching policies. That’s why we have providers in the first place — to hide that complexity.

But there’s a difference between hiding complexity and hiding changes.

When you use a database, you expect it to have a version number and for things not to break without warning. You don’t need to know exactly which CPU is running your code, but you have some idea what you’re buying.

Inference needs that same maturity.

We need model aliases that actually mean something stable. Abstractions are great for hiding mess, but they shouldn’t be used to swap out the product without telling anyone.

An abstraction is useful when it hides complexity. It becomes misleading when it hides substitution.

There’s also a business case for this.

If the serving layer is where the real value is, providers should want to show it off. Without clear provenance, it’s impossible to tell a high-quality service from a cheap, overstuffed one, and everything just becomes about price.

Giving a service an identity doesn’t mean giving away trade secrets. You can identify the checkpoint, precision, and behavioral limits without publishing your proprietary code.

Making quality measurable is how you actually get paid for it.

The Silent Failures

When normal infrastructure breaks, it’s obvious; you get a 500 error or a timeout.

But quality degradation is much more insidious.

Your dashboard stays green, but everything is slightly worse.

Maybe the tool selection is a bit flakier, or the code it generates needs an extra fix. The endpoint is still returning a 200 OK, so you might not even realize your success rate is dropping.

This can happen after an engine upgrade, a quantization change, a capacity move, a new scheduling configuration, or a fallback to another deployment. The change may improve aggregate performance while hurting one customer’s workload.

An unchanged model identifier doesn’t prove that the served system remains unchanged.

Until we have better standards, you have to benchmark the endpoint yourself.

For agents, focus on three things: latency under load, total task success rate, and how well it handles repeated prefixes.

And you probably need to repeat at least some of that evaluation after deployment. Pin revisions where possible, run a few representative canaries, and maintain an alternative endpoint that has been tested recently.

Yesterday’s benchmark has an expiration date.

We Already Know How to Do This

Versioned identities aren’t a new idea. The big players are already doing it to some extent.

OpenAI lets you pin specific snapshots like gpt-4.1-2025-04-14 so your app doesn’t suddenly change its behavior. Anthropic does the same with dated model names. (OpenAI) (Anthropic)

But even these pinned versions don’t tell you anything about the hardware or scheduler behind them.

You can pin the model, but you’re still guessing about parts of the system that actually run it.

For open-weight models through third-party providers, you may not even be able to pin the exact model configuration being served.

The industry already accepts that model behavior needs versioning.

Extending stable identity to the served configuration is overdue.

We’re seeing early signs of progress, though. Fireworks AI, for example, is starting to distinguish the model from the “deployment,” letting you choose things like accelerator type and precision. It’s not full provenance yet, but it’s a step in the right direction. (Fireworks AI)

The Road Ahead

In the next year or two, I think the market is going to stop treating models as interchangeable widgets.

We’re going to start asking for change logs, immutable revisions, and real consistency commitments. If you’re a big buyer, you’ll want to know exactly what checkpoint and precision you’re getting.

I think the way we pay for all of this has to change too.

Per-token billing works reasonably well when you’re thinking about individual prompts. It gets much weirder with agents.

Agents try things. They take wrong turns. They retry. They explore multiple approaches. Sometimes the expensive run is the one that fails, and sometimes that failure is exactly what teaches you what to try next.

If every experiment and every failed attempt increments a meter, you naturally start optimizing for using fewer tokens. That’s not necessarily the same thing as optimizing for getting better work done.

I think that’s an important distinction.

When compute is available as capacity instead of a token budget, the question changes. You already have the capacity, so you’re free to use it. Let the agent try again. Run three approaches instead of one. Evaluate a different quantization. Test another model. Give the agent more context. Let it keep working.

Failure still consumes compute, of course. Capacity isn’t infinite. But it doesn’t show up as another surprise line item on a token bill.

That’s a big part of what we’re trying to do with LunaRoute. We want people to be able to experiment with models and serving configurations, run agents that do a lot of work, and accept that some of those attempts won’t work without having to think about the cost of every additional token.

And it ties back to the rest of this post.

If you can’t know ahead of time whether a different quantization, scheduler, model, or serving configuration will be better for your workload, then you need to be able to try them.

You need room to experiment.

I expect we’ll see more explicit service classes around that idea: reserved capacity for those who need it, low-latency service for interactive apps, throughput-optimized capacity for work that can wait, and fixed-cost capacity where teams can use what they’ve bought without metering every token.

Some workloads may eventually be priced around tasks or outcomes too. There probably won’t be one model that fits everything.

Look, I could be wrong. Maybe everyone just keeps fighting over price per token.

But as agents take over more work, I think the thing that matters is the cost of getting that work finished — and whether developers or knowledge workers feel free to let the system do enough work to actually get there.

That creates room for a different kind of inference company.

The Future Belongs to Systems Companies

The first phase of AI was all about who had the best model.

The next phase is about who can run those models best.

The real winners will be the companies that master quantization, cache reuse, and scheduling — the ones that can take expensive silicon and turn it into a reliable, predictable outcome for your app.

I have a clear bias here: this is exactly why we’re building LunaRoute.

It’s also why the data in this essay comes from a stack we actually control. We don’t think inference is just a routing problem; we think every decision about how context is reused and how capacity is managed is part of the product.

And we think the pricing model is part of that product too. If the only way to find out what works is to experiment, the infrastructure should make experimentation easier, not make you think twice about every retry.

Don’t assume two endpoints are the same just because they have the same name.

A benchmark for the weights isn’t a benchmark for the service.

And never assume that the endpoint you called yesterday is the one you’re calling today.

The model is just a component. The endpoint is the product.

I’d like to thank Jesse Robbins and Gur Brosh for reviewing early drafts of this post.