← All guides

Concurrency vs token pricing: what's the difference?

With token pricing, you pay for every token you send to a model and every token it sends back, so the bill grows with usage. With concurrency pricing, you pay a fixed monthly price for how many requests can run at the same time, rather than for tokens. On LunaRoute, that buys priority capacity with automatic overflow, and the bill for included inference stays the same however much you use. Token pricing suits light, occasional use. Concurrency pricing suits agents and other work that runs a lot and unpredictably.

At a glance

Token pricing compared with concurrency pricing
What changesToken pricingConcurrency pricing
What you pay forEach input and output tokenPriority concurrent requests, with automatic overflow
Monthly billGrows with usage, known only afterwardsFixed, known in advance
A retry or a failed runAdds to the billAdds nothing
Long contextsCost more on every callCost the same
Cached inputCheaper, but only while the provider keeps your cacheFaster responses, same price
Limit you hit when busyYour budget and provider rate limitsPriority capacity, overflow and queue limits
Best forLight or occasional useAgents, batch jobs, heavy daily use

The concurrency column describes LunaRoute's included inference. Model and plan limits still apply; metered add-ons and external provider charges are separate.

How token pricing works

Most LLM APIs charge per million tokens, with separate prices for input and output and often a discount for cached input. For example, Featherless lists GLM 5.3 at $1.40 per million input tokens and $4.40 per million output tokens on its per-token plan (Featherless, checked October 6, 2026).

The appeal is that you pay only for what you use. The catch is that you can't know what you'll use until you've used it.

Cached tokens

Most providers charge much less for input they've seen recently. Featherless lists cached GLM 5.3 input at $0.26 per million tokens instead of $1.40, about 80% less. Agents resend the same instructions and files on every step, so caching can cut an agent's input bill dramatically.

It also makes the bill harder to predict. A cache only lasts as long as the provider keeps it, and that can be just a few minutes without activity. Stop to read a diff or step away for lunch, and the next request pays full price to rebuild the cache. To avoid that, you sometimes have to keep sending requests just to keep the cache warm, which means paying for requests whose only job is to save money on other requests. How long a cache lives, how much it holds and how it's priced all differ from provider to provider, so the same workload can cost very different amounts depending on timing.

How concurrency pricing works

A concurrent request is one model call that's running right now. Requests waiting in a queue don't count. With concurrency pricing, you buy a number of these slots per month. On LunaRoute that's 2 to 10 priority concurrent requests depending on the plan, from $99 to $499.99 a month.

Inside those slots you can send as many requests as you like, with context up to the model's supported limit. What limits you is priority capacity, overflow availability and queue limits, not token counts. Metered add-ons and external provider charges are separate from included inference.

Why agents make token bills hard to predict

A person chatting sends a short message, reads the answer, and sends another. An agent works differently:

  • It resends context on every step. A coding agent sends its instructions, the relevant files and the conversation so far each time it calls the model. A task that takes 30 steps with 40,000 tokens of context each is 1.2 million input tokens before the model writes anything.
  • It retries and explores. Agents take wrong turns, run tests, fail, and try again. The same task can take 10 steps one day and 60 the next.
  • The expensive run is often the one that failed. Sometimes that failure is exactly what tells you what to try next, and you still pay for every token of it.
  • Its cache comes and goes. Cached input is cheap only while it stays cached, so pauses between steps can quietly raise the price of the next one.
30steps×40,000tokens of context=1.2Minput tokens
Input tokens before the model writes anything.

So the cost of an agent depends on decisions the agent makes while it runs, and on timing it doesn't control. Cheaper per-token prices shrink the numbers, but the meter keeps running with every step.

When every attempt adds to the bill, people start optimizing for fewer tokens. They give the model less context, cut retries, and skip the experiment that might not work. That's not the same as optimizing for getting the work done.

I wrote more about this in The Model Is Not the Product.

What concurrency pricing changes

When you already have the capacity, the question becomes how to use it well. We cache shared context too, which makes repeated calls faster, but it doesn't change your bill, so there's nothing to keep warm. You can let the agent try again, run three approaches instead of one, give it the whole codebase instead of a trimmed summary, or leave a batch job running overnight. None of that changes the bill for included inference; model context limits still apply.

That's the idea behind LunaRoute. We'd rather you spend your attention on what you're building than on a token meter.

The limits of concurrency pricing

Capacity is finite, and a fair explanation has to say what happens when you hit it. On LunaRoute (concurrency docs):

  • Priority requests run first, up to your plan's number. If priority and overflow capacity are busy, a new request waits up to 60 seconds for a free priority slot.
  • Overflow is included on every plan. Requests beyond your priority number run automatically on overflow capacity, at background priority, and may wait during busy periods.
  • Background requests, marked with a -background model suffix, are for work that can wait. They fill gaps around priority traffic and can wait up to 240 seconds.
  • When the queue is full, the request returns HTTP 429 with a Retry-After header. Your client should honor it and use bounded retries with backoff.
  • Priority isn't a speed guarantee. It gets you ahead in the queue. If you want capacity that isn't shared with other customers, talk to us about Enterprise.

In short, a busy month doesn't cost more, but very large bursts can wait in a queue. The Usage & Metrics page in your dashboard at app.lunaroute.com shows what percentage of the time your requests were queued and how long they waited. If both climb during your busy hours, add priority requests.

When token pricing is the better deal

  • You use models lightly. If you send a few requests a day, per-token billing will cost less than any monthly plan.
  • You need closed frontier models. Concurrency plans like ours serve open-weight models.
  • You need huge parallelism for a short time. A one-off job that fans out to hundreds of simultaneous requests fits a per-token API better than a plan sized for everyday use.

How to choose

  • If you can predict your usage and it's small, use token pricing.
  • If your usage is heavy, bursty or driven by agents, use concurrency pricing and size it to how many requests you run at once.
  • Many teams use both: a frontier model per token for the hardest steps, and fixed-price capacity for the volume. That's how a lot of people use Claude Code with LunaRoute.

Not sure how many concurrent requests you need? See How much concurrency do I need?

FAQ

What is concurrency pricing for LLM APIs?

You pay a fixed monthly price for a number of requests that can run at the same time, instead of paying per token. On LunaRoute, that buys priority capacity with automatic overflow. The bill for included inference doesn't change with how many tokens you use.

Is concurrency pricing cheaper than token pricing?

For heavy or agent-driven use, usually. For light or occasional use, token pricing is cheaper. The bigger difference is predictability: a concurrency plan costs the same every month.

What happens when all my concurrent requests are busy?

On LunaRoute, extra requests run on overflow capacity automatically at background priority. If capacity is full, priority requests wait up to 60 seconds, and a request that can't be queued returns HTTP 429 with a Retry-After header.

How do cached tokens affect token pricing?

Cached input usually costs far less, often around 80% less. But a cache expires after a period of inactivity that varies by provider, so pauses can send you back to full price, and keeping the cache warm can mean sending extra requests. With concurrency pricing, caching makes requests faster without changing the bill.

Does a long context cost more with concurrency pricing?

Not on the bill. A long context makes a request take longer, so it holds a concurrent slot for longer, but the monthly price stays the same. The model's supported context limit still applies.

Do retries and failed agent runs cost extra?

Not for included inference with concurrency pricing. With token pricing, every token in a failed run is billed.

Stop counting tokens

Pick how many requests you want running at once, and keep building.

Get started