How much concurrency do I need?
Count the model calls you have running at the same moment on a busy day, not the number of tools or users. One Claude Code session usually needs one. Every subagent or extra session running in parallel needs one more. Work that can wait should go to the background tier, where it doesn't use your priority slots. Then pick the plan whose number of priority concurrent requests covers that peak.
Quick sizing by workload
| Plan | Concurrent capacity | Price per month | Fits when you run at most |
|---|---|---|---|
| Personal Agent | 1, scheduled as capacity becomes available; no priority scheduling | $19.99 | One personal agent (OpenClaw, Hermes) handling your tasks |
| Pro | 2 priority | $99 | One Claude Code session, plus an occasional second session or subagent |
| Max | 4 priority | $199.99 | Up to 4 sessions or subagents working at the same time |
| Apex | 6 priority | $299 | Up to 6 agents at the same time, with batch work on the background tier |
| Teams | 10 priority, plus bundles of 10 | $499.99, plus $499.99 per bundle | Your team's peak across everyone, with per-member limits |
Every coding plan also includes overflow capacity, so a short spike above your priority number can run at lower priority when capacity is available. Queue limits and timeouts still apply.
What counts as a concurrent request
A concurrent request is one model call that's running right now. It starts when the request reaches the model and ends when the response finishes. Requests waiting in a queue don't count, and neither does the time an agent spends running tests, editing files or waiting for you.
That's why a single developer rarely needs as many slots as they expect. An agent spends a lot of its time doing things other than calling the model.
Common setups
Claude Code, one session.
Claude Code mostly makes one model call at a time, so one session uses one concurrent request while it's thinking. Pro gives you a second slot for a quick parallel session or a burst of subagents. See Using Claude Code with LunaRoute.
Claude Code with subagents.
When Claude Code fans a task out to subagents, each one running in parallel holds a slot while it generates. If you regularly have up to four running at once, Max fits.
Several sessions or worktrees.
Running one agent per branch is where concurrency adds up fastest. Count the sessions you typically have working at the same time. People rarely have all of them calling the model at once, so the peak is usually lower than the number of open terminals.
Batch and background jobs.
Document processing, evaluations and overnight runs don't need priority. Add -background to the model name (for example, glm-5.3-background) and those requests fill gaps around your interactive work instead of competing with it. They can wait up to 240 seconds for a slot (concurrency docs).
Teams.
Size Teams for the team's busiest hour, not for the number of people. Admins can cap each member's concurrency from the Members page, so one person's batch job can't take the whole pool. Add capacity in bundles of 10 priority requests as you grow.
How to tell if you need more
Start with the dashboard. The Usage & Metrics page in your dashboard at app.lunaroute.com shows what percentage of the time your requests were queued and how long they waited. If both climb during your busy hours, add priority requests.
You'll also see it in the API responses:
- Overflowed responses. When a request runs on overflow, the response carries
X-LunaRoute-Warning: overflowed-to-background. An occasional one is normal. If most of your busy hours look like that, move up a plan. - 429 errors. A request that can't be queued returns HTTP
429withCONCURRENT_REQUEST_LIMIT_EXCEEDEDandRetry-After: 60. Seeing these regularly can mean your peak is above your plan, or that a per-member cap is being reached. Check both before upgrading. - Waiting on interactive work. If your own sessions feel slow while batch jobs run, move the batch work to the background tier before you upgrade.
Upgrades take effect immediately, prorated, and downgrades start at the next billing cycle (usage and billing). So it's cheap to start small and move up.
Our suggestion
Start with the plan that covers your typical peak, watch the Usage & Metrics page for a week, and adjust. Upgrades apply immediately, so starting small costs little. For teams, size for the busiest hour rather than one request per person, since people rarely all run agents at the same moment.
Wondering why we price by concurrency at all? See Concurrency vs token pricing.
FAQ
How many concurrent requests does Claude Code use?
One per session while it's calling the model, plus one for each subagent running in parallel. Time spent running tools or waiting for you doesn't use a slot.
What happens if I go over my plan's concurrency?
On coding plans, extra requests can run on available overflow capacity automatically, at background priority, and may wait during busy periods. If capacity is full, priority requests wait up to 60 seconds, and a request that can't be queued gets HTTP 429 with a Retry-After header.
Does a team need one concurrent request per person?
Usually not. People rarely run agents at exactly the same moment, so size Teams for the busiest hour. Admins can set per-member limits so nobody takes the whole pool.
Should batch jobs use priority capacity?
No. Use the background tier by adding -background to the model name. Background requests fill gaps around priority work and can wait up to 240 seconds for a slot.
Can I change plans later?
Yes. Upgrades apply immediately with proration, and downgrades take effect at the next billing cycle.
Pick a starting point
Start small, watch for overflow, and move up when you need to.