Works with Claude Code

Keep Claude Code. Turn off the token meter.

Point Claude Code at LunaRoute and it runs on open-weight models such as GLM 5.3, for a fixed monthly price. Keep your Claude plan for the work that needs it, and run the long sessions, parallel agents and background jobs here.

npx @lunaroute/cli setup claude-code
Claude Code
CLAUDE.md · skills · hooks · MCP
LunaRoute
GLM 5.3DeepSeek 4.1 Flash
Claude Code → LunaRoute → open-weight models

Can Claude Code use models other than Claude?

Yes, through a gateway. Claude Code sends its requests to whatever address is set in the ANTHROPIC_BASE_URL environment variable (Anthropic’s gateway docs). LunaRoute is a gateway that answers in Anthropic’s format with open-weight models such as GLM 5.3 and DeepSeek 4.1 Flash, for a fixed monthly price. The LunaRoute CLI sets the variables for you. Anthropic doesn’t support non-Claude models in Claude Code, so support for this setup comes from LunaRoute.

What stays the same, and what changes

What stays the same

You keep the Claude Code you already use. Your CLAUDE.md files, skills, hooks, slash commands and MCP servers live on your machine, and they work the same way when the model calls go to LunaRoute. Claude Code is the most common way people use LunaRoute, and it’s how we use it ourselves.

What changes

Two things change: requests go to models that LunaRoute runs on its own GPUs instead of to Anthropic, and you pay for a number of priority concurrent requests per month instead of for tokens. When a task needs Claude itself, run that session on your Claude plan.

How much does it cost to run Claude Code on LunaRoute?

The price stays the same whether a session runs for ten minutes or all night. Retries, long contexts and abandoned experiments don’t add to the bill, so you can stop doing the math before you try something. A cheap per-token API still charges for every one of those, which means the cost of an agent loop depends on how long the agent decides to run.

Web search is the one thing billed separately. Your organization gets 100 searches a month at no charge. Above that, you can opt in to paid search or bring your own Brave, Exa or Kagi key (see the FAQ).

Concurrency vs token pricing, explained

Price per month and priority concurrent requests for each plan
PlanPrice per monthPriority concurrent requests
Pro$99.002
Max$199.994
Apex$299.006
Teams$499.9910, with per-member limits

Your plan sets how many requests get priority scheduling. Extra requests run on overflow capacity automatically, at background priority, and may wait during busy periods. Priority requests wait up to 60 seconds for a free slot. If the queue is full, the request gets an HTTP 429 with a Retry-After header.

Use both, one session at a time

Claude Code talks to one backend per session. If you keep your Claude plan as the default, the split is easy to control:

Open two terminals and use whichever fits the task. Keep Claude for the steps where you’ve seen it do noticeably better, such as a hard design decision or a gnarly bug you’ve already burned an hour on. The simplest way to find your own split is to give the same task to both and compare.

your normal Claude plan

claude

this session runs on LunaRoute

lunaroute run claude

Work that fits LunaRoute well

  • Long refactors and test-and-fix loops that run for hours
  • Several sessions in parallel, each working on its own branch
  • Subagent-heavy workflows, where one task fans out into many model calls
  • Headless runs in scripts and CI (lunaroute run claude -- -p "...")
  • The "let’s just see if this works" branch you’d otherwise skip because of what it might cost

Where it’s not the right fit

  • You need Claude models for most of your work. LunaRoute runs open-weight models. Try your own tasks on them before you commit.
  • You use Claude Code lightly. If your Claude plan rarely runs out, a second subscription may not be worth it.
  • You need a guaranteed response time. Priority gets you ahead in the queue, but it doesn’t guarantee a particular speed. If you want capacity that isn’t shared with other customers, talk to us about Enterprise.

How many concurrent requests does Claude Code use?

A concurrent request is one model call that is running right now. Requests waiting in a queue don’t count.

A single Claude Code session mostly makes one model call at a time, so it uses one concurrent request while it’s working and none while it waits for you. Subagents running in parallel each use one while they generate, and so does every extra session you have open.

Claude Code’s small auxiliary requests, such as the auto-mode safety check, are routed separately so they don’t queue behind your session’s main work.

The Usage & Metrics page in your dashboard at app.lunaroute.com shows what percentage of the time your requests were queued and how long they waited. If both climb during your busy hours, add priority requests. Not sure where to start? See How much concurrency do I need?

How to set up Claude Code with LunaRoute

Setup is one command:

npx @lunaroute/cli setup claude-code

It signs you in through your browser and prints a few environment variables. Paste them into your shell profile (~/.zshrc or ~/.bashrc), open a new terminal, and run claude as usual. Those lines read your key through the LunaRoute CLI, so install it once with npm install -g @lunaroute/cli (Node.js 20+).

Once the variables are set, they take precedence over your Claude login, so every claude session goes to LunaRoute. Remove them to switch back.

Or keep your Claude plan as the default

If you’d rather keep your Claude plan as the default, skip the profile step and start LunaRoute sessions only when you want them with lunaroute run claude. Pick a model with --model, or switch inside Claude Code with /model, where the picker lists LunaRoute’s models. Anything after -- goes straight to Claude Code.

Full details are in the Claude Code setup docs.

Does LunaRoute keep my code?

No. Claude Code sends a lot of your code and context with every request. On LunaRoute, prompts and responses are processed in memory and then discarded. We keep billing metadata (timestamps, model, token counts) and nothing else from the request, and we never use your data to train models. How zero data retention works

FAQ

Claude Code questions, answered.

Run your next long session on LunaRoute

Setup is one command, and you keep your Claude plan for the hard parts.

npx @lunaroute/cli setup claude-code