← All comparisons

LunaRoute vs OpenRouter

OpenRouter gives you one API key for hundreds of models, including Claude, GPT and Gemini, routed to dozens of providers and billed per token at each provider's price. LunaRoute serves a curated set of open models on infrastructure we run ourselves, for a fixed monthly price for included inference.

Use OpenRouter when you want breadth, closed frontier models, or occasional use. Use LunaRoute when you run agents on open models all day and want a predictable bill and consistent serving. Plenty of people use both.

At a glance

OpenRouter details are from its pricing page, docs and public model API, checked October 9, 2026. The source list links the references for pricing, routing, tools and privacy.

LunaRoute and OpenRouter at a glance
What changesLunaRouteOpenRouter
How you payBase coding plans: $99 to $499.99 per month, sized by priority concurrent requestsPer token at each provider's list price, plus a 5.5% fee when you buy Standard plan credits ($0.80 minimum per purchase)
ModelsA curated set: GLM 5.3, GLM 5.3 Flash, DeepSeek 4.1 Flash, plus decision, embedding and image modelsHundreds of models, including closed frontier models
Who serves the modelLunaRoute, on our own GPUsOne of dozens of providers, subject to your routing settings
GLM 5.3One serving setup we run and tune: NVFP4 with vision, 512K context41 endpoints from 32 providers, at FP4, NVFP4, FP8 or unstated precision
Claude Code with open modelsSupported by LunaRoute; one command generates the settings to applyPossible, but only guaranteed with the Anthropic first-party provider
Prompt retentionZero data retention by default for LunaRoute-hosted inferenceNot stored by OpenRouter unless you opt in. Provider policies vary, with a ZDR routing setting
Tools around the modelsWeb search, OCR, document conversion, image generation, embeddingsModel access plus web search and PDF processing
Free optionNo free hosted-inference planFree models, rate limited
EnterpriseDedicated infrastructureSSO, SLAs, policy controls

The LunaRoute column describes hosted inference and base coding plans; other plans differ. Metered add-ons and external providers are separate. Model and capacity limits still apply. Provider counts, prices and features can change.

How you pay

OpenRouter passes through each provider's list price with no markup on requests. It charges a 5.5% platform fee when you buy credits on its standard pay-as-you-go plan, with a $0.80 minimum per purchase, and offers free models capped at 50 requests a day (1,000 after you've bought $10 of credits). That's a good deal for light use and for trying many models. See OpenRouter pricing and its fee and free-tier explanation.

For agents, per-token pricing has the usual problem: the bill grows with every step, retry and long context, and you only know the total afterwards. On OpenRouter it also depends on which provider handles each request. For GLM 5.3, non-cached list prices across providers ranged from $0.05 to $2.80 per million input tokens and $2.20 to $8.80 per million output tokens in the endpoint API snapshot checked October 9, 2026.

LunaRoute's base coding plans are a flat monthly price for 2 to 10 priority concurrent requests, with available overflow included. Long sessions and retries don't add to the included-inference bill, within model and plan limits. Overflow depends on available capacity; queue limits and timeouts still apply, and clients should use bounded retries. Metered add-ons and external providers are separate. More in Concurrency vs token pricing.

Same model name, different service

This is the less obvious difference, and for agents it can matter more than price.

When you call GLM 5.3 on OpenRouter, the model name can resolve to endpoints from 32 providers. On October 9, 2026, those providers listed 41 endpoints: 14 at FP8, 12 at FP4 or NVFP4, and 15 that didn't state a precision. Context limits ranged from about 262,000 tokens to 1 million. Each provider also runs its own engine, parsers, caching and scheduling. The public endpoint API reports the listed precisions, contexts and prices; these are configuration differences, not a quality ranking.

You can set provider preferences on OpenRouter, restrict requests to a provider allowlist, disable fallbacks and filter by precision (provider routing docs). Those controls can narrow the variation, but the model name alone doesn't tell you what you're getting.

That variation shows up in agent work. A different quantization or a slightly different tool-call parser can change how often a 30-step task finishes. I wrote more about this in The Model Is Not the Product.

LunaRoute runs one serving setup per model on our own GPUs, and we tune and test it ourselves. Our GLM 5.3 is a vision-capable NVFP4 build with a 512K context window, and the weights are public. You get the same system on every request, and when we change it, we test it first.

Claude Code

OpenRouter exposes an Anthropic-compatible endpoint, so you can point Claude Code at it. Its own guide says Claude Code there “is only guaranteed to work with the Anthropic first-party provider”, and other models may have compatibility limitations (OpenRouter docs). It's a good way to pay per token for Claude models themselves.

LunaRoute is built for running Claude Code on open models. Start setup with one command, and we support it:

npx @lunaroute/cli setup claude-code

The command prints the environment settings to apply; it does not write Claude Code's configuration or your shell profile. Follow the printed instructions, then restart Claude Code. See the setup steps and Using Claude Code with LunaRoute.

Privacy

OpenRouter says it doesn't store your prompts or responses unless you opt in. It retains request metadata, and offers separate opt-in logging and data-use settings. The providers it routes to have their own policies, and OpenRouter offers a Zero Data Retention routing setting to restrict model requests to endpoints with a ZDR policy. In-memory caching can still occur, and ZDR endpoint filtering does not universally cover plugin backends.

For LunaRoute-hosted inference there's only one provider: us. Prompts and responses are processed in memory and discarded by default. We retain account, usage and billing metadata. Generated images and URL-based document artifacts can be retained for up to seven days; third-party traffic, including BYOK, follows the upstream provider's policy. See zero data retention.

Choose OpenRouter if

  • You want one API key for many models, including Claude, GPT and Gemini.
  • Your usage is light or occasional, or you want free models to experiment with.
  • You want to compare many models and providers quickly.
  • You need Enterprise SSO and contractual SLAs across many providers.

Choose LunaRoute if

  • You run agents on open models for hours a day and want a fixed monthly bill for included inference.
  • You want the same serving setup on every request, tuned for agent work.
  • You want to run Claude Code on open models with a setup that's supported.
  • You want web search, OCR, document conversion and images from the same key.

Using both

A common setup is OpenRouter (or Anthropic directly) for occasional calls to closed frontier models, and LunaRoute for the high-volume agent work on open models. With the LunaRoute CLI installed and signed in, Claude Code makes this easy: keep one session on your Claude plan and run lunaroute run claude for the long, heavy sessions.

Switching from OpenRouter

Both offer OpenAI-compatible endpoints, so most clients need a new base URL, key and model ID: use https://gw.lunaroute.com/v1 with your lr_ key and a model from the LunaRoute catalog.

LunaRoute also accepts Anthropic Messages at /v1/messages and Responses requests at /v1/responses on the same gateway. For Anthropic clients that append /v1/messages themselves, use the gateway root, https://gw.lunaroute.com, rather than the OpenAI /v1 base URL. For coding agents, the LunaRoute CLI generates the appropriate configuration; the Claude Code setup prints environment settings for you to apply.

FAQ

Is LunaRoute cheaper than OpenRouter?

It can be for heavy, agent-driven use of open models, but the answer depends on the model, token volume and caching. A fixed price makes the included-inference bill predictable; it does not guarantee the lowest cost. For light or occasional use, OpenRouter's per-token pricing or free models usually cost less. Compare a month of your actual workload using Concurrency vs token pricing.

Does OpenRouter mark up model prices?

No. OpenRouter bills inference at each provider's list price and charges a 5.5% platform fee when you buy credits on its standard pay-as-you-go plan, with a $0.80 minimum per purchase. See OpenRouter pricing and its fee explanation.

Is GLM 5.3 the same on every provider?

No. On OpenRouter, GLM 5.3 was listed by 32 providers at different precisions (FP4, NVFP4, FP8 or unstated), context limits and prices on October 9, 2026. Each provider's engine and settings also differ. You can use provider and precision filters to narrow the options; check the endpoint API for the current list.

Can I use Claude Code with OpenRouter?

Yes, through its Anthropic-compatible endpoint. OpenRouter’s guide only guarantees compatibility with the Anthropic first-party provider. LunaRoute supports Claude Code on open models; one setup command prints the settings to apply.

Does LunaRoute offer Claude or GPT models?

Not in our hosted-inference catalog: LunaRoute serves open-weight models. For closed frontier models, use their providers or a router like OpenRouter alongside LunaRoute. Separately connected third-party providers, including BYOK, are outside the included hosted-inference plan and follow their own pricing and retention terms.

Sources

OpenRouter sources checked October 9, 2026. Counts and prices are a dated snapshot, not a promise about future availability.

Run your agents on one stack, for one price

Start setup with one command. The bill for included inference stays the same however long your agents run.

See plans and pricing