LunaRoute vs OpenRouter
OpenRouter gives you one API key for hundreds of models, including Claude, GPT and Gemini, routed to dozens of providers and billed per token at each provider's price. LunaRoute serves a curated set of open models on infrastructure we run ourselves, for a fixed monthly price for included inference.
Use OpenRouter when you want breadth, closed frontier models, or occasional use. Use LunaRoute when you run agents on open models all day and want a predictable bill and consistent serving. Plenty of people use both.
At a glance
OpenRouter details are from its pricing page, docs and public model API, checked October 9, 2026. The source list links the references for pricing, routing, tools and privacy.
| What changes | LunaRoute | OpenRouter |
|---|---|---|
| How you pay | Base coding plans: $99 to $499.99 per month, sized by priority concurrent requests | Per token at each provider's list price, plus a 5.5% fee when you buy Standard plan credits ($0.80 minimum per purchase) |
| Models | A curated set: GLM 5.3, GLM 5.3 Flash, DeepSeek 4.1 Flash, plus decision, embedding and image models | Hundreds of models, including closed frontier models |
| Who serves the model | LunaRoute, on our own GPUs | One of dozens of providers, subject to your routing settings |
| GLM 5.3 | One serving setup we run and tune: NVFP4 with vision, 512K context | 41 endpoints from 32 providers, at FP4, NVFP4, FP8 or unstated precision |
| Claude Code with open models | Supported by LunaRoute; one command generates the settings to apply | Possible, but only guaranteed with the Anthropic first-party provider |
| Prompt retention | Zero data retention by default for LunaRoute-hosted inference | Not stored by OpenRouter unless you opt in. Provider policies vary, with a ZDR routing setting |
| Tools around the models | Web search, OCR, document conversion, image generation, embeddings | Model access plus web search and PDF processing |
| Free option | No free hosted-inference plan | Free models, rate limited |
| Enterprise | Dedicated infrastructure | SSO, SLAs, policy controls |
The LunaRoute column describes hosted inference and base coding plans; other plans differ. Metered add-ons and external providers are separate. Model and capacity limits still apply. Provider counts, prices and features can change.
How you pay
OpenRouter passes through each provider's list price with no markup on requests. It charges a 5.5% platform fee when you buy credits on its standard pay-as-you-go plan, with a $0.80 minimum per purchase, and offers free models capped at 50 requests a day (1,000 after you've bought $10 of credits). That's a good deal for light use and for trying many models. See OpenRouter pricing and its fee and free-tier explanation.
For agents, per-token pricing has the usual problem: the bill grows with every step, retry and long context, and you only know the total afterwards. On OpenRouter it also depends on which provider handles each request. For GLM 5.3, non-cached list prices across providers ranged from $0.05 to $2.80 per million input tokens and $2.20 to $8.80 per million output tokens in the endpoint API snapshot checked October 9, 2026.
LunaRoute's base coding plans are a flat monthly price for 2 to 10 priority concurrent requests, with available overflow included. Long sessions and retries don't add to the included-inference bill, within model and plan limits. Overflow depends on available capacity; queue limits and timeouts still apply, and clients should use bounded retries. Metered add-ons and external providers are separate. More in Concurrency vs token pricing.
Same model name, different service
This is the less obvious difference, and for agents it can matter more than price.
When you call GLM 5.3 on OpenRouter, the model name can resolve to endpoints from 32 providers. On October 9, 2026, those providers listed 41 endpoints: 14 at FP8, 12 at FP4 or NVFP4, and 15 that didn't state a precision. Context limits ranged from about 262,000 tokens to 1 million. Each provider also runs its own engine, parsers, caching and scheduling. The public endpoint API reports the listed precisions, contexts and prices; these are configuration differences, not a quality ranking.
You can set provider preferences on OpenRouter, restrict requests to a provider allowlist, disable fallbacks and filter by precision (provider routing docs). Those controls can narrow the variation, but the model name alone doesn't tell you what you're getting.
That variation shows up in agent work. A different quantization or a slightly different tool-call parser can change how often a 30-step task finishes. I wrote more about this in The Model Is Not the Product.
LunaRoute runs one serving setup per model on our own GPUs, and we tune and test it ourselves. Our GLM 5.3 is a vision-capable NVFP4 build with a 512K context window, and the weights are public. You get the same system on every request, and when we change it, we test it first.
Claude Code
OpenRouter exposes an Anthropic-compatible endpoint, so you can point Claude Code at it. Its own guide says Claude Code there “is only guaranteed to work with the Anthropic first-party provider”, and other models may have compatibility limitations (OpenRouter docs). It's a good way to pay per token for Claude models themselves.
LunaRoute is built for running Claude Code on open models. Start setup with one command, and we support it:
npx @lunaroute/cli setup claude-code
The command prints the environment settings to apply; it does not write Claude Code's configuration or your shell profile. Follow the printed instructions, then restart Claude Code. See the setup steps and Using Claude Code with LunaRoute.
Privacy
OpenRouter says it doesn't store your prompts or responses unless you opt in. It retains request metadata, and offers separate opt-in logging and data-use settings. The providers it routes to have their own policies, and OpenRouter offers a Zero Data Retention routing setting to restrict model requests to endpoints with a ZDR policy. In-memory caching can still occur, and ZDR endpoint filtering does not universally cover plugin backends.
For LunaRoute-hosted inference there's only one provider: us. Prompts and responses are processed in memory and discarded by default. We retain account, usage and billing metadata. Generated images and URL-based document artifacts can be retained for up to seven days; third-party traffic, including BYOK, follows the upstream provider's policy. See zero data retention.
Choose OpenRouter if
- You want one API key for many models, including Claude, GPT and Gemini.
- Your usage is light or occasional, or you want free models to experiment with.
- You want to compare many models and providers quickly.
- You need Enterprise SSO and contractual SLAs across many providers.
Choose LunaRoute if
- You run agents on open models for hours a day and want a fixed monthly bill for included inference.
- You want the same serving setup on every request, tuned for agent work.
- You want to run Claude Code on open models with a setup that's supported.
- You want web search, OCR, document conversion and images from the same key.
Using both
A common setup is OpenRouter (or Anthropic directly) for occasional calls to closed frontier models, and LunaRoute for the high-volume agent work on open models. With the LunaRoute CLI installed and signed in, Claude Code makes this easy: keep one session on your Claude plan and run lunaroute run claude for the long, heavy sessions.
Switching from OpenRouter
Both offer OpenAI-compatible endpoints, so most clients need a new base URL, key and model ID: use https://gw.lunaroute.com/v1 with your lr_ key and a model from the LunaRoute catalog.
LunaRoute also accepts Anthropic Messages at /v1/messages and Responses requests at /v1/responses on the same gateway. For Anthropic clients that append /v1/messages themselves, use the gateway root, https://gw.lunaroute.com, rather than the OpenAI /v1 base URL. For coding agents, the LunaRoute CLI generates the appropriate configuration; the Claude Code setup prints environment settings for you to apply.
FAQ
Is LunaRoute cheaper than OpenRouter?
It can be for heavy, agent-driven use of open models, but the answer depends on the model, token volume and caching. A fixed price makes the included-inference bill predictable; it does not guarantee the lowest cost. For light or occasional use, OpenRouter's per-token pricing or free models usually cost less. Compare a month of your actual workload using Concurrency vs token pricing.
Does OpenRouter mark up model prices?
No. OpenRouter bills inference at each provider's list price and charges a 5.5% platform fee when you buy credits on its standard pay-as-you-go plan, with a $0.80 minimum per purchase. See OpenRouter pricing and its fee explanation.
Is GLM 5.3 the same on every provider?
No. On OpenRouter, GLM 5.3 was listed by 32 providers at different precisions (FP4, NVFP4, FP8 or unstated), context limits and prices on October 9, 2026. Each provider's engine and settings also differ. You can use provider and precision filters to narrow the options; check the endpoint API for the current list.
Can I use Claude Code with OpenRouter?
Yes, through its Anthropic-compatible endpoint. OpenRouter’s guide only guarantees compatibility with the Anthropic first-party provider. LunaRoute supports Claude Code on open models; one setup command prints the settings to apply.
Does LunaRoute offer Claude or GPT models?
Not in our hosted-inference catalog: LunaRoute serves open-weight models. For closed frontier models, use their providers or a router like OpenRouter alongside LunaRoute. Separately connected third-party providers, including BYOK, are outside the included hosted-inference plan and follow their own pricing and retention terms.
Sources
OpenRouter sources checked October 9, 2026. Counts and prices are a dated snapshot, not a promise about future availability.
- Pricing and plan features: billing model, free models and Enterprise controls.
- Credit fees and free-tier limits: the 5.5% fee, $0.80 minimum and daily free-model allowances.
- Model catalog API and provider API: model and provider breadth.
- GLM 5.3 endpoint API: provider/endpoint counts, precision, context and non-cached prices.
- Provider routing: provider allowlists, fallback controls and precision filters.
- Claude Code integration: endpoint support and the first-party compatibility guarantee.
- Web search and plugins: search and PDF-processing capabilities.
- Data collection and ZDR routing: retention controls and their scope.
Run your agents on one stack, for one price
Start setup with one command. The bill for included inference stays the same however long your agents run.