LUNAROUTEENTERPRISE[ BOOK A DEMO ]
// LUNAROUTE ENTERPRISE

Your whole AI stack,
under one gateway.

Two ways to scale with LunaRoute: a unified gateway over every model you use, and private inference that squeezes more throughput out of dedicated servers you control.

one gateway: BYO keys or use our providers
private inference scheduler for dedicated GPUs
budgets & policy by user, team, key or workload
MARE SERENITATIS
ONE GATEWAY · EVERY ROUTE
OCEANUS PROCELLARUM
PRIVATE DEDICATED CAPACITY
MARE NUBIUM
SCHEDULER · LEDGER · NO LOGS
// 01 - THE GATEWAY

One gateway for every model your teams already use.

Put LunaRoute in front of your provider keys or use ours. Teams get one OpenAI-compatible endpoint, finance gets one ledger, and platform engineering gets policy that follows users, keys, agents, and workloads.

BYO provider keys, LunaRoute-hosted providers, or both
budgets by user, team, key, CI job, or workload
failover and routing without changing agent config
one ledger for public and private inference spend
enterprise-gateway — topology
teams + agents
   |
   |  OpenAI-compatible or Anthropic-compatible requests
   v
+--------------------------+
| LUNAROUTE ENTERPRISE    |
|  |- policy + budgets     |
|  |- routing + failover   |
|  '- ledger + accounting  |
+-----------+--------------+
            |
  +---------+---------+---------+
  v                   v         v
BYO keys         LunaRoute   private
providers        providers  inference
// 02 - PRIVATE INFERENCE

Private inference, without idle dedicated GPUs.

Bring dedicated capacity under the same gateway. LunaRoute can run private models on servers you control, route sensitive workloads there first, and spill only the traffic you allow.

private-models — roster
glm-5.2            agentic coding default
kimi-k2.7          long-horizon planning
qwen-3.6-plus      fast, cheap workhorse
minimax-m2.7       whole-repo context
minimax-m3         m2.7's smarter successor
deepseek-v4-pro    deep reasoning & review
Single-tenant servers
Dedicated GPUs, dedicated caches, and deployment boundaries that match your risk model.
Gateway policy stays central
Keep users on one endpoint while routing sensitive workloads to your private fleet.
Utilization without routing sprawl
Keep dedicated capacity busy through the same gateway policy, ledger, and scheduler instead of sending teams to separate endpoints.
// 03 - THE SCHEDULER

Idle GPUs are expensive GPUs.

Once the rig is a fixed cost, utilization is savings. The LunaRoute scheduler does not push past the server limit; it controls who enters the backend, keeps realtime and agent work responsive, and lets flex or ballast work use the gaps.

scheduler — live
$ lunaroute-cli sched status srv-east-1

srv-east-1   8 x B200   single-tenant

admission cap ............... max_in_flight = 64
efficient target ............ b_star = 56
backend waiting ceiling ..... 0

realtime / agents ........... priority 0
flex delivery work .......... priority 10
ballast background .......... priority 100

capacity rule ............... never admit past max_in_flight
gap fill .................... use spare seats, yield on deadline risk
Priority classes
Realtime and agent calls get priority; flex and ballast work wait behind them when deadlines compete.
Headroom admission
Normal work enters only while the backend has headroom; no admission crosses the hard max_in_flight cap.
Queue with backpressure
Excess work waits in the scheduler queue or fails fast when the configured deadline can no longer be met.
Gap-filling work
Flex and ballast jobs can fill spare slots between realtime bursts, raising useful work without pretending the server has more capacity.
// 04 - DATA & PRIVACY

NO LOGS, at enterprise scale.

The same posture as the public product, contractually: request and response bodies are processed in memory and never stored by LunaRoute. On private inference, traffic stays on your single-tenant path.

No request-body storage
Prompts and completions are never written to disk by LunaRoute. We keep timestamps, model, token counts and cost - your ledger.
US-based providers
Our hosted inference runs on US-incorporated providers in US datacenters. Private inference runs wherever you place it.
Single-tenant isolation
Dedicated servers are yours alone - no shared GPUs, no shared caches, no co-mingled traffic.

Let's size your fleet.

Tell us your models, your traffic, and whether you're bringing GPUs or renting ours. We'll map it to a gateway + private inference plan.

LUNAROUTE© 2026 · 384,400 km ahead of the competition
no request bodies were logged in the making of this website