← All guides

Should you rent GPUs or use managed inference?

Renting GPUs gets you hardware. What you and your agents actually feel is the serving system on top of it: how the model is quantized, which engine runs it, how requests are scheduled and cached, and what happens under load. Building that system takes specialist skills in model weights, GPU memory, caching, load balancing and monitoring, and keeping it good is a permanent job. Unless serving models is your business, a managed service is usually the better use of your time and money.

Quick comparison

Renting GPUs compared with LunaRoute managed inference
What changesRent GPUsManaged inference (LunaRoute)
What you getGPUs and a blank serverA tuned serving system with models ready to call
What you pay forGPU hours, whether they're busy or notA fixed monthly plan, sized by concurrent requests
Quantization, engine, parsersYou choose, test and maintain themWe do
Scheduling and cachingYou build and tune itBuilt in: priority, background and overflow
Upgrades and regressionsYou test every changeWe test, and run the stack ourselves
Expertise neededInference engineering: weights, memory, caching, load balancing, monitoringCalling an API
Setup timeDays to weeks to tune and validate for productionMinutes to the first API call
Model choiceAny open model, including your own fine-tunesThe models in our catalog

The managed column describes LunaRoute coding plans and included inference; metered add-ons are separate. Setup times are estimates: prebuilt images can shorten a self-hosted start, and either approach still needs evaluation on your tasks.

Renting gets you GPUs, not an inference service

I like building several projects at once, and juggling multiple Claude Max accounts and Codex just to stay under each provider's limits got exhausting. Running your own GPUs is the obvious way out of other people's limits, and I enjoyed building my own rig. Running models for real work is a different job, though. This is the list of what you're signing up for:

The performance figures below are reported examples from our serving-stack tests, described in The Model Is Not the Product. They are not performance guarantees or benchmarks of the H200 rental configuration priced below.

  1. Fitting the model. GLM 5.3 has 753 billion parameters (model card). At FP8 that's roughly 750 GB of weights before you leave room for context, which is more than an 8× 80 GB H100 node holds (640 GB). An 8× H200 node is one option; other GPU configurations or more aggressive quantization can also work.
  2. Choosing a precision. You can run the full-precision weights, FP8, or a heavier quantization that fits on fewer GPUs. Compression can hurt a hard coding task, make no difference, or occasionally help. You only find out by testing on your own workload.
  3. Configuring the engine. vLLM and SGLang both work, but the result depends on configuration and patches: the chat template, sampling defaults, the tool-call parser, the reasoning parser. A tool-call parser that's slightly wrong doesn't throw errors. Your agent just picks the wrong tool more often.
  4. Scheduling. When a long batch job and an interactive session hit the same GPUs, one of them waits. On our own 8× B200 cluster, a single change to how background work competes with interactive requests cut interactive latency from 2.7 seconds to 0.4 seconds, without touching the model, the weights or the hardware. Background jobs got slower in exchange. That's a decision someone has to make, measure and own. I wrote more about this in The Model Is Not the Product.
  5. Finding your saturation point. More simultaneous requests isn't always more output. In one of our tests, 87 concurrent requests produced 42,700 tokens per second, and pushing to 188 dropped it to 25,100. Past a certain point, the overhead of juggling requests eats the throughput.
  6. Caching. Agents send the same instructions and files over and over. Keeping that shared prefix in GPU memory cut our time to first token from 206 ms to 68 ms, and even a copy in system memory beat recomputing it (0.96 s vs 2.54 s). How much you cache, where, and when you evict it is yours to design.
  7. Starting and scaling. Large models take minutes to load. One vLLM user reported about 8 minutes from launch to ready for a DeepSeek model on an 8-GPU node (GitHub issue). So you either keep nodes warm and pay for idle time, or your users wait.
  8. Catching silent regressions. When an engine upgrade or a config change makes the model slightly worse, nothing goes red. Requests still return 200 OK while your agent's success rate drifts down. You need evaluations and canary tasks running all the time to notice.
  9. Doing it again for every new model. Open models improve every few weeks. Each release means new weights, new templates and parsers, new tuning, and another round of testing.

It takes a specialist skill set

Each item on that list is its own area of expertise, and running a model well takes all of them at once:

Inference engineering skills and operational risks
AreaWhat you need to knowWhat goes wrong without it
Weights and precisionCheckpoints, quantization formats, and how compression affects accuracy on different tasksQuality drops on hard tasks and nobody can say why
GPU memoryHow weights, KV cache, context length and batch size share memory across tensor-parallel GPUsOut-of-memory crashes under load, or capacity left unused
Serving enginesvLLM or SGLang internals: chat templates, parsers, kernels and the patches each new model needsBroken tool calls and slow output from a "working" server
CachingPrefix caching, eviction, and whether to keep cache in GPU or system memoryAgents recompute the same context on every call
Scheduling and load balancingRouting requests across replicas and nodes, prioritizing interactive work, finding the saturation pointSlow responses for some users while GPUs sit half idle
Monitoring and evaluationLatency and throughput dashboards plus ongoing quality tests on real tasksRegressions that only show up as failed agent runs
OperationsCapacity planning, failover, upgrades and being on callOutages at 2 a.m. and a model nobody dares to upgrade

Few engineers have all of these. Most teams that self-host need a dedicated inference engineer plus infrastructure support, and that's people's time taken away from the product they set out to build.

Why small serving problems hurt agents most

A chatbot that's a little off still gives you a usable answer. As a simplified illustration, suppose an agent task needs 30 independent steps to all succeed. If each step is 95% reliable, the whole task succeeds about 21% of the time. At 97% per step, it succeeds about 40% of the time. Two points of per-step reliability nearly double the number of finished tasks in this illustration; real agent errors can be correlated, and agents can recover from mistakes.

That's why a homegrown stack that's "almost right" is expensive. A slightly worse quantization or a flaky tool-call parser shows up as retries, failed runs and time spent figuring out what went wrong. The number to watch is the cost of a finished task, which the GPU's hourly price doesn't tell you.

What the hardware alone costs

These hardware-only estimates use on-demand rates listed on October 9, 2026, for an illustrative GLM 5.3 FP8 configuration with 8× NVIDIA H200, running all month (730 hours). RunPod totals multiply the advertised per-GPU starting rate by eight; actual eight-GPU node availability and pricing can differ.

Estimated hardware-only cost of eight H200 GPUs
ProviderPer hourPer month
RunPod, Community Cloud (8 × $3.59)$28.72about $21,000
RunPod, Secure Cloud (8 × $5.29)$42.32about $30,900
AWS p5en.48xlarge$63.30about $46,200

Sources: RunPod pricing, AWS p5en pricing via Vantage. Rates vary by region, availability and date. Marketplaces, spot and reserved capacity can lower these, in exchange for interruptions or long commitments. Those figures cover one node and one model. They don't include a second node for redundancy, storage, or the engineering time from the list above.

A dedicated eight-GPU node is not capacity-equivalent to a shared LunaRoute subscription. Compare the capacity and total cost your own workload actually needs.

You also pay for every idle hour. Cast AI's 2026 report found average GPU utilization of 5% in its sample of non-optimized Kubernetes clusters (Cast AI). That is a broad infrastructure sample, not a forecast of your agent workload's utilization. Agent work comes in bursts, so a node sized for your busiest afternoon can sit mostly idle the rest of the day.

What you get with LunaRoute

We run the whole stack ourselves, from the GPUs to the scheduler. That's where the performance measurements above come from. The work on that list is our full-time job: choosing precisions, tuning engines, separating interactive from background work, caching shared context, and testing every change before it reaches you.

You pick how many priority requests you want running at once, and the price for included inference stays the same however busy your month is. Requests above your priority allowance can run automatically on available overflow capacity on coding plans; queue limits and timeouts still apply. Because retries and long runs don't add to the included-inference bill, you can let an agent try a second approach, or compare models on your own tasks, without doing the math first. Metered add-ons and external providers are separate. Coding plans start at $99 a month.

When renting your own GPUs still makes sense

It's the right call for some teams, and we'd tell you so. Usually it's when most of these are true:

  • Serving models is core to your product, and you have engineers who want to own the serving system full time.
  • Your load is steady and high enough to keep nodes busy around the clock.
  • You need a model we don't serve, or your own fine-tuned weights.

If what you need is dedicated hardware or data kept in a specific place, you don't have to run it yourself. Our Enterprise plans include dedicated infrastructure that we operate for you.

FAQ

What's the hardest part of running your own inference server?

Keeping it good. Getting a model to respond can be a weekend's work. Running it well takes a mix of skills that few engineers have together: choosing precision for the weights, planning GPU memory, designing caching, balancing load across nodes, and monitoring quality after every upgrade. It's ongoing work, usually for a dedicated inference engineer.

Is it cheaper to self-host an open model than to use an API?

It can be at high, steady utilization, even after accounting for engineering. Compare the total cost, including your team's time. The eight-H200 on-demand examples above work out to roughly $21,000 to $46,000 a month before operations, depending on the provider, and rented capacity is billed even while idle.

What GPUs do I need to run GLM 5.3?

GLM 5.3 has 753 billion parameters, so the FP8 weights alone take about 750 GB. That fits on an 8× H200 node (1,128 GB in total), but not on 8× 80 GB H100s (640 GB) unless you use a more aggressive quantization. Other GPU configurations can also provide enough memory.

Are two providers serving the same model the same?

No. Providers can serve the same model name with different precisions, engines, parsers, hardware, caching and scheduling, and the results differ. A model's benchmark scores tell you about the weights. Test the endpoint itself on your own tasks. More in The Model Is Not the Product.

How is LunaRoute different from renting a GPU?

You don't rent hardware or build a serving system. You pick a coding plan by the number of priority concurrent requests, from 2 to 10, and we run and tune the models, the engine, the scheduler and the fleet. Requests above your priority allowance can run automatically on available overflow capacity; queue limits and timeouts still apply.

Does my data stay private with managed inference?

For LunaRoute-hosted inference, prompts and responses are processed in memory and then discarded, and we never use them for training. We retain account, usage and billing metadata. Generated images, document artifacts and external providers have separate retention terms. How zero data retention works

Spend your time on what you're building

Let us run the serving system. You keep the fixed monthly bill and the room to experiment.

See plans and pricing