LunaRoute vs Together AI
Together AI is a full AI cloud: serverless inference on 200+ models, with chat billed per token, dedicated GPU endpoints, GPU clusters, fine-tuning and code sandboxes. LunaRoute does one thing: it runs a curated set of open models for agents, for a fixed monthly price for included inference, with zero data retention by default on LunaRoute-hosted inference.
Choose Together if you need its breadth, fine-tuning or your own GPUs. Choose LunaRoute if you run agents on open models every day and want a predictable bill without building on a full cloud platform.
At a glance
Together details are from its pricing page, model docs and privacy policy, checked October 9, 2026. The source list links the references for billing, retention, Together Link and platform features.
| What changes | LunaRoute | Together AI |
|---|---|---|
| How you pay | Base coding plans: $99 to $499.99 per month, sized by priority concurrent requests | Per token for serverless chat; per GPU-hour on dedicated endpoints |
| GLM 5.3 | Included in coding plans. NVFP4 with vision, 512K context | $1.40 in / $4.40 out per million tokens ($0.26 cached), FP4, 1M context |
| DeepSeek 4.1 Flash | Included in coding plans. Original weights, 1M context | $0.30 in / $1.20 out per million tokens, FP8, 1M context |
| Catalog | A curated set of open models, plus decision, embedding and image models | 200+ models across chat, image, video, audio and embeddings |
| Claude Code | One command generates settings to apply; native Anthropic format | Together Link (beta), which routes through a hosted gateway |
| Prompt retention | Zero data retention by default for LunaRoute-hosted inference | Zero data retention is optional and off until you turn it on |
| Fine-tuning | No | Yes (SFT and DPO) |
| Dedicated GPUs | Enterprise | Dedicated endpoints and GPU clusters, self-serve |
| Tools around the models | Web search, OCR, document conversion, image generation | Code sandbox, Code Interpreter |
| Free trial | No | No. Minimum $15 for an organization's first credit purchase; $5 after that |
The LunaRoute column describes hosted inference and base coding plans; other plans differ. Model and capacity limits still apply. Metered add-ons and external providers are separate. Together's image, video and audio billing units vary by model. Prices, features and availability can change.
How you pay
Together's serverless prices are competitive. GLM 5.3 is $1.40 per million input tokens and $4.40 per million output tokens, and DeepSeek V4.1 Flash is $0.30 and $1.20. For light or occasional use, that's hard to beat.
For agents, the bill follows the agent. Each step can resend context, repeated inference calls add billed tokens, and a long session can cost more than a short one. Cached-input discounts can reduce that cost. With dedicated endpoints you pay for GPU hours instead, which brings its own question of how busy you can keep them (see Should you rent GPUs or use managed inference?).
LunaRoute's base coding plans are a fixed monthly price for 2 to 10 priority concurrent requests, with available overflow included. How much your agents run doesn't change the included-inference bill, within model and plan limits. Overflow depends on available capacity; queue limits and timeouts still apply, and clients should use bounded retries. Metered add-ons and external providers are separate. More in Concurrency vs token pricing.
Data retention
This is a real difference if your agents work on private code or documents.
Together says it doesn't train on your data without your explicit opt-in. Its privacy and security docs say prompt and response storage is on by default. To enable zero data retention for organization API keys, an organization admin must open Organization Settings → Privacy and set Store prompts and model responses to No. The personal-account toggle is separate; it does not control an organization's traffic.
The ZDR docs say the change applies from that point forward, not retroactively, and disables passthrough models. Together still retains account and usage metadata, and files explicitly uploaded for fine-tuning or batch jobs remain until you delete them. The privacy policy, last updated December 17, 2025, describes these data practices.
On LunaRoute, zero data retention is the default for LunaRoute-hosted inference. Prompts and responses are processed in memory and discarded. We retain account, usage and billing metadata. Generated images and URL-based document artifacts can be retained for up to seven days; third-party traffic, including BYOK, follows the upstream provider's policy. See zero data retention.
Claude Code
Together has put real work into coding agents. Together Link, currently in beta, launches Claude Code with temporary settings that route through Together's hosted gateway. Its model menu maps Claude's tiers to open models (Opus to Kimi K3, Fable to GLM 5.3, Sonnet to DeepSeek V4.1 Flash), and usage on Together-hosted models is billed at serverless per-token rates (Together Link docs).
Link uses Auto routing by default. In Claude Code sessions with an Anthropic API key configured, Auto can send difficult requests to Claude Opus and bill the Anthropic account behind that key separately. Without that key, those sessions stay on Together-hosted models; you can also pin a model rather than use Auto.
LunaRoute starts Claude Code setup with one command and supports the setup directly:
npx @lunaroute/cli setup claude-code
The command prints the environment settings to apply; it does not write Claude Code's configuration or your shell profile. Follow the printed instructions, then restart Claude Code. See the setup steps.
The difference is less about setup and more about what happens after: Together-hosted model usage is billed per token, and on LunaRoute it's part of your fixed plan's included inference. See Using Claude Code with LunaRoute.
Breadth vs focus
Together is the broader platform. You can fine-tune a model, deploy it on dedicated GPUs that you can scale to zero, rent a GPU cluster, run code in a sandbox, and generate images, video and audio, all in one account. If you need several of those, Together is a strong choice.
LunaRoute is narrower on purpose. We serve a small set of open models on our own stack and tune it for agent work: interactive requests ahead of background jobs, shared context cached between calls, and available overflow when you go past your coding plan's priority capacity. We also bundle the tools agents reach for most often, such as web search, OCR and document conversion.
Choose Together AI if
- You want to fine-tune models and deploy them on dedicated GPUs.
- You need image, video or audio models alongside chat.
- Your usage is light, and per-token pricing costs less than a monthly plan for your workload.
- You want raw GPU clusters as well as inference.
Choose LunaRoute if
- You run agents on open models every day and want a fixed monthly bill for included inference.
- You want zero data retention on hosted inference without having to remember a setting.
- You want Claude Code on open models as part of a flat plan.
- You'd rather use one focused service than build on a full cloud platform.
Switching from Together AI
Both offer OpenAI-compatible endpoints, so most clients need a new base URL, key and model ID: use https://gw.lunaroute.com/v1 with your lr_ key and a model from the LunaRoute catalog.
LunaRoute also accepts Anthropic Messages at /v1/messages and Responses requests at /v1/responses on the same gateway. For Anthropic clients that append /v1/messages themselves, use the gateway root, https://gw.lunaroute.com, rather than the OpenAI /v1 base URL. For coding agents, the LunaRoute CLI generates the appropriate configuration; the Claude Code setup prints environment settings for you to apply.
FAQ
Is LunaRoute cheaper than Together AI?
It can be for heavy, agent-driven use of open models, but savings depend on the model, token volume and caching. A fixed price makes the included-inference bill predictable; it does not guarantee the lowest cost. For light use, Together's per-token prices will often cost less. Compare your actual workload using Concurrency vs token pricing.
Does Together AI keep my prompts?
By default, Together stores prompts and responses. You can enable zero data retention in your organization's privacy settings; it applies going forward and disables passthrough models. Training is a separate opt-in and is off by default. See Together's ZDR docs. LunaRoute applies zero data retention by default to LunaRoute-hosted inference, with the metadata, artifact and third-party exceptions described above.
Can I use Claude Code with Together AI?
Yes, through Together Link (in beta), which can route Claude Code to open models such as GLM 5.3 and bill per token. Auto can also call Claude Opus when you supply an Anthropic API key, billed separately to that account. LunaRoute supports Claude Code on open models as part of included inference; one setup command prints the settings to apply.
Does LunaRoute offer fine-tuning?
No. LunaRoute serves a curated set of open models. If you need fine-tuning, Together AI or a dedicated setup is a better fit.
What precision does Together serve GLM 5.3 at?
Together's model docs list GLM 5.3 at FP4 with a 1M-token context, checked October 9, 2026. LunaRoute serves a vision-capable NVFP4 build of GLM 5.3 with a 512K context; the weights are public. Precision and context are configuration differences, not a quality ranking.
Sources
Together sources checked October 9, 2026. Prices, model configurations and beta features are a dated snapshot, not a promise about future availability.
- Pricing and serverless model specs: token prices, cached-input rates, precision, context limits and GPU-hour pricing.
- Model catalog: Together's advertised 200+ models and supported modalities.
- Privacy and security, zero data retention and privacy policy: storage defaults, training opt-in, organization settings and retention exceptions.
- Together Link: beta status, temporary configuration, model mappings and Auto routing/billing.
- Credits: no general free trial; $15 first credit purchase and $5 minimum thereafter.
- Preference fine-tuning: SFT and DPO workflows.
- Fine-tuned model deployment: dedicated inference billing and scaling to zero. The pricing page also covers GPU clusters, sandboxes and Code Interpreter.
Keep your agents running for one monthly price
Setup starts with one command, and zero data retention is on from the start for LunaRoute-hosted inference.