LunaRoute vs Featherless
Both serve open-weight models, and both talk about predictable costs. The biggest difference is what a flat price covers. Featherless's $25 Chat plan is for interactive chat by a person; its shared Developer API plan bills agents and automation per token. LunaRoute's coding plans are flat for included inference, including coding agents like Claude Code. Featherless's dedicated Business and hosted-agent offerings are separate, as explained below.
If you want the widest choice of open models or dedicated GPUs, Featherless is a strong option. If you want to run agents all day for a fixed monthly price, LunaRoute is built for that.
At a glance
Featherless details are from its pricing page, terms and docs, checked October 10, 2026. The source list links the references for plans, model specs, tools, integrations and privacy.
| What changes | LunaRoute | Featherless |
|---|---|---|
| Flat-price plan | Base coding plans: $99 to $499.99 a month, sized by priority concurrent requests | Chat: $25 a month, 4 concurrent units, 32K context |
| Agents and automation on the flat plan | Yes, on coding plans, subject to plan terms | Not on Chat, which is for interactive, human-driven use |
| Plan for agents | Coding plans; Personal Agent has different capacity and model scope | Shared API: Developer, from $50 a month in credits, billed per token. Business and hosted agents are separate |
| GLM 5.3 | Included in coding plans. NVFP4 with vision, 512K context | $1.40 in / $4.40 out per million tokens on Developer ($0.26 cached), FP8, 256K context |
| Model catalog | A curated set: GLM 5.3, GLM 5.3 Flash, DeepSeek 4.1 Flash, plus decision, embedding and image models | 40,000+ open models advertised |
| API formats | OpenAI, Anthropic Messages, Responses | OpenAI-compatible |
| Claude Code | One command generates settings to apply; native Anthropic format | Documented setup uses claude-code-router, a translation proxy |
| Tools around the models | Web search, OCR, document conversion, image generation, embeddings | Embeddings and speech APIs; Developer includes one hosted agent environment |
| Prompt retention | Zero data retention by default for LunaRoute-hosted inference | Says it does not log API chats, prompts or completions; metadata and hosted-app storage are separate |
| Dedicated hardware | Enterprise | Business plan, with an SLA |
The LunaRoute column describes base coding plans and hosted inference unless noted. Personal Agent starts at $19.99 with different capacity and model scope. Featherless concurrent units are model-weighted, not a count of requests. Model, plan and capacity limits apply; metered add-ons and external providers are separate. Prices and availability can change.
What “flat price” covers
This is the difference that matters most if you run agents.
Featherless's Chat plan is $25 a month with no per-token billing, but it's for a person chatting. Its terms say individual plans are “for interactive use or proto-typing and experimentation by the purchaser”, and other use can get the subscription terminated without a refund. Its own guide to coding agents puts it plainly: “The Chat plan is for human-driven use only and excludes automation, so don't try to run an agent fleet on it.” (Featherless blog, August 2026)
For Claude Code, opencode or other agents using the shared API, Featherless points you to the Developer plan. That plan starts at $50 a month in credits, and usage is billed per token. GLM 5.3 costs $1.40 per million input tokens and $4.40 per million output tokens, or $0.26 per million for cached input. Featherless only bills successful requests, which is a nice touch, and unused credits carry over without expiring. See Developer and Credits.
The plan limits differ too: Chat has 32K context and four concurrent units; Developer offers up to 256K and 100 units. Those are not 100 simultaneous requests. Each model consumes a number of units while its request runs; the documented examples use one, two or four units per request, with exceptions. Four units can mean four small-model requests or one very-large-model request. Requests above the budget receive HTTP 429. See concurrency limits.
That is the Chat-versus-Developer API comparison, not a claim that every Featherless agent workload must be token-billed. Business offers dedicated hardware at a contracted price. Its Agent Marketplace also describes hosted sandboxes above the entry Chat tier with inference bundled into the subscription. The pricing page includes one agent environment on Developer, while the Developer billing docs explicitly meter API requests. Check the terms for the particular hosted-agent offering; “bundled” should not be read as unlimited Developer API credits.
LunaRoute's coding plans are flat for agents too. You pick how many priority concurrent requests you want, from 2 on Pro ($99) to 10 on Teams ($499.99), and the price for included inference stays the same however much your agents run, within model and plan limits. Requests above your priority allowance can use available overflow capacity automatically; queue limits and timeouts still apply, and requests can be rejected when those limits are reached.
Web search includes 100 searches a month per organization. Above that you opt in to pay or bring your own key, subject to your organization's search limits. External providers and optional add-ons are separate from included inference; see web search.
Which works out cheaper depends on how much your agents run. If they run occasionally, per-token billing can cost less. If they run for hours a day, or you want to stop thinking about the meter, a flat plan is easier to live with. More in Concurrency vs token pricing.
LunaRoute also has Personal Agent at $19.99 a month for one non-priority request at a time, with its own model scope. Featherless's $25 Chat plan is cheaper than our base coding plans, not every LunaRoute plan; compare intended use, models and capacity rather than the entry price alone.
Using it with Claude Code and other coding agents
LunaRoute speaks Anthropic's Messages format natively, so Claude Code connects directly. Setup starts with one command:
npx @lunaroute/cli setup claude-code
The command prints the environment settings to apply; it does not write Claude Code's configuration or your shell profile. Follow the printed instructions, then restart Claude Code. See the setup steps.
The same CLI sets up Codex, opencode, pi, Copilot CLI, Hermes and OpenClaw. See Using Claude Code with LunaRoute.
Featherless offers an OpenAI-compatible API. Agents that speak the OpenAI format can use it directly. For Claude Code, Featherless's own guide routes requests through claude-code-router, a proxy that translates Anthropic-format requests into OpenAI format. That works, but it's one more piece to install and keep running.
Models
Featherless's catalog is much bigger: it advertises over 40,000 open models, including large coding models and many fine-tunes. If you want to try a niche model, or switch models often, that breadth is a real advantage.
LunaRoute serves a smaller, curated set and tunes how each one runs on our own GPUs: GLM 5.3, GLM 5.3 Flash and DeepSeek 4.1 Flash, plus decision, embedding and image models. With open models, the serving setup affects speed and reliability as much as the model name does, so we'd rather run a few models well. Our GLM 5.3 is a vision-enabled NVFP4 build with a 512K context window, and its weights are public on Hugging Face. Featherless describes its GLM 5.3 serving configuration as FP8 with a 256K context on Developer; see its model comparison.
Featherless also documents embeddings and speech generation, alongside its hosted agent environment.
Privacy
Both services say they don't keep API inference prompts or completions. Featherless states that it “does not log chats, prompts, or completions sent through our API” (privacy docs). Its privacy policy retains account, payment and model-usage information and reserves the right to store sampler settings. The API policy is not a blanket promise about hosted-app storage: its Marketplace docs say sandbox files and configuration are preserved.
On LunaRoute, prompts and responses for LunaRoute-hosted inference are processed in memory and discarded by default. We retain account, usage and billing metadata. Generated images and URL-based document artifacts can be retained for up to seven days; third-party traffic, including BYOK, follows the upstream provider's policy. See zero data retention.
Choose Featherless if
- You chat with models yourself and want a $25 flat plan for that, with a broad catalog.
- You need a model outside our catalog, or want to browse thousands of fine-tunes.
- You want per-token billing with a lot of concurrency headroom (100 model-weighted units on Developer, not 100 requests).
- You need dedicated GPUs with an SLA.
Choose LunaRoute if
- You run coding agents or other automation and want a flat monthly price for included inference.
- You use Claude Code and want it connected natively, without a separate translation proxy.
- You want web search, OCR, document conversion and images from the same API key.
- You'd rather not watch a token meter while agents retry and explore.
Switching from Featherless
Both use API keys and an OpenAI-compatible endpoint, so most clients need a new base URL, key and model ID: use https://gw.lunaroute.com/v1 with your lr_ key and a model from the LunaRoute catalog.
For Anthropic clients that append /v1/messages themselves, use the gateway root, https://gw.lunaroute.com, rather than the OpenAI /v1 base URL. LunaRoute accepts Anthropic Messages at /v1/messages and Responses requests at /v1/responses. For coding agents, the CLI generates the appropriate configuration; Claude Code setup prints settings to apply, and connecting directly to LunaRoute removes the need for claude-code-router.
FAQ
Can I use Featherless's $25 plan for coding agents?
No. Featherless says its Chat plan is for interactive, human-driven use and excludes automation. For shared API agent traffic, it points to the Developer plan, which is billed per token. Dedicated Business and hosted-agent offerings have their own terms; see pricing.
Is LunaRoute cheaper than Featherless?
It depends on how you use it. Featherless's $25 Chat plan costs less than LunaRoute's base coding plans, but LunaRoute also offers Personal Agent at $19.99 with one non-priority request and a different model scope. For shared API agents, Featherless Developer bills per token; LunaRoute's coding plans start at $99 a month for included inference, subject to model and capacity limits. Occasional use can favor per-token pricing; sustained usage can favor a flat plan. Compare your workload, not just the entry price.
Does Featherless work with Claude Code?
Yes. Featherless’s guide uses claude-code-router, a proxy that translates Claude Code's Anthropic-format requests to Featherless's OpenAI-compatible API. LunaRoute supports the Anthropic format directly.
Which has more models?
Featherless, by far. It advertises over 40,000 open models. LunaRoute serves a curated set, including GLM 5.3 and DeepSeek 4.1 Flash.
Do both keep my prompts private?
Both say they don't store API inference prompts or completions, but that does not mean neither service retains any data. Featherless keeps account and usage information, and hosted-app storage is separate. LunaRoute-hosted inference processes prompts in memory and retains account, usage and billing metadata; generated image and URL document artifacts can remain for up to seven days, and third-party traffic follows upstream policies. See zero data retention and Featherless privacy and logging.
Sources
Featherless sources checked October 10, 2026. Prices, plan limits, model availability and integration guidance are a dated snapshot, not a promise about future availability.
- Pricing: Chat, Developer and Business pricing, intended use, context limits, catalog size, hosted agent environment and SLA.
- Terms of service: the interactive-use restriction and termination language. The terms use Individual/Scale naming; the current pricing page identifies Chat/Developer/Business.
- Developer and Credits and concurrency limits: successful-request billing, credit rollover and model-weighted concurrent units.
- Coding-agent guide, published August 24, 2026: the Chat/automation distinction and documented Claude Code proxy setup.
- Model comparison: GLM 5.3 token rates and the vendor-described FP8/256K serving configuration.
- Embeddings and audio and speech: documented APIs beyond chat completions.
- Agent Marketplace: hosted sandboxes, bundled-inference wording and retained app state. Confirm the applicable offering's billing terms rather than assuming unlimited Developer API usage.
- Privacy and logging and privacy policy: API prompt/completion handling and retained metadata.
Run your agents on a flat monthly price
Setup starts with one command, and Claude Code connects without a separate translation proxy.