Build until it works. Not until credits run out.

Iteration is part of building. Run your coding agents on supported open models with fixed-price inference, so another attempt doesn’t add a token charge for included usage.

Zero retention for inference prompts and responses. Never used for training.

From $99/mo. No monthly token cap. Queue limits apply; external providers are billed separately.

Trusted by teams at

Open models.
One fixed price.

Run your agents, applications, and automations on supported open models. Your subscription covers included inference, with no per-token charges.

Keep your existing tools. Connect through our OpenAI-compatible API.

Explore supported models

Included inference

  • GLM 5.3
  • GLM 5.3 Flash
  • DeepSeek 4.1 Flash

Model access varies by plan.

Decide. Search. Create.

  • Djev

    A System One decision model for classification, choices, and scoring.

  • 4 embedding models

    Find related content with semantic search.

  • 2 image models

    Generate new images and edit existing ones.

Give your agents more to work with.

  • Search the web

    Bring fresh context into your workflows with Exa or Brave.

  • Read scanned documents

    Extract usable text with AI-powered OCR.

  • Convert documents

    Turn PDFs, documents, slides, and spreadsheets into Markdown your agents can use.

Your concurrency. Room to burst.

Choose how many requests get priority. Additional requests use available overflow capacity automatically.

Max plan example

4 priority requests

Preferential scheduling

  1. 1
  2. 2
  3. 3
  4. 4

Automatic overflow

Additional requests · as capacity allows

  1. 5
  2. 6
  3. 7
  4. 8

4 requests receive priority. 4 more use available overflow capacity.

$199.99 / month

Max · monthly billing

Same monthly price.

No per-token charges for included inference.

Each numbered square is one request in this example. Overflow runs as capacity allows; plan limits and timeouts apply. Priority is not a speed guarantee. Read the concurrency terms.

Put your agents to work.

Code and iterate

Explore code, try ideas, and review changes without counting tokens.

Run background jobs

Evaluate, index, and process documents when completion times are flexible.

Share with your team

Share concurrent capacity across developers, agents, and CI. Teams adds per-member concurrency controls.

Choose your concurrent requests.

Choose your priority concurrency. Overflow included. No per-token charges for included inference.

Every coding plan includes: private prompts and responses, curated models, an OpenAI-compatible API, API keys, usage reports, and member management.

Pro

$99.00/mo
Priority scheduling
2 concurrent requestsThis many requests can run at once with priority scheduling. Priority does not guarantee a particular response time or generation speed.
Overflow included
Email support
Change plans anytime
Start with Pro

Max

$199.99/mo
Priority scheduling
4 concurrent requestsThis many requests can run at once with priority scheduling. Priority does not guarantee a particular response time or generation speed.
Overflow included
Email support
Change plans anytime
Start with Max

Teams

$499.99/mo
Priority scheduling
10 concurrent requestsThis many requests can run at once with priority scheduling. Priority does not guarantee a particular response time or generation speed.
Overflow included
Dedicated Slack support
Per-member controls
10-request bundles
Start with Teams
Compare all plan features

Teams adds per-member concurrency controls and expandable capacity bundles.

Swipe sideways to compare all four plans.

Features of Pro, Max, Apex, and Teams
FeatureProMaxApexTeams
Priority concurrent requests24610
Add and remove membersIncludedIncludedIncludedIncluded
Set priority and overflow concurrency per memberNot includedNot includedNot includedIncluded
Add capacity to your current planMove to a larger packageMove to a larger packageMove to a larger packageBundles of 10 priority concurrent requests plus overflow capacity
API keys, routing, and usage reportingIncludedIncludedIncludedIncluded
SupportEmail supportEmail supportPriority email supportDedicated Slack support

Overflow limits are additional to priority concurrency. Overflow uses available capacity and may wait during busy periods; queue limits and timeouts apply. All coding plans share infrastructure and the same curated models.

One personal agent?

One request at a time, scheduled as capacity becomes available.

Need dedicated infrastructure?

Enterprise offers dedicated deployments and custom routing.

Explore Enterprise

Shared infrastructure. Overflow may wait during busy periods; capacity and queue limits apply. External providers and optional tools cost extra.

Private prompts. Clear data handling.

We don't store your inference prompts or responses, and we never use them to train models.

Prompts and responses

Processed in memory, then discarded. We do not store your prompts or responses.

Usage metadata

Timestamps, models, and token counts are kept for operations, billing, and your usage ledger.

Never used for training

Your inference prompts and responses are never used to train models.

Read the privacy policy
private inference · the default
your agent
    |
    v
process prompt in memory
    |
    v
return response to your agent
    |
    +-- prompt & response: not retained
    '-- usage metadata: your ledger

Connect in three steps.

  1. Find your fit

    Check that our models suit the work you do, from coding and research to everyday agent tasks.

    Explore use cases →
  2. Get your API key

    Choose a plan, create an account, and connect your agent.

    Setup guide →
  3. Try a real task

    Test an included model, then check the response and usage ledger.

    First request →
FAQ

Before you choose.

Your bill stays fixed.

Priority for everyday work. Room to burst when work picks up.

We have special plans for crustaceans!

Explore Personal Agent plans