Skip to content

Pay less for open models.

Per-token prices for the API and the chat, below the cheapest listed provider.

Prices

Pay only for the tokens you use.

  • Llama 3.1 8B Instruct

    32K context

    Input / 1M
    $0.017
    Output / 1M
    $0.034

    15% below DeepInfra.

    Input / 1M
    $0.02DeepInfra
    $0.017Peryn
    Output / 1M
    $0.04DeepInfra
    $0.034Peryn

USD per million tokens, excluding VAT. Prices are dynamic and can change every hour.

Read the API docs

Chat

Free chat, every day.

Free tier

50,000

tokens a day after you sign in

For the web chat only. After that, chat requests use your credits at the prices above.

Start chatting

Privacy tiers

Two tiers. One is never mined.

Neither tier stores prompt or completion content. The difference is mining.

Standard

A GPU node in our network answers, and the same work also mines PRL.

The prices above.

Private

Never used for mining, and routed only to vetted nodes.

Talk to us about Private

To verify mining, our pool briefly checks a small slice of intermediate model values in memory, then discards them. Compute privacy

GPU partners

Your GPUs earn twice.

Want jobs from our API on your GPUs? Talk to us.

Payouts begin at launch.

GPU Clouds

Teams

Private deployment for teams: talk to us.

Dedicated GPUs, your own routing rules, no mining on your traffic.

We use your name, work email, company and message only to reply and to discuss working with you. We delete the stored enquiry 12 months after your last message, and keep our email thread with you for 2 years after our last reply. Privacy

Questions

Billing, credits and cold starts.

How do I pay for the API?

The API is prepaid in US dollars, excluding VAT, and each request is charged against your credit. Card checkout opens at launch; to get API credit now, talk to us. The free chat tokens need no credit.

How is a request charged?

By tokens, counted with each model’s own tokenizer. When a request starts, we hold enough credit for its input plus its longest possible answer (max_tokens). When it ends, we charge what it used and release the rest. Without enough credit, the request is refused with HTTP 402.

Do credits expire?

Unused credits expire after 12 months without activity on your account. Activity means signing in, a request charged to your credits, or buying credits. We email you at least 30 days before.

Are there rate limits?

Each API key allows 60 requests and 200,000 tokens per minute by default, and the console shows each key’s limits. Over a limit you get HTTP 429 with a Retry-After header.

Need more? Talk to us

Why can the first request take longer?

GPUs with nothing to do are switched off. The first request after a quiet period can take a few minutes. We hold it for up to 4 minutes for a GPU to come online. If none does, you get HTTP 503 and pay nothing.

What happens to my prompts?

We do not log or store prompt or completion content, and never train on it. How mining affects your data:

Compute privacy
How a request is charged: the worst case is held, what it used is charged, the rest goes back.