Pay less for open models.
Per-token prices for the API and the chat, below the cheapest listed provider. Per-token prices for open models, for the API and the chat.
Prices
Pay only for the tokens you use.
-
Llama 3.1 8B Instruct
32K context
- Input / 1M
- $0.017
- Output / 1M
- $0.034
15% below DeepInfra.
Input / 1MOutput / 1M
USD per million tokens, excluding VAT. Prices are dynamic and can change every hour.
Read the API docsChat
Free chat, every day.
Free tier
50,000
tokens a day after you sign in
For the web chat only. After that, chat requests use your credits at the prices above.
Start chattingPrivacy tiers
Two tiers. One is never mined.
Neither tier stores prompt or completion content. The difference is mining.
Standard
A GPU node in our network answers, and the same work also mines PRL.
The prices above.
To verify mining, our pool briefly checks a small slice of intermediate model values in memory, then discards them. Compute privacy
GPU partners
Your GPUs earn twice.
Teams
Private deployment for teams: talk to us.
Dedicated GPUs, your own routing rules, no mining on your traffic.
Questions
Billing, credits and cold starts.
How do I pay for the API?
The API is prepaid in US dollars, excluding VAT, and each request is charged against your credit. Card checkout opens at launch; to get API credit now, talk to us. The free chat tokens need no credit.
How is a request charged?
By tokens, counted with each model’s own tokenizer. When a request starts, we hold enough credit for its input plus its longest possible answer (max_tokens). When it ends, we charge what it used and release the rest. Without enough credit, the request is refused with HTTP 402.
Do credits expire?
Unused credits expire after 12 months without activity on your account. Activity means signing in, a request charged to your credits, or buying credits. We email you at least 30 days before.
Are there rate limits?
Each API key allows 60 requests and 200,000 tokens per minute by default, and the console shows each key’s limits. Over a limit you get HTTP 429 with a Retry-After header.
Need more? Talk to us Why can the first request take longer?
GPUs with nothing to do are switched off. The first request after a quiet period can take a few minutes. We hold it for up to 4 minutes for a GPU to come online. If none does, you get HTTP 503 and pay nothing.
What happens to my prompts?
We do not log or store prompt or completion content, and never train on it. How mining affects your data:
Compute privacy