Standard
Every key, by default
The same GPU work also mines PRL, which keeps prices low.
An OpenAI-compatible API for open models. Change the base URL and your key; keep your code.
Python
from openai import OpenAI client = OpenAI(base_url= "https:// api. peryn. ai/ v1") chat = client. chat. completionsstream = chat. create(model=MODEL, messages=messages, stream=True)Quickstart
01
Get a key
Sign in to the console and create a key. For API credit, talk to us; card checkout opens at launch.
02
Set the base URL
https://api.peryn.ai/v1
03
Send a request
Any OpenAI SDK. Streaming works the same way.
Code
Python
2 lines changed
import osfrom openai import OpenAI client = OpenAI(Removed: base_url="https:// api. openai. com/ v1",Removed: api_key=os. environ["OPENAI_API_KEY"],Added: base_url="https:// api. peryn. ai/ v1",Added: api_key=os. environ["PERYN_API_KEY"],) reply = client. chat. completions. create( model=MODEL, messages=messages,)JavaScript
2 lines changed
import OpenAI from "openai"; const client = new OpenAI({Removed: baseURL: "https:// api. openai. com/ v1",Removed: apiKey: process. env. OPENAI_API_KEY,Added: baseURL: "https:// api. peryn. ai/ v1",Added: apiKey: process. env. PERYN_API_KEY,}); const reply = await client. chat. completions. create({ model: MODEL, messages,});
MODEL is a model id from the table below, for example llama-3.1-8b-instruct.
Models
Prices in USD per million tokens.
| Model | Input / 1M | Output / 1M | |
|---|---|---|---|
Llama 3.1 8B Instructllama-3.1-8b-instruct | 32K | $0.017 | $0.034 |
15% below DeepInfra, the cheapest listed provider for this model. | |||
The first request after a quiet period can take a few minutes.
Endpoints
OpenAI-compatible. Send your key as a Bearer token.
https://api.peryn.ai/v1
/chat/completionsChat, streamed or in one piece.API key/completionsText completion from a prompt.API key/modelsEvery model: id, context, price.No key/models/{id}One model.No key/catalogPrices per 1M tokens and context, per model.No keyStreaming
Set stream to true. Tokens arrive as server-sent events, and the stream ends with data: [DONE].
import osfrom openai import OpenAI client = OpenAI(base_ url= "https:// api. peryn. ai/ v1",api_ key= os. environ[ "PERYN_ API_ KEY"],) stream = client. chat. completions. create(model= "llama-3.1-8b-instruct",messages=[ {"role": "user", "content": "Hello"}],stream=True,)for chunk in stream:if chunk. choices:print( chunk. choices[ 0]. delta. content or "", end="")text/event-stream
: ping
data: {…, "choices": [{"index": 0, "delta": {"role": "assistant", "content": "Hel"}}]}
data: {…, "choices": [{"index": 0, "delta": {"content": "lo!"}}]}
data: {…, "choices": [], "usage": {"prompt_tokens": 9, "completion_tokens": 2, "total_tokens": 11}}
data: [DONE]
Add stream_options.include_usage for token counts in the last chunk.
Lines that start with a colon keep the connection open; OpenAI SDKs skip them.
Tiers
Each API key has a tier. Neither stores prompts or answers, and we never train on them.
Every key, by default
The same GPU work also mines PRL, which keeps prices low.
On request
Never mined. Vetted nodes only, priced from GPU cost.
Talk to us about PrivateStandard requests run on GPUs that also mine: our pool briefly checks a small slice of intermediate model values in memory, then discards them. Private requests are never mined. Compute privacy
Set region rules, such as EU only, for your account or one key in the console. Every response a GPU served names its
country in x-pc-node-country.
Credits
You pay for the tokens each request used.
Prepaid credit
Too little credit for the hold: 402. A lower max_tokens keeps holds small.
Limits per key
The defaults. A request counts its prompt plus max_tokens, and gets back what it did not use. Over a limit: 429 with Retry-After.
Errors
Every error has the same JSON body. OpenAI SDKs raise it with the status code.
Response
HTTP/1.1 402 Payment Required
{ "error": { "message": "Not enough credit for this request. Lower max_tokens, or get more credit: see the console.",For people. Log it or show it. "type": "insufficient_quota",The kind of error. "param": null,The field at fault, when there is one. "code": "insufficient_quota"Stable. Branch on it. }}| Status | Code and what to do |
|---|---|
| 400 | invalid_jsonunsupported_parameter Fix the request. A retry gives the same answer. |
| 401 | missing_api_keyinvalid_api_keyrevoked_api_key Check the key in the Authorization header. |
| 402 | insufficient_quota Get more credit (talk to us), or lower max_tokens. |
| 404 | model_not_found Use a model id from /models. |
| 429 | rate_limit_exceeded Wait for Retry-After, then retry. |
| 429 | no_capacity Every GPU for this model is busy. Retry shortly. |
| 503 | no_capacity The wait for a GPU ran out. Retry shortly. |
| 504 | timeout Retry, or lower max_tokens. |
Every code, in full: API docs
Contact