Skip to content

Open models. Lower prices. Two lines of code.

An OpenAI-compatible API for open models. Change the base URL and your key; keep your code.

Python

from openai import OpenAI client = OpenAI(base_url=  "https://api.peryn.ai/v1") chat = client.chat.completionsstream = chat.create(model=MODEL,    messages=messages, stream=True)

Quickstart

Three steps to a first answer.

  1. 01

    Get a key

    Sign in to the console and create a key. For API credit, talk to us; card checkout opens at launch.

  2. 02

    Set the base URL

    https://api.peryn.ai/v1

  3. 03

    Send a request

    Any OpenAI SDK. Streaming works the same way.

Open the console

Code

The same client, two lines changed.

Python

2 lines changed

import osfrom openai import OpenAI client = OpenAI(Removed:     base_url="https://api.openai.com/v1",Removed:     api_key=os.environ["OPENAI_API_KEY"],Added:     base_url="https://api.peryn.ai/v1",Added:     api_key=os.environ["PERYN_API_KEY"],) reply = client.chat.completions.create(    model=MODEL,    messages=messages,)

JavaScript

2 lines changed

import OpenAI from "openai"; const client = new OpenAI({Removed:   baseURL: "https://api.openai.com/v1",Removed:   apiKey: process.env.OPENAI_API_KEY,Added:   baseURL: "https://api.peryn.ai/v1",Added:   apiKey: process.env.PERYN_API_KEY,}); const reply = await client.chat.completions.create({  model: MODEL,  messages,});

MODEL is a model id from the table below, for example llama-3.1-8b-instruct.

Models

Open models, priced per token.

Prices in USD per million tokens.

Models on the API: context and prices per million tokens, USD
ModelInput / 1MOutput / 1M
Llama 3.1 8B Instructllama-3.1-8b-instruct$0.017$0.034
Input
$0.017 vs $0.02
Output
$0.034 vs $0.04

15% below DeepInfra, the cheapest listed provider for this model.

The first request after a quiet period can take a few minutes.

Endpoints

The endpoints your SDK calls.

OpenAI-compatible. Send your key as a Bearer token.

https://api.peryn.ai/v1

  • POST/chat/completionsChat, streamed or in one piece.API key
  • POST/completionsText completion from a prompt.API key
  • GET/modelsEvery model: id, context, price.No key
  • GET/models/{id}One model.No key
  • GET/catalogPrices per 1M tokens and context, per model.No key

Streaming

Streams like OpenAI.

Set stream to true. Tokens arrive as server-sent events, and the stream ends with data: [DONE].

import osfrom openai import OpenAI client = OpenAI(base_url="https://api.peryn.ai/v1",api_key=os.environ["PERYN_API_KEY"],) stream = client.chat.completions.create(model="llama-3.1-8b-instruct",messages=[{"role": "user", "content": "Hello"}],stream=True,)for chunk in stream:if chunk.choices:print(chunk.choices[0].delta.content or "", end="")

text/event-stream

: ping
data: {…,"choices":[{"index":0,"delta":{"role":"assistant","content":"Hel"}}]}
data: {…,"choices":[{"index":0,"delta":{"content":"lo!"}}]}
data: {…,"choices":[],"usage":{"prompt_tokens":9,"completion_tokens":2,"total_tokens":11}}
data: [DONE]

Add stream_options.include_usage for token counts in the last chunk. Lines that start with a colon keep the connection open; OpenAI SDKs skip them.

Tiers

Two privacy tiers.

Each API key has a tier. Neither stores prompts or answers, and we never train on them.

Standard

Every key, by default

The same GPU work also mines PRL, which keeps prices low.

Private

On request

Never mined. Vetted nodes only, priced from GPU cost.

Talk to us about Private

Standard requests run on GPUs that also mine: our pool briefly checks a small slice of intermediate model values in memory, then discards them. Private requests are never mined. Compute privacy

Set region rules, such as EU only, for your account or one key in the console. Every response a GPU served names its country in x-pc-node-country.

Credits

Prepaid credit. Clear limits.

You pay for the tokens each request used.

Prepaid credit

  • Held while a request runs: its input plus max_tokens, at the model’s price
  • Charged when it ends: the tokens it used
  • The rest of the hold goes back to your credit

Too little credit for the hold: 402. A lower max_tokens keeps holds small.

Limits per key

  • 60 requests per minute
  • 200,000 tokens per minute

The defaults. A request counts its prompt plus max_tokens, and gets back what it did not use. Over a limit: 429 with Retry-After.

Where you see them

Errors

Errors in the OpenAI shape.

Every error has the same JSON body. OpenAI SDKs raise it with the status code.

Response

HTTP/1.1 402 Payment Required

{
"error": {
"message": "Not enough credit for this request. Lower max_tokens, or get more credit: see the console.",For people. Log it or show it.
"type": "insufficient_quota",The kind of error.
"param": null,The field at fault, when there is one.
"code": "insufficient_quota"Stable. Branch on it.
}
}
HTTP status codes, error codes and what to do
Status Code and what to do
400 invalid_jsonunsupported_parameter Fix the request. A retry gives the same answer.
401 missing_api_keyinvalid_api_keyrevoked_api_key Check the key in the Authorization header.
402 insufficient_quota Get more credit (talk to us), or lower max_tokens.
404 model_not_found Use a model id from /models.
429 rate_limit_exceeded Wait for Retry-After, then retry.
429 no_capacity Every GPU for this model is busy. Retry shortly.
503 no_capacity The wait for a GPU ran out. Retry shortly.
504 timeout Retry, or lower max_tokens.

Headers worth logging

x-request-id
On every response. Quote it when you write to us.
retry-after
On 429 and 503: seconds to wait.
x-pc-node-country
The country of the GPU node that served the request.

Every code, in full: API docs

Contact

Need another open model? Ask us.

We use your name, work email, company and message only to reply and to discuss working with you. We delete the stored enquiry 12 months after your last message, and keep our email thread with you for 2 years after our last reply. Privacy