BYOK is live — 1,000,000 free BYOK requests every month, no top-up requiredLearn more

Docs

Chat Completions

POST/v1/chat/completions

OpenAI-compatible chat completions for every text model in the catalog. The official OpenAI SDKs work with the base URL set to https://api.boostrail.com/v1.

Example

curl https://api.boostrail.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4.1-flash",
    "messages": [{"role": "user", "content": "Hello"}]
  }'
from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.boostrail.com/v1")
stream = client.chat.completions.create(
    model="deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "Write a haiku about trains."}],
    max_tokens=200,
    stream=True,
)
for chunk in stream:
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="")
    if chunk.usage:
        print("\n", chunk.usage)
import OpenAI from "openai";

const client = new OpenAI({ apiKey: "YOUR_API_KEY", baseURL: "https://api.boostrail.com/v1" });
const completion = await client.chat.completions.create({
  model: "deepseek-v4.1-flash",
  messages: [{ role: "user", content: "Hello" }],
});
console.log(completion.choices[0].message.content);

Request body

FieldTypeNotes
modelstring, requiredA model id from the Models page.
messagesarray, requiredRoles system, user, assistant and tool. content is a string, an array of parts (text, image_url), or null on assistant turns that only carry tool_calls. Image parts need a model that accepts images.
max_tokens, max_completion_tokensintegerUpper limit on output tokens; max_completion_tokens is used when max_tokens is absent. It also sets how much balance is reserved: see Balance and max_tokens.
streambooleanServer-sent events. See Streaming.
stream_optionsobjectAccepted. Usage is always sent in the final chunk, whatever this says.
temperature, top_p, stop, frequency_penalty, presence_penalty, seedPassed on unchanged.
tools, tool_choice, parallel_tool_callsPassed on unchanged. Tool-call history in messages is supported.
response_formatobjectPassed on unchanged: JSON mode or json_schema, on models that support them.
reasoning_effort, thinking, reasoning, enable_thinkingFour ways to switch thinking on or off or set its budget. Send whichever your client uses; it is translated for the route serving the model.
logprobs, top_logprobsPassed on unchanged.
userstringUsed as a session id for routing (see Session affinity). Not passed to the model.
nintegerOnly 1. Any other value returns 400 invalid_request.

Fields not listed above

  • audio, web_search_options, and modalities other than ["text"] return 400 unsupported_parameter; param names the field.
  • Any other field is not passed on. The request still runs, and the x-boostrail-degraded response header names the dropped fields.

Streaming

With stream: true the response is a series of data: lines in the OpenAI chunk format, ending with data: [DONE]. The last chunk before [DONE] carries usage.

If the route serving the request sends nothing within 45 seconds, the request moves to another route before any data reaches you. Once data has been sent, the request stays where it is. Details are on Errors and failover.

Response

A standard chat.completion object. usage has prompt_tokens, completion_tokens and total_tokens, plus prompt_tokens_details.cached_tokens when part of the prompt was read from cache. Cached tokens are billed at the model's cache-read price.

Checks before a request starts

  • Unknown model id: 400 model_not_found.
  • Input too long: a request estimated to exceed the model's input limit returns 400 context_length_exceeded before any balance is reserved. The message states the limit and the estimate, so clients that compact history on this error keep working.
  • Non-streaming requests have a 180-second limit and end with 504 provider_timeout when it runs out. Stream long generations instead.

Balance and max_tokens

Each request reserves balance for its input plus its maximum output, then settles at actual usage and releases the rest.

  • With max_tokens set: if the balance cannot cover the input plus max_tokens, the request returns 402 insufficient_balance.
  • Without max_tokens: if the balance is short, the output limit is lowered to what the balance covers and the request runs. When even a short reply is not covered, it returns 402 insufficient_balance.
  • With your own provider keys (BYOK) added for the model: when the balance cannot cover the request, it runs on your keys only, without platform fallback, instead of returning 402. If your keys fail, the error message says that only your keys were tried.
  • A request that fails is not charged; its reservation is released.

Last updated: 2026-10-03