Docs
Chat Completions
POST/v1/chat/completions
OpenAI-compatible chat completions for every text model in the catalog. The official OpenAI SDKs work with the base URL set to https://api.boostrail.com/v1.
Example
curl https://api.boostrail.com/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4.1-flash",
"messages": [{"role": "user", "content": "Hello"}]
}'from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.boostrail.com/v1")
stream = client.chat.completions.create(
model="deepseek-v4.1-flash",
messages=[{"role": "user", "content": "Write a haiku about trains."}],
max_tokens=200,
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")
if chunk.usage:
print("\n", chunk.usage)import OpenAI from "openai";
const client = new OpenAI({ apiKey: "YOUR_API_KEY", baseURL: "https://api.boostrail.com/v1" });
const completion = await client.chat.completions.create({
model: "deepseek-v4.1-flash",
messages: [{ role: "user", content: "Hello" }],
});
console.log(completion.choices[0].message.content);Request body
| Field | Type | Notes |
|---|---|---|
model | string, required | A model id from the Models page. |
messages | array, required | Roles system, user, assistant and tool. content is a string, an array of parts (text, image_url), or null on assistant turns that only carry tool_calls. Image parts need a model that accepts images. |
max_tokens, max_completion_tokens | integer | Upper limit on output tokens; max_completion_tokens is used when max_tokens is absent. It also sets how much balance is reserved: see Balance and max_tokens. |
stream | boolean | Server-sent events. See Streaming. |
stream_options | object | Accepted. Usage is always sent in the final chunk, whatever this says. |
temperature, top_p, stop, frequency_penalty, presence_penalty, seed | Passed on unchanged. | |
tools, tool_choice, parallel_tool_calls | Passed on unchanged. Tool-call history in messages is supported. | |
response_format | object | Passed on unchanged: JSON mode or json_schema, on models that support them. |
reasoning_effort, thinking, reasoning, enable_thinking | Four ways to switch thinking on or off or set its budget. Send whichever your client uses; it is translated for the route serving the model. | |
logprobs, top_logprobs | Passed on unchanged. | |
user | string | Used as a session id for routing (see Session affinity). Not passed to the model. |
n | integer | Only 1. Any other value returns 400 invalid_request. |
Fields not listed above
audio,web_search_options, andmodalitiesother than["text"]return400 unsupported_parameter;paramnames the field.- Any other field is not passed on. The request still runs, and the
x-boostrail-degradedresponse header names the dropped fields.
Streaming
With stream: true the response is a series of data: lines in the OpenAI chunk format, ending with data: [DONE]. The last chunk before [DONE] carries usage.
If the route serving the request sends nothing within 45 seconds, the request moves to another route before any data reaches you. Once data has been sent, the request stays where it is. Details are on Errors and failover.
Response
A standard chat.completion object. usage has prompt_tokens, completion_tokens and total_tokens, plus prompt_tokens_details.cached_tokens when part of the prompt was read from cache. Cached tokens are billed at the model's cache-read price.
Checks before a request starts
- Unknown model id:
400 model_not_found. - Input too long: a request estimated to exceed the model's input limit returns
400 context_length_exceededbefore any balance is reserved. The message states the limit and the estimate, so clients that compact history on this error keep working. - Non-streaming requests have a 180-second limit and end with
504 provider_timeoutwhen it runs out. Stream long generations instead.
Balance and max_tokens
Each request reserves balance for its input plus its maximum output, then settles at actual usage and releases the rest.
- With
max_tokensset: if the balance cannot cover the input plusmax_tokens, the request returns402 insufficient_balance. - Without
max_tokens: if the balance is short, the output limit is lowered to what the balance covers and the request runs. When even a short reply is not covered, it returns402 insufficient_balance. - With your own provider keys (BYOK) added for the model: when the balance cannot cover the request, it runs on your keys only, without platform fallback, instead of returning
402. If your keys fail, the error message says that only your keys were tried. - A request that fails is not charged; its reservation is released.