BYOK is live — 1,000,000 free BYOK requests every month, no top-up requiredLearn more

Docs

Responses

POST/v1/responses

The OpenAI Responses format, served statelessly for every text model in the catalog. Codex CLI and other Responses clients work with the base URL https://api.boostrail.com/v1.

Example

curl https://api.boostrail.com/v1/responses \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4.1-flash",
    "input": "Hello"
  }'
from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.boostrail.com/v1")
response = client.responses.create(
    model="deepseek-v4.1-flash",
    instructions="Answer in one sentence.",
    input="What is a token?",
)
print(response.output_text)

Stateless only

  • Nothing is stored between requests. Send the full conversation in input every time.
  • previous_response_id returns 400 unsupported_parameter.
  • store: true returns 400 unsupported_parameter. Leave store out or set it to false; responses always report store: false.

Request body

FieldTypeNotes
modelstring, requiredA model id from the Models page.
inputstring or array, requiredItems of type message (roles user, assistant, system, developer), function_call and function_call_output. reasoning items are skipped.
instructionsstringSent as the system message.
toolsarrayfunction tools only. Other tool types, such as web_search, are dropped and named in x-boostrail-degraded, so clients that always send them keep working.
tool_choicestring or objectauto, none, required, or {"type": "function", "name": ...}.
max_output_tokensintegerOutput limit. When it is reached, incomplete_details.reason is max_output_tokens.
reasoningobjectreasoning.effort is applied; other keys are ignored.
textobjecttext.format of type json_object or json_schema.
temperature, top_p, parallel_tool_calls, streamPassed on.
userstringUsed as a session id for routing (see Session affinity).

Message content parts: input_text, output_text, and input_image with image_url as a URL string or {"url": ...}. Other part types return 400 invalid_request.

Fields not listed above

Any other top-level field, for example include or prompt_cache_key, is dropped and named in x-boostrail-degraded.

Streaming

With stream: true the events follow the Responses streaming format: response.created, response.in_progress, output item, content part, text delta and function-call argument events, then response.completed, or response.incomplete when max_output_tokens was reached.

Errors

Errors use the OpenAI format and the codes on Errors and failover. Balance, input-length and time limits are the same as on Chat Completions.

Last updated: 2026-10-03