Docs
Responses
POST/v1/responses
The OpenAI Responses format, served statelessly for every text model in the catalog. Codex CLI and other Responses clients work with the base URL https://api.boostrail.com/v1.
Example
curl https://api.boostrail.com/v1/responses \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4.1-flash",
"input": "Hello"
}'from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.boostrail.com/v1")
response = client.responses.create(
model="deepseek-v4.1-flash",
instructions="Answer in one sentence.",
input="What is a token?",
)
print(response.output_text)Stateless only
- Nothing is stored between requests. Send the full conversation in
inputevery time. previous_response_idreturns400 unsupported_parameter.store: truereturns400 unsupported_parameter. Leavestoreout or set it tofalse; responses always reportstore: false.
Request body
| Field | Type | Notes |
|---|---|---|
model | string, required | A model id from the Models page. |
input | string or array, required | Items of type message (roles user, assistant, system, developer), function_call and function_call_output. reasoning items are skipped. |
instructions | string | Sent as the system message. |
tools | array | function tools only. Other tool types, such as web_search, are dropped and named in x-boostrail-degraded, so clients that always send them keep working. |
tool_choice | string or object | auto, none, required, or {"type": "function", "name": ...}. |
max_output_tokens | integer | Output limit. When it is reached, incomplete_details.reason is max_output_tokens. |
reasoning | object | reasoning.effort is applied; other keys are ignored. |
text | object | text.format of type json_object or json_schema. |
temperature, top_p, parallel_tool_calls, stream | Passed on. | |
user | string | Used as a session id for routing (see Session affinity). |
Message content parts: input_text, output_text, and input_image with image_url as a URL string or {"url": ...}. Other part types return 400 invalid_request.
Fields not listed above
Any other top-level field, for example include or prompt_cache_key, is dropped and named in x-boostrail-degraded.
Streaming
With stream: true the events follow the Responses streaming format: response.created, response.in_progress, output item, content part, text delta and function-call argument events, then response.completed, or response.incomplete when max_output_tokens was reached.
Errors
Errors use the OpenAI format and the codes on Errors and failover. Balance, input-length and time limits are the same as on Chat Completions.