开发文档
Responses
POST/v1/responses
以无状态方式提供 OpenAI Responses 格式,目录里的所有文本模型都能用。Codex CLI 等 Responses 客户端把 base URL 设为 https://api.boostrail.com/v1 即可使用。
示例
curl https://api.boostrail.com/v1/responses \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4.1-flash",
"input": "你好"
}'from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.boostrail.com/v1")
response = client.responses.create(
model="deepseek-v4.1-flash",
instructions="用一句话回答。",
input="什么是 token?",
)
print(response.output_text)只支持无状态
- 请求之间不保存任何内容,每次都要在
input里带上完整对话。 previous_response_id返回400 unsupported_parameter。store: true返回400 unsupported_parameter。不传store或设为false即可;响应里始终是store: false。
请求体
| 字段 | 类型 | 说明 |
|---|---|---|
model | string,必填 | 模型页 上的模型 ID。 |
input | string 或 array,必填 | 支持 message(角色 user、assistant、system、developer)、function_call 和 function_call_output 三类项。reasoning 项会被跳过。 |
instructions | string | 作为 system 消息发送。 |
tools | array | 只支持 function 工具。其他类型(例如 web_search)会被丢掉,并列在 x-boostrail-degraded 里,所以总是带着这些工具的客户端也能照常工作。 |
tool_choice | string 或 object | auto、none、required,或 {"type": "function", "name": ...}。 |
max_output_tokens | integer | 输出上限。达到上限时,incomplete_details.reason 为 max_output_tokens。 |
reasoning | object | reasoning.effort 会生效,其他键忽略。 |
text | object | text.format 支持 json_object 或 json_schema。 |
temperature, top_p, parallel_tool_calls, stream | 原样传给模型。 | |
user | string | 用作路由的会话 ID(见 会话保持)。 |
消息内容分段支持 input_text、output_text,以及 input_image,其中 image_url 可以是 URL 字符串或 {"url": ...}。其他分段类型返回 400 invalid_request。
上表以外的字段
其他顶层字段(例如 include 或 prompt_cache_key)会被丢掉,并列在 x-boostrail-degraded 里。
流式输出
设置 stream: true 后,事件遵循 Responses 的流式格式:先是 response.created、response.in_progress,然后是输出项、内容分段、文本增量和函数调用参数事件,最后是 response.completed;达到 max_output_tokens 时以 response.incomplete 结束。
报错
报错使用 OpenAI 格式,错误码见 报错与换线。余额、输入长度和时限的规则与 Chat Completions 相同。