RESPONSES

Create a Response

Create a model response. Provide text or structured input items to generate text or JSON outputs. Supports function calling, web search, code interpreter, multi-turn conversations via previous_response_id, and server-sent event streaming.

POST/v1/responses

Authorization

Authorizationstringheaderrequired

Bearer token — your API key. Example: Bearer sk-...

Request body

application/json
modelstringrequired

The model to use (e.g. gpt-oss-120b, Llama-4-Maverick-17B-128E-Instruct).

inputstring | object[]required

Text or array of input items to generate a response for. Can be a simple string or a structured array of messages, function call outputs, and function calls from previous turns.

instructionsstring

A system (or developer) message inserted at the top of the conversation. Use this to set the model's behavior, persona, or output format.

toolsobject[]

An array of tools the model may call while generating a response. Supports function tools (your custom code), web search, and code interpreter.

tool_choice"auto" | "required" | "none" | object

Controls which tool (if any) the model calls. 'auto' lets the model decide, 'required' forces tool use, 'none' disables tools, or specify a function by name.

streamboolean

Whether to stream the response using server-sent events. When true, events are emitted progressively: response.created, response.output_text.delta, response.completed, etc.

previous_response_idstring

The ID of a previous response to continue from. Creates multi-turn conversations by automatically prepending the conversation history. Cannot be used with conversation.

temperaturenumber

Sampling temperature between 0 and 2. Higher values (e.g. 0.8) produce more random output, lower values (e.g. 0.2) produce more deterministic output. Defaults to 1.

top_pnumber

Nucleus sampling: only consider tokens whose cumulative probability exceeds this value. Recommended to use temperature or top_p, but not both.

max_output_tokensinteger

The maximum number of output tokens to generate. The model will stop once this limit is reached.

metadataobject

Key-value pairs of metadata to attach to the response. Useful for tracking, filtering, or annotating responses.

reasoningobject

Configuration for the model's reasoning behavior. Controls effort and output summarization.

textobject

Configuration for text output, including response format (plain text, JSON, or JSON Schema).

storeboolean

Whether to store the response for later retrieval via GET /v1/responses/{id}. Defaults to true.

Response

idstringrequired

Unique response identifier (e.g. resp_abc123).

object"response"required

Object type — always "response".

created_atintegerrequired

Unix timestamp (seconds) of when the response was created.

status"completed" | "failed" | "incomplete" | "in_progress"required

The status of the response. 'completed': finished successfully, 'failed': an error occurred, 'incomplete': hit token limit or content filter, 'in_progress': currently generating.

completedfailedincompletein_progress
outputobject[]required

Array of output items generated by the model. Can include assistant messages and function calls.

output_textstringrequired

Convenience field: all output_text content parts concatenated into a single string.

modelstringrequired

The model that generated the response.

usageobjectrequired
errorobjectrequired

Error details if the response failed, otherwise null.

temperaturenumberrequired

The temperature used for this response.

top_pnumberrequired

The top_p value used for this response.

max_output_tokensintegerrequired

The max output tokens limit used for this response.

previous_response_idstringrequired

The ID of the previous response this continues from, if any.

instructionsstringrequired

The system instructions used for this response.

toolsobject[]required

The tools that were available for this response.

metadataobjectrequired

Metadata attached to this response.

Request

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: "your-scx-api-key",
  baseURL: "https://api.scx.ai/v1",
});

const response = await client.responses.create({
  model: "gpt-oss-120b",
  input: "Explain quantum computing in simple terms.",
});

console.log(response.output_text);

Response

{
  "id": "resp_abc123def456",
  "object": "response",
  "created_at": 1709251200,
  "status": "completed",
  "output": [
    {
      "type": "message",
      "id": "msg_abc123",
      "role": "assistant",
      "status": "completed",
      "content": [
        {
          "type": "output_text",
          "text": "Quantum computing uses quantum bits (qubits) that can exist in multiple states simultaneously, unlike classical bits which are either 0 or 1. This allows quantum computers to process certain types of problems exponentially faster than classical computers.",
          "annotations": []
        }
      ]
    }
  ],
  "output_text": "Quantum computing uses quantum bits (qubits) that can exist in multiple states simultaneously, unlike classical bits which are either 0 or 1. This allows quantum computers to process certain types of problems exponentially faster than classical computers.",
  "model": "gpt-oss-120b",
  "usage": {
    "input_tokens": 18,
    "output_tokens": 42,
    "total_tokens": 60
  },
  "error": null,
  "temperature": 1,
  "top_p": 1,
  "max_output_tokens": null,
  "previous_response_id": null,
  "instructions": null,
  "tools": [],
  "metadata": {}
}