REALTIME

Realtime Audio

OpenAI Realtime-compatible WebSocket endpoint for low-latency bidirectional audio. Supports streaming speech-to-text (with optional server-side voice activity detection), text-to-speech with multiple wire formats, and STT→TTS roundtrips on a single connection.

WS/v1/realtime

Authorization

The realtime upgrade accepts your API key via either an HTTP header or a WebSocket subprotocol. Use whichever fits your client.

Authorizationheaderrecommended

Bearer token — your API key. Example: Bearer sk-.... Only works from non-browser clients (Node, Python, curl); the WebSocket API in browsers doesn't expose headers.

openai-insecure-api-key.<key>subprotocol

Pass the API key as a WebSocket subprotocol alongside realtime. This is the OpenAI SDK pattern for browser usage. Example: new WebSocket(url, ["realtime", "openai-insecure-api-key.sk-..."]).

Connection & handshake

Open a WebSocket to wss://<host>/v1/realtime. Authentication has two paths depending on whether the client can set HTTP headers:

ClientAuth path
Node.js, Python, curl, server-to-serverAuthorization: Bearer <key> request header.
Browser WebSocket (which can't send custom headers)The realtime subprotocol plus one of openai-insecure-api-key.<key> or bearer.<key> as a second subprotocol.

Immediately on open, the server sends a session.created event carrying a session ID and default config. Send a session.update next to configure the session — audio formats, transcription model, output modalities, TTS model. The server replies with session.updated — this is your signal that the effective config is active and it is safe to start streaming.

Clients should not rely on any audio or response event sent before the first session.updated; wait for the ack before starting real work. Binary WebSocket frames are rejected — all events are JSON text frames.

Session configuration

One session can do both directions on a single socket:

ConfiguresFieldDirection
Speech-to-textaudio.input.format, audio.input.transcription, audio.input.turn_detectionMic → server.
Text-to-speechaudio.output.format, output_modalities, tts_modelServer → speaker.

You don't have to use both. A transcription-only client just sets audio.input. A TTS-only client sets output_modalities: ['audio'] and (optionally) audio.output.format.

session.update is idempotent and merges into the existing session — fields you don't send keep their prior value, so you can change audio.input.transcription.language mid-stream without re-sending the whole config.

If a field is rejected (unsupported_feature for an unsupported codec, invalid_request_error for a malformed value), the server emits an error event and that field stays at its prior value. Other fields in the same update are still applied.

Lifecycle of a transcription utterance

Manual commit (turn_detection: null)

  1. Send one or more input_audio_buffer.append events carrying base64-encoded audio chunks (~85–250 ms each is typical). No server ack on success.
  2. Send input_audio_buffer.commit. Server replies immediately with input_audio_buffer.committed (carrying an item_id), then begins transcribing.
  3. Server streams conversation.item.input_audio_transcription.delta events as partial text arrives.
  4. Server emits a terminal ...transcription.completed with the canonical transcript and a usage block — or ...failed if transcription errored.

Server-side VAD (turn_detection: { type: 'server_vad' })

The server detects utterance boundaries from the audio itself:

  • input_audio_buffer.speech_started fires when speech crosses the threshold.
  • input_audio_buffer.speech_stopped fires after silence_duration_ms of trailing silence.
  • The server then runs the same committed → deltas → completed flow above, without you sending commit yourself.
KnobRangeDefaultUse
threshold0–1tuned for clean room audioHigher = stricter ("must really be speech"); lower = more sensitive.
prefix_padding_ms0–5000 ms300 msHow much pre-speech audio to retain so the transcriber sees the leading edge.
silence_duration_ms0–5000 ms500 msHow long after the speaker pauses before the server declares end-of-turn.

Note: server_vad requires audio/pcm input — G.711 codecs are rejected since the VAD needs linear samples.

Buffers are hard-capped at 25 MiB between commits. Appends that would exceed the cap produce error with code buffer_overflow — commit or clear before appending more.

Lifecycle of a TTS response

  1. Send session.update with output_modalities: ['audio'], optionally audio.output.format (defaults to PCM 24 kHz) and tts_model (defaults to scx-tts). Wait for session.updated.

  2. Send conversation.item.create with a user message containing the text to synthesise:

    Server acks with conversation.item.created.

  3. Send response.create. Server emits the preamble: response.created → response.output_item.added → response.content_part.added.

  4. Audio streams as response.audio.delta events, base64-encoded in the configured format. Begin playback as soon as the first delta arrives.

  5. Text-form transcript events (response.audio_transcript.delta → response.audio_transcript.done) and structural close events (response.audio.done, response.content_part.done, response.output_item.done).

  6. Terminal response.done with status: 'completed' | 'cancelled' | 'failed' and a usage block when completed.

Long inputs are split at sentence boundaries and synthesised so audio length scales with input length — the client sees one logical response envelope, no manual chunking required.

Cancellation. Send response.cancel at any point during steps 3–5. The server stops mid-stream and emits a final response.done with status: 'cancelled'.

Concurrency. At most one response is in flight per connection. A response.create arriving while another is active returns error with code response_in_progress.

Audio formats

The same codec enum applies to input (audio.input.format) and output (audio.output.format):

typeCodecSample ratePer-sample size
audio/pcmLinear PCM, signed 16-bit LE, mono8 000–48 000 Hz2 bytes
audio/pcmuG.711 μ-law (North American telephony)8 000 Hz fixed1 byte
audio/pcmaG.711 A-law (European telephony)8 000 Hz fixed1 byte

The server resamples internally so any rate the underlying model produces is delivered in the requested format. PCM 24 kHz is a good output default for browser AudioContext playback; choose 16 kHz to halve bandwidth, 8 kHz μ-law/A-law for telephony bridges.

Requesting G.711 at any rate other than 8000 Hz returns invalid_request_error.

Voice and TTS model selection

tts_model defaults to scx-tts — the canonical SCX text-to-speech model. Configure TTS model selection via session.update; per-response overrides on response.create are reserved and not honoured today.

audio.output.voice accepts the exact stored voice_id returned by POST /v1/audio/voices, for example qwen-tts-vc-voice_.... Omit audio.output.voice to use the configured David default voice. Preset voice names such as ito or alice-bennett are not an SCX voice catalog.

Realtime TTS uses either the configured David default or a stored voice_id. It does not accept inline voice_ref_wav_b64, ref_text, or x_vector_only_mode fields.

Error reference

Returned in the error.code field on error events:

CodeWhen
invalid_request_errorMalformed event, missing required field, type mismatch.
unsupported_featureFeature is in the wire spec but this deployment doesn't implement it (e.g. semantic_vad, server-side noise_reduction, realtime inline voice cloning).
buffer_overflowinput_audio_buffer.append would exceed the 25 MiB cap. Commit or clear first.
response_in_progressresponse.create arrived while another response is streaming. Send response.cancel first or wait for response.done.
upstream_errorTranscription or synthesis failed mid-call. Non-fatal — the session stays open.
missing_api_key / invalid_api_keySurfaced on the upgrade response (HTTP 401), not as an error event.
internal_errorServer bug. Please report.

error events do not tear down the session unless severe — clients may continue sending events after an error.

Features not yet supported

Requesting any of these returns an error event with code unsupported_feature — the session stays open and you can retry with a valid config:

FeatureReasonWorkaround
audio.input.turn_detection: { type: 'semantic_vad' }Semantic VAD is not implemented in this deployment.Use server_vad or manual commits.
audio.input.noise_reduction (any non-null value)Server-side DSP is not provided.Apply client-side: getUserMedia({ noiseSuppression: true }).
Inline zero-shot voice cloning fields on realtime (voice_ref_wav_b64, ref_text, x_vector_only_mode)Realtime TTS uses stored voice IDs, not per-response reference clips.Enroll a stored voice with POST /v1/audio/voices and pass the returned voice_id as audio.output.voice in session.update.
conversation.item.delete and rolling-history retrievalThe session keeps a small in-memory window of recent items but does not expose list/get/delete events.Each TTS request reads only the most recent user-text item; nothing is persisted across connections.
Per-response overrides on response.create (response.modalities, response.model)Parsed but not yet honoured.Configure modalities and the TTS model via session.update instead.

Usage & billing

OperationEndpoint labelInput tokensOutput tokens
Transcription commit/v1/audio/transcriptionsAudio duration.Transcribed text length.
Translation commit (session.type: 'translation')/v1/audio/translationsAudio duration.English transcript length.
TTS response (response.create)/v1/audio/speechInput character count.Synthesised audio duration.

Each commit and each response.create counts as one API call. The internal sentence-chunking of long TTS inputs is invisible to billing.

Usage events appear in your dashboard with the appropriate endpoint label so realtime traffic is grouped alongside the corresponding HTTP traffic. The HTTP transcription endpoint is file/SSE-oriented; low-latency microphone streaming should use this WebSocket endpoint.

Client events

JSON over WebSocket
session.update"session.update"

Configure or re-configure the session. Send this immediately after `session.created` and wait for `session.updated` before sending audio or creating responses. May be sent again to mutate config (e.g. switch language, swap output format) mid-session.

input_audio_buffer.append"input_audio_buffer.append"

Append audio bytes to the current commit buffer. The server sends no ack on success — only an `error` on failure (invalid base64, buffer overflow).

input_audio_buffer.commit"input_audio_buffer.commit"

Flush the accumulated audio for transcription. Commits are serialised per-connection — a second commit waits for the first's `conversation.item.input_audio_transcription.completed` (or `.failed`) before starting. The server replies with `input_audio_buffer.committed` immediately, then streams `delta` events, then a terminal `completed` or `failed`. Under `turn_detection: server_vad` you don't normally send this — the server auto-commits.

input_audio_buffer.clear"input_audio_buffer.clear"

Drop the accumulated audio without transcribing. Useful after a false start. Server acks with `input_audio_buffer.cleared`.

conversation.item.create"conversation.item.create"

Append a message to the conversation buffer. The session retains a rolling window of the last 64 items; older items roll off as new ones arrive. `response.create` reads the most recent `user`-role text item to drive synthesis, so a typical TTS request is `conversation.item.create` (with `role: 'user'`, `content: [{ type: 'input_text', text }]`) followed immediately by `response.create`.

response.create"response.create"

Trigger TTS synthesis on the most recent user-text item in the conversation. Errors with `invalid_request_error` if `output_modalities` doesn't include `'audio'`, if no user text is available, or with `response_in_progress` if a response is already streaming. Successful invocations stream a `response.created` → `response.audio.delta`* → `response.done` envelope.

response.cancel"response.cancel"

Abort the in-flight response. The server stops streaming audio deltas and emits a terminal `response.done` with `status: 'cancelled'`. Returns `invalid_request_error` ("No response in progress to cancel.") if no response is active.

Server events

session.created"session.created"

First event emitted after the WebSocket opens, before any client event. Carries the session ID and default config. No client action required — but you should send `session.update` next.

session.updated"session.updated"

Acknowledgement of a successful `session.update`. Carries the merged effective config. Safe to start sending audio (or `conversation.item.create` + `response.create` for TTS) once this arrives.

input_audio_buffer.committed"input_audio_buffer.committed"

Immediate ack of `input_audio_buffer.commit` (or auto-commit under server_vad). Sent before transcription begins — track progress via `conversation.item.input_audio_transcription.delta` / `.completed`.

input_audio_buffer.cleared"input_audio_buffer.cleared"

Ack of `input_audio_buffer.clear`.

input_audio_buffer.speech_started"input_audio_buffer.speech_started"

Emitted by `turn_detection: server_vad` when speech first crosses the threshold for `min_speech_duration_ms`. Use it to drive UI affordances like a 'listening' indicator.

input_audio_buffer.speech_stopped"input_audio_buffer.speech_stopped"

Emitted by `turn_detection: server_vad` after `silence_duration_ms` of trailing silence. The server auto-commits the buffered audio immediately after — no `input_audio_buffer.commit` needed.

conversation.item.created"conversation.item.created"

Acknowledgement of `conversation.item.create`. Mirrors the submitted item, with a server-assigned `id` if the client didn't supply one.

conversation.item.input_audio_transcription.delta"conversation.item.input_audio_transcription.delta"

Partial transcript chunk streamed as the transcript is produced. Typically one-to-several words per delta. Emission rate is bounded by the model; don't assume a fixed cadence.

conversation.item.input_audio_transcription.completed"conversation.item.input_audio_transcription.completed"

Terminal event for a successful commit. After this, the session is ready for the next commit — send more `input_audio_buffer.append` chunks followed by another `commit` (or let server_vad auto-commit).

conversation.item.input_audio_transcription.failed"conversation.item.input_audio_transcription.failed"

Terminal event for a commit that failed transcription (model error, invalid audio, etc.). Non-fatal to the session — the next commit can proceed normally.

response.created"response.created"

First event emitted after `response.create`. Carries the response ID propagated through every subsequent `response.*` event in this cycle.

response.output_item.added"response.output_item.added"

Structural event signalling the assistant message item has been added to the response. Most clients can ignore this and rely on `response.audio.delta` instead.

response.content_part.added"response.content_part.added"

Structural event signalling the audio content part started. Pairs with `response.content_part.done` at the end of the response.

response.audio.delta"response.audio.delta"

Streamed audio chunk. Begin playback as soon as the first delta arrives — don't wait for `response.audio.done`.

response.audio.done"response.audio.done"

Terminal audio event. After this, no more `response.audio.delta` events for this response.

response.audio_transcript.delta"response.audio_transcript.delta"

Streamed text-form transcript of the synthesised audio. Useful for rendering captions alongside playback.

response.audio_transcript.done"response.audio_transcript.done"

Terminal transcript event.

response.content_part.done"response.content_part.done"

Structural event signalling the audio content part is complete.

response.output_item.done"response.output_item.done"

Structural event carrying the completed assistant message item. The same shape appears nested in `response.done.response.output[]`.

response.done"response.done"

Terminal envelope for a TTS response. After this, the connection is ready for another `conversation.item.create` + `response.create` cycle. At most one response is in flight per connection at any time — concurrent `response.create` events return `response_in_progress` until the prior `response.done` fires.

error"error"

Protocol-level error — does NOT tear down the session. The client may continue sending events. Hard auth failures are surfaced on the upgrade response (HTTP 401), not as `error` events.

Request

// Browser — TTS-only via /v1/realtime. Connect, configure output modalities,
// add a user message, fire response.create, stream audio.delta back.

const ws = new WebSocket("wss://api.scx.ai/v1/realtime", [
  "realtime",
  "openai-insecure-api-key.your-scx-api-key",
]);
ws.binaryType = "arraybuffer";

ws.onopen = () => {
  ws.send(JSON.stringify({
    type: "session.update",
    session: {
      output_modalities: ["audio"],
      tts_model: "scx-tts",
      audio: {
        output: {
          format: { type: "audio/pcm", rate: 24000 },
          // Optional: use a stored voice created with POST /v1/audio/voices.
          // voice: "qwen-tts-vc-voice_...",
        },
      },
    },
  }));
};

const ctx = new AudioContext({ sampleRate: 24000 });
let playhead = ctx.currentTime;

ws.onmessage = (msg) => {
  const ev = JSON.parse(msg.data);
  switch (ev.type) {
    case "session.updated":
      // Trigger one TTS request.
      ws.send(JSON.stringify({
        type: "conversation.item.create",
        item: {
          role: "user",
          content: [{ type: "input_text", text: "Hello from the realtime endpoint." }],
        },
      }));
      ws.send(JSON.stringify({ type: "response.create" }));
      return;

    case "response.audio.delta": {
      // Decode base64 PCM16 LE → Int16 → Float32 → schedule for playback.
      const bytes = Uint8Array.fromBase64(ev.delta);
      const i16 = new Int16Array(bytes.buffer, bytes.byteOffset, bytes.byteLength / 2);
      const f32 = new Float32Array(i16.length);
      for (let i = 0; i < i16.length; i++) f32[i] = i16[i] / 32768;

      const buf = ctx.createBuffer(1, f32.length, 24000);
      buf.copyToChannel(f32, 0);
      const src = ctx.createBufferSource();
      src.buffer = buf;
      src.connect(ctx.destination);
      src.start(Math.max(playhead, ctx.currentTime));
      playhead = Math.max(playhead, ctx.currentTime) + buf.duration;
      return;
    }

    case "response.done":
      console.log("response", ev.response.status, "usage:", ev.response.usage);
      ws.close(1000, "done");
      return;

    case "error":
      console.error("realtime error:", ev.error);
      return;
  }
};

Response

{
  "type": "session.created",
  "event_id": "event_7f4c2a1b9e6d8f0a4c2b5e1d",
  "session": {
    "id": "sess_9a8b7c6d5e4f3a2b1c0d9e8f",
    "object": "realtime.session",
    "type": "transcription",
    "audio": {
      "input": {
        "format": {
          "type": "audio/pcm",
          "rate": 24000
        },
        "transcription": {
          "model": "scx-stt"
        },
        "turn_detection": null,
        "noise_reduction": null
      },
      "output": {
        "format": {
          "type": "audio/pcm",
          "rate": 24000
        }
      }
    },
    "output_modalities": [],
    "tts_model": "scx-tts"
  }
}