AUDIO

Audio Translations

Translate non-English audio into English text. Supports streaming via `stream=true`, emitting the same `transcript.text.delta` / `transcript.text.done` server-sent event shape as the transcriptions endpoint. Streaming is an SCX extension — OpenAI does not define it for translations, so the official OpenAI SDK will not parse the stream; consume the SSE directly instead.

POST/v1/audio/translations

Authorization

Authorizationstringheaderrequired

Bearer token — your API key. Example: Bearer sk-...

Streaming output (SCX extension)

OpenAI does not define streaming for the translations endpoint. SCX adds it as a parity feature with /v1/audio/transcriptions — when stream=true you get the same server-sent event shape:

EventFrame shapeWhen
transcript.text.delta{ type, delta }One per partial chunk. Concatenate all delta values to reconstruct the running English translation.
transcript.text.done{ type, text, usage }Terminal frame with the canonical English translation.
[DONE]sentinelCloses the stream.

Because OpenAI's official SDK doesn't define streaming for translations, the SDK will not parse the stream natively. Consume the text/event-stream body directly with fetch + ReadableStream (see the JavaScript example) or httpx in Python.

Streaming is only available with response_format=json.

When to use /v1/realtime instead

This endpoint takes a complete audio file and returns the English translation. It's the right choice when:

  • The audio already exists on disk (recordings, batch pipelines).
  • Total recording length is under the 25 MB / ~90 min cap.

For live cross-language captions or voice agents that translate as the speaker talks, use the realtime WebSocket endpoint at /v1/realtime with session.type: "translation" instead. It accepts streaming audio chunks, supports server-side voice activity detection, and returns translation deltas with sub-second first-token latency.

Request body

multipart/form-data
fileanyrequired

The audio file to translate to English, in one of these formats: FLAC, MP3, MP4, MPEG, MPGA, M4A, Ogg, WAV, or WebM. File size limit is 25 MB.

modelstringrequired

The model ID to use. Example: `Whisper-Large-v3`.

promptstring
language"en" | "zh" | "de" | "es" | "ru" | "ko" | "fr" | "ja" | "pt" | "tr" | "pl" | "ca" | "nl" | "ar" | "sv" | "it" | "id" | "hi" | "fi" | "vi" | "he" | "uk" | "el" | "ms" | "cs" | "ro" | "da" | "hu" | "ta" | "no" | "th" | "ur" | "hr" | "bg" | "lt" | "la" | "mi" | "ml" | "cy" | "sk" | "te" | "fa" | "lv" | "bn" | "sr" | "az" | "sl" | "kn" | "et" | "mk" | "br" | "eu" | "is" | "hy" | "ne" | "mn" | "bs" | "kk" | "sq" | "sw" | "gl" | "mr" | "pa" | "si" | "km" | "sn" | "yo" | "so" | "af" | "oc" | "ka" | "be" | "tg" | "sd" | "gu" | "am" | "yi" | "lo" | "uz" | "fo" | "ht" | "ps" | "tk" | "nn" | "mt" | "sa" | "lb" | "my" | "bo" | "tl" | "mg" | "as" | "tt" | "haw" | "ln" | "ha" | "ba" | "jw" | "su" | "yue"
response_format"json" | "text" | "srt" | "verbose_json" | "vtt"default: "json"

Output format. `json` (default) returns `{ text }`. `text` returns the translation as a plain string. `srt` and `vtt` return subtitle files. `verbose_json` adds segment timestamps. Note: when `stream=true`, only `json` is supported.

jsontextsrtverbose_jsonvtt
temperaturenumber
streambooleandefault: false

If true, the response is streamed as a `text/event-stream` of server-sent events using the same event shape as `/v1/audio/transcriptions`: a sequence of `transcript.text.delta` events followed by a terminal `transcript.text.done` event with `{ text, usage }`, then `data: [DONE]`. Note: OpenAI does not define streaming for the translations endpoint — this is an SCX extension, so the official OpenAI SDK cannot parse it natively; consume the SSE directly.

Response

textstringrequired

The English translation of the audio. Returned for non-streaming requests.

Request

import fs from "node:fs";
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: "your-scx-api-key",
  baseURL: "https://api.scx.ai/v1",
});

// Non-streaming: returns the full English translation JSON
const response = await client.audio.translations.create({
  model: "Whisper-Large-v3",
  file: fs.createReadStream("audio.mp3"),
});
console.log(response.text);

// Streaming: SCX-extension SSE (OpenAI SDK won't parse this for you,
// so fall back to fetch + ReadableStream manually).
const form = new FormData();
form.set("model", "Whisper-Large-v3");
form.set("stream", "true");
form.set("file", fs.createReadStream("audio.mp3"));

const res = await fetch("https://api.scx.ai/v1/audio/translations", {
  method: "POST",
  headers: { Authorization: "Bearer your-scx-api-key" },
  body: form,
});

const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = "";

while (true) {
  const { value, done } = await reader.read();
  if (done) break;
  buffer += decoder.decode(value, { stream: true });
  const parts = buffer.split("\n\n");
  buffer = parts.pop() ?? "";
  for (const block of parts) {
    const line = block.trim();
    if (!line.startsWith("data:")) continue;
    const data = line.slice(5).trim();
    if (data === "[DONE]") continue;
    const event = JSON.parse(data);
    if (event.type === "transcript.text.delta") {
      process.stdout.write(event.delta);
    } else if (event.type === "transcript.text.done") {
      console.log("\nfinal:", event.text);
    }
  }
}

Response

{
  "text": "Hello. This is an audio transcription system test. The fast brown fox jumps over the lazy dog."
}