AUDIO
Audio Translations
Translate non-English audio into English text. Supports streaming via `stream=true`, emitting the same `transcript.text.delta` / `transcript.text.done` server-sent event shape as the transcriptions endpoint. Streaming is an SCX extension — OpenAI does not define it for translations, so the official OpenAI SDK will not parse the stream; consume the SSE directly instead.
Authorization
AuthorizationstringheaderrequiredBearer token — your API key. Example: Bearer sk-...
Streaming output (SCX extension)
OpenAI does not define streaming for the translations endpoint. SCX adds it as a parity feature with /v1/audio/transcriptions — when stream=true you get the same server-sent event shape:
| Event | Frame shape | When |
|---|---|---|
transcript.text.delta | { type, delta } | One per partial chunk. Concatenate all delta values to reconstruct the running English translation. |
transcript.text.done | { type, text, usage } | Terminal frame with the canonical English translation. |
[DONE] | sentinel | Closes the stream. |
Because OpenAI's official SDK doesn't define streaming for translations, the SDK will not parse the stream natively. Consume the text/event-stream body directly with fetch + ReadableStream (see the JavaScript example) or httpx in Python.
Streaming is only available with response_format=json.
When to use /v1/realtime instead
This endpoint takes a complete audio file and returns the English translation. It's the right choice when:
- The audio already exists on disk (recordings, batch pipelines).
- Total recording length is under the 25 MB / ~90 min cap.
For live cross-language captions or voice agents that translate as the speaker talks, use the realtime WebSocket endpoint at /v1/realtime with session.type: "translation" instead. It accepts streaming audio chunks, supports server-side voice activity detection, and returns translation deltas with sub-second first-token latency.
Request body
multipart/form-datafileanyrequiredThe audio file to translate to English, in one of these formats: FLAC, MP3, MP4, MPEG, MPGA, M4A, Ogg, WAV, or WebM. File size limit is 25 MB.
modelstringrequiredThe model ID to use. Example: `Whisper-Large-v3`.
promptstringlanguage"en" | "zh" | "de" | "es" | "ru" | "ko" | "fr" | "ja" | "pt" | "tr" | "pl" | "ca" | "nl" | "ar" | "sv" | "it" | "id" | "hi" | "fi" | "vi" | "he" | "uk" | "el" | "ms" | "cs" | "ro" | "da" | "hu" | "ta" | "no" | "th" | "ur" | "hr" | "bg" | "lt" | "la" | "mi" | "ml" | "cy" | "sk" | "te" | "fa" | "lv" | "bn" | "sr" | "az" | "sl" | "kn" | "et" | "mk" | "br" | "eu" | "is" | "hy" | "ne" | "mn" | "bs" | "kk" | "sq" | "sw" | "gl" | "mr" | "pa" | "si" | "km" | "sn" | "yo" | "so" | "af" | "oc" | "ka" | "be" | "tg" | "sd" | "gu" | "am" | "yi" | "lo" | "uz" | "fo" | "ht" | "ps" | "tk" | "nn" | "mt" | "sa" | "lb" | "my" | "bo" | "tl" | "mg" | "as" | "tt" | "haw" | "ln" | "ha" | "ba" | "jw" | "su" | "yue"response_format"json" | "text" | "srt" | "verbose_json" | "vtt"default: "json"Output format. `json` (default) returns `{ text }`. `text` returns the translation as a plain string. `srt` and `vtt` return subtitle files. `verbose_json` adds segment timestamps. Note: when `stream=true`, only `json` is supported.
jsontextsrtverbose_jsonvtttemperaturenumberstreambooleandefault: falseIf true, the response is streamed as a `text/event-stream` of server-sent events using the same event shape as `/v1/audio/transcriptions`: a sequence of `transcript.text.delta` events followed by a terminal `transcript.text.done` event with `{ text, usage }`, then `data: [DONE]`. Note: OpenAI does not define streaming for the translations endpoint — this is an SCX extension, so the official OpenAI SDK cannot parse it natively; consume the SSE directly.
Response
textstringrequiredThe English translation of the audio. Returned for non-streaming requests.
Request
import fs from "node:fs";
import OpenAI from "openai";
const client = new OpenAI({
apiKey: "your-scx-api-key",
baseURL: "https://api.scx.ai/v1",
});
// Non-streaming: returns the full English translation JSON
const response = await client.audio.translations.create({
model: "Whisper-Large-v3",
file: fs.createReadStream("audio.mp3"),
});
console.log(response.text);
// Streaming: SCX-extension SSE (OpenAI SDK won't parse this for you,
// so fall back to fetch + ReadableStream manually).
const form = new FormData();
form.set("model", "Whisper-Large-v3");
form.set("stream", "true");
form.set("file", fs.createReadStream("audio.mp3"));
const res = await fetch("https://api.scx.ai/v1/audio/translations", {
method: "POST",
headers: { Authorization: "Bearer your-scx-api-key" },
body: form,
});
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
while (true) {
const { value, done } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const parts = buffer.split("\n\n");
buffer = parts.pop() ?? "";
for (const block of parts) {
const line = block.trim();
if (!line.startsWith("data:")) continue;
const data = line.slice(5).trim();
if (data === "[DONE]") continue;
const event = JSON.parse(data);
if (event.type === "transcript.text.delta") {
process.stdout.write(event.delta);
} else if (event.type === "transcript.text.done") {
console.log("\nfinal:", event.text);
}
}
}import fs from "node:fs";
import OpenAI from "openai";
const client = new OpenAI({
apiKey: "your-scx-api-key",
baseURL: "https://api.scx.ai/v1",
});
// Non-streaming: returns the full English translation JSON
const response = await client.audio.translations.create({
model: "Whisper-Large-v3",
file: fs.createReadStream("audio.mp3"),
});
console.log(response.text);
// Streaming: SCX-extension SSE (OpenAI SDK won't parse this for you,
// so fall back to fetch + ReadableStream manually).
const form = new FormData();
form.set("model", "Whisper-Large-v3");
form.set("stream", "true");
form.set("file", fs.createReadStream("audio.mp3"));
const res = await fetch("https://api.scx.ai/v1/audio/translations", {
method: "POST",
headers: { Authorization: "Bearer your-scx-api-key" },
body: form,
});
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
while (true) {
const { value, done } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const parts = buffer.split("\n\n");
buffer = parts.pop() ?? "";
for (const block of parts) {
const line = block.trim();
if (!line.startsWith("data:")) continue;
const data = line.slice(5).trim();
if (data === "[DONE]") continue;
const event = JSON.parse(data);
if (event.type === "transcript.text.delta") {
process.stdout.write(event.delta);
} else if (event.type === "transcript.text.done") {
console.log("\nfinal:", event.text);
}
}
}Response
{
"text": "Hello. This is an audio transcription system test. The fast brown fox jumps over the lazy dog."
}{
"text": "Hello. This is an audio transcription system test. The fast brown fox jumps over the lazy dog."
}