AUDIO
Create Speech
Generate audio from text with `scx-tts`. Supports the configured default voice and stored voices; `pcm` streams as it synthesises.
Authorization
AuthorizationstringheaderrequiredBearer token — your API key. Example: Bearer sk-...
Model
| Model | Voice mechanism | When to use |
|---|---|---|
scx-tts | Configured default voice, or stored voice IDs. | Use the default voice for production TTS, or enroll a reusable speaker with POST /v1/audio/voices. |
The public SCX speech surface is scx-tts. Other model aliases may exist internally, but client-facing integrations should target scx-tts.
Voice selection
Omit voice to use the configured default voice.
For reusable cloned voices, first enroll the reference clip with POST /v1/audio/voices, then pass the returned stored voice_id in voice. This is the recommended production path because clients upload the reference clip once and reuse the ID across HTTP speech and realtime TTS.
Preset voice names such as ito or alice-bennett are not an SCX voice catalog.
Inline request-time clone fields (voice_ref_wav_b64, ref_text, x_vector_only_mode) are no longer accepted on any surface and return invalid_request_error. Enroll a stored voice first.
Stored voices and realtime
For realtime TTS, create a stored voice with POST /v1/audio/voices, then pass the returned voice_id in session.audio.output.voice on WS /v1/realtime.
The realtime WebSocket endpoint does not accept inline voice_ref_wav_b64, ref_text, or x_vector_only_mode in session.update / response.create; attempting that returns unsupported_feature.
Reference clip tips
- Use 5-15 seconds of clean, single-speaker speech.
- Provide the exact full transcript in
ref_textwhen enrolling the voice. - WAV is recommended; MP3 and M4A/MP4 audio are also accepted.
Output formats and streaming
response_format | Content-Type | Use case |
|---|---|---|
wav (default) | audio/wav | Lossless 16-bit PCM in a WAV container. Easy to play with any decoder; ~44 bytes of header overhead. |
pcm | audio/pcm | Raw signed 16-bit little-endian PCM, mono, 24 kHz. No container framing — feed bytes directly into an AudioContext, FFmpeg pipeline, or sample-rate converter. Best for low-latency playback. |
mp3 is not available on scx-tts and returns invalid_request_error. Encode client-side from wav if you need it.
pcm streams as it synthesises — first bytes arrive while later text is still being spoken. wav needs the total length for its container header, so it is delivered once synthesis completes. Set response_format explicitly in production so clients do not depend on defaults.
Speed
The speed parameter is accepted by scx-tts and can shorten output at faster values. Treat it as a model hint rather than a precise linear multiplier: test the exact values you expose in production and avoid promising exact duration scaling.
Long inputs
input is hard-capped at 5000 characters per request — anything longer returns invalid_request_error before any synthesis happens.
Long input strings are split at sentence boundaries internally and stitched back into one response, so audio length scales linearly with input length — you don't need to chunk client-side.
Sentences over ~400 characters fall back to clause-level splits (,;:), and anything still over the 600-character per-pass cap is split at the nearest word boundary.
Most callers can ignore max_new_tokens. Tune it down if you want to cap synthesis latency for unbounded user inputs; tune it up only if you genuinely need single-call audio longer than ~39 s and your input has no natural sentence boundaries to chunk on.
Usage & billing
Each request counts as one TTS call regardless of how many internal chunks the synthesis pipeline emits. input_tokens is estimated from input character count, output_tokens from synthesised audio duration. Usage is recorded with the endpoint: /v1/audio/speech label so HTTP-mode and realtime-mode TTS traffic are grouped together in your dashboard.
Stored voice enrollment is handled by POST /v1/audio/voices; synthesis calls are billed as TTS calls.
Request body
multipart/form-datamodelstringrequiredTTS model ID. Use **`scx-tts`** for SCX text-to-speech. Use the configured default voice or a stored `voice_id` from `POST /v1/audio/voices`.
inputstringThe text to synthesise. Maximum 5000 characters. The server splits long inputs at sentence boundaries internally so audio length scales with input length — you don't need to chunk client-side.
voicestringOptional stored voice ID returned by `POST /v1/audio/voices`, e.g. `qwen-tts-vc-voice_...`. Omit this field for the configured default voice. Preset voice names such as `ito` are not an SCX voice catalog.
response_format"wav" | "pcm"default: "wav"Audio container/codec. `wav` is the default and is lossless 16-bit PCM in a WAV container. `pcm` = raw signed 16-bit LE PCM, mono, 24 kHz — feed directly into a low-latency playback pipeline, and the lowest-latency choice because it streams as it synthesises. `mp3` is not available on `scx-tts` and returns `invalid_request_error`.
wavpcmspeednumberdefault: 1Speech speed hint between 0.25 and 4.0. Validated for range but not currently applied by `scx-tts`, which synthesises at its natural rate. Resample client-side if you need a specific tempo.
voice_ref_wav_b64string**No longer accepted.** Returns `invalid_request_error`. Enroll the reference clip once with `POST /v1/audio/voices` and pass the returned voice id in `voice` instead.
voice_ref_wav_format"mp3" | "wav" | "pcm"**No longer accepted.** Returns `invalid_request_error` alongside `voice_ref_wav_b64`.
mp3wavpcmref_textstring**No longer accepted.** Returns `invalid_request_error`. Supply the transcript when enrolling a stored voice instead.
x_vector_only_modeboolean**No longer accepted.** Returns `invalid_request_error`. Stored voices are the only cloning path.
max_new_tokensinteger**Ignored.** Accepted for backwards compatibility and no longer has any effect — there is no per-call audio-token cap. Long inputs are chunked internally regardless.
Request
// Default scx-tts voice — recommended for most use cases.
import OpenAI from "openai";
import fs from "fs";
const client = new OpenAI({
apiKey: "your-scx-api-key",
baseURL: "https://api.scx.ai/v1",
});
const response = await client.audio.speech.create({
model: "scx-tts",
input: "Hello! This is the SCX text-to-speech model.",
response_format: "wav",
});
fs.writeFileSync("output.wav", Buffer.from(await response.arrayBuffer()));// Default scx-tts voice — recommended for most use cases.
import OpenAI from "openai";
import fs from "fs";
const client = new OpenAI({
apiKey: "your-scx-api-key",
baseURL: "https://api.scx.ai/v1",
});
const response = await client.audio.speech.create({
model: "scx-tts",
input: "Hello! This is the SCX text-to-speech model.",
response_format: "wav",
});
fs.writeFileSync("output.wav", Buffer.from(await response.arrayBuffer()));Response
{
"note": "Streamed binary audio in the requested format.",
"contentType": {
"wav": "audio/wav",
"pcm": "audio/pcm"
}
}{
"note": "Streamed binary audio in the requested format.",
"contentType": {
"wav": "audio/wav",
"pcm": "audio/pcm"
}
}