Speech
The Audio API provides a speech endpoint for converting text into lifelike spoken audio. Three models are available: scx-tts (the canonical SCX text-to-speech model with a configured default voice and reusable stored voices), and tts-1 / tts-1-hd (preset voice catalogs with six built-in accents). All three share the same wire format, so you can switch between them by changing the model field.
Quick start
The speech endpoint takes three key inputs: the model, the text to convert to audio, and the voice to use. A simple request looks like this:
from openai import OpenAI
from pathlib import Path
client = OpenAI(
base_url="https://api.scx.ai/v1",
api_key="your-scx-api-key",
)
response = client.audio.speech.create(
model="tts-1",
voice="ito",
input="Hello! Welcome to SCX.ai text-to-speech.",
)
Path("output.mp3").write_bytes(response.content)
from openai import OpenAI
from pathlib import Path
client = OpenAI(
base_url="https://api.scx.ai/v1",
api_key="your-scx-api-key",
)
response = client.audio.speech.create(
model="tts-1",
voice="ito",
input="Hello! Welcome to SCX.ai text-to-speech.",
)
Path("output.mp3").write_bytes(response.content)
By default, the endpoint outputs an MP3 file of the spoken audio, but it can be configured to output other supported formats.
Choosing a model
| Model | Voice mechanism | When to use |
|---|---|---|
scx-tts | Configured default voice, or stored Qwen-backed voice IDs. | Use the default voice for production TTS, or enroll a reusable speaker with POST /v1/audio/voices. Pick this when the preset voices don't match your brand or character. |
tts-1 | Preset voice catalog — pick by name or UUID. | Latency‑tuned. Right for chat/agent UX where speed matters and any of the six preset accents fit. |
tts-1-hd | Same preset catalog as tts-1. | Quality‑tuned. Slightly slower, marginally cleaner output for offline / podcast / narration use cases. |
The three models share the same request body, response stream, and error shape — you swap between them by changing only the model field.
Preset voices (tts-1, tts-1-hd)
Pick a preset by passing its name in the voice field. These apply to tts-1 and tts-1-hd only. For scx-tts, omit voice to use the configured default or pass the stored voice_id returned by POST /v1/audio/voices.
| Voice | Description |
|---|---|
australian-sam | Male, Australian accent - natural and friendly |
friendly-kiwi | Male, New Zealand accent - casual and approachable |
likeable-aussie | Female, Australian/NZ accent - likeable and pleasing |
ito | Male, American accent - warm and conversational |
serene-assistant | Female, American accent - calm and professional |
alice-bennett | Female, British accent - professional and articulate |
Stored voices with scx-tts
scx-tts uses the configured default voice when voice is omitted. To use another reusable speaker, enroll a short reference clip once with POST /v1/audio/voices, then pass the returned voice_id in voice for HTTP speech or session.audio.output.voice for realtime TTS.
Realtime TTS does not accept inline voice_ref_wav_b64, ref_text, or x_vector_only_mode fields. If you need a custom realtime voice, create the stored voice first.
Default reference voice (no cloning)
If you just want to use scx-tts without supplying your own voice, omit voice entirely. The endpoint uses the configured default voice:
from openai import OpenAI
from pathlib import Path
client = OpenAI(
base_url="https://api.scx.ai/v1",
api_key="your-scx-api-key",
)
response = client.audio.speech.create(
model="scx-tts",
input="Hello! This is the SCX text-to-speech model.",
response_format="wav",
)
Path("output.wav").write_bytes(response.content)
from openai import OpenAI
from pathlib import Path
client = OpenAI(
base_url="https://api.scx.ai/v1",
api_key="your-scx-api-key",
)
response = client.audio.speech.create(
model="scx-tts",
input="Hello! This is the SCX text-to-speech model.",
response_format="wav",
)
Path("output.wav").write_bytes(response.content)
Creating a stored voice
Upload a clean reference clip and its exact transcript to POST /v1/audio/voices. SCX enrolls the speaker through Qwen's DashScope customization API and returns a reusable provider voice ID such as qwen-tts-vc-voice_....
import requests
from pathlib import Path
with open("reference-voice.wav", "rb") as f:
voice_res = requests.post(
"https://api.scx.ai/v1/audio/voices",
headers={"Authorization": "Bearer your-scx-api-key"},
files={"audio_sample": ("reference-voice.wav", f, "audio/wav")},
data={
"consent": "I own or have permission to use this voice.",
"name": "David",
"ref_text": "Exact transcript of the full reference clip.",
},
)
voice_res.raise_for_status()
voice_id = voice_res.json()["voice_id"]
response = requests.post(
"https://api.scx.ai/v1/audio/speech",
headers={"Authorization": "Bearer your-scx-api-key"},
json={
"model": "scx-tts",
"voice": voice_id,
"input": "This sentence will be spoken in the stored voice.",
"response_format": "wav",
},
)
response.raise_for_status()
Path("stored-voice.wav").write_bytes(response.content)
import requests
from pathlib import Path
with open("reference-voice.wav", "rb") as f:
voice_res = requests.post(
"https://api.scx.ai/v1/audio/voices",
headers={"Authorization": "Bearer your-scx-api-key"},
files={"audio_sample": ("reference-voice.wav", f, "audio/wav")},
data={
"consent": "I own or have permission to use this voice.",
"name": "David",
"ref_text": "Exact transcript of the full reference clip.",
},
)
voice_res.raise_for_status()
voice_id = voice_res.json()["voice_id"]
response = requests.post(
"https://api.scx.ai/v1/audio/speech",
headers={"Authorization": "Bearer your-scx-api-key"},
json={
"model": "scx-tts",
"voice": voice_id,
"input": "This sentence will be spoken in the stored voice.",
"response_format": "wav",
},
)
response.raise_for_status()
Path("stored-voice.wav").write_bytes(response.content)
Reference clip tips
- Clip length: use 5-15 seconds of clean, single-speaker speech.
- Transcript: provide the exact full transcript in
ref_text; partial or mismatched transcripts can produce unstable output. - Clean recordings beat long ones: a 6-second clip recorded in a quiet room with a decent mic is better than a 30-second clip with HVAC hum or a phone speaker.
- Container format: WAV is recommended. Qwen also accepts MP3 and M4A/MP4 audio.
Streaming raw PCM for low‑latency playback
scx-tts supports the same streaming behaviour as the preset models. Combined with response_format: pcm, you get raw 16‑bit signed LE samples at 24 kHz mono — perfect for piping straight into a playback pipeline without any container parsing.
import requests, sounddevice as sd, numpy as np
resp = requests.post(
"https://api.scx.ai/v1/audio/speech",
headers={"Authorization": "Bearer your-scx-api-key"},
json={
"model": "scx-tts",
"input": "Streaming PCM straight to the speaker.",
"response_format": "pcm",
},
stream=True,
)
resp.raise_for_status()
# Raw 16-bit LE PCM at 24 kHz mono.
with sd.OutputStream(samplerate=24000, channels=1, dtype="int16") as out:
for chunk in resp.iter_content(chunk_size=4096):
if chunk:
out.write(np.frombuffer(chunk, dtype="<i2"))
import requests, sounddevice as sd, numpy as np
resp = requests.post(
"https://api.scx.ai/v1/audio/speech",
headers={"Authorization": "Bearer your-scx-api-key"},
json={
"model": "scx-tts",
"input": "Streaming PCM straight to the speaker.",
"response_format": "pcm",
},
stream=True,
)
resp.raise_for_status()
# Raw 16-bit LE PCM at 24 kHz mono.
with sd.OutputStream(samplerate=24000, channels=1, dtype="int16") as out:
for chunk in resp.iter_content(chunk_size=4096):
if chunk:
out.write(np.frombuffer(chunk, dtype="<i2"))
Supported output formats
The default response format is WAV. Format availability depends on the model: scx-tts produces wav and pcm (requesting mp3 returns invalid_request_error), while tts-1 and tts-1-hd additionally offer mp3.
| Format | Models | Description |
|---|---|---|
wav | all | Default. Lossless 16-bit PCM in a WAV container. |
pcm | all | Raw 16-bit signed LE samples, mono, 24 kHz. Lowest latency — streams as it synthesises. |
mp3 | tts-1, tts-1-hd | Lossy, smallest payload. Not available on scx-tts. |
Request parameters
| Parameter | Type | Description | Default |
|---|---|---|---|
model | String | The model to use: scx-tts, tts-1, or tts-1-hd. | Required |
input | String | The text to generate audio for. Maximum length is 5000 characters. | Required |
voice | String | Preset voice name for tts-1 / tts-1-hd, or the stored voice_id returned by POST /v1/audio/voices for scx-tts. Omit for the scx-tts default. | - |
response_format | String | The audio format: wav or pcm (mp3 additionally for tts-1 / tts-1-hd). | wav |
speed | Number | The speed of the generated audio. Range: 0.25 to 4.0. Treated as a model hint; scx-tts currently synthesises at its natural rate. | 1.0 |
Retired scx-tts inline fields
Inline voice cloning is no longer supported on any scx-tts surface. Sending voice_ref_wav_b64, ref_text, or x_vector_only_mode returns invalid_request_error. Enroll the reference clip once with POST /v1/audio/voices and pass the returned voice ID in voice instead. max_new_tokens is accepted for backwards compatibility but has no effect.
Streaming real-time audio
The Speech API supports real-time audio streaming using chunk transfer encoding. This means audio can begin playing before the full file is generated.
from openai import OpenAI
client = OpenAI(
base_url="https://api.scx.ai/v1",
api_key="your-scx-api-key",
)
response = client.audio.speech.create(
model="tts-1",
voice="serene-assistant",
input="This is a streaming test. The audio will start playing before generation completes.",
)
# Stream to a file
with open("stream_output.mp3", "wb") as f:
for chunk in response.iter_bytes():
f.write(chunk)
from openai import OpenAI
client = OpenAI(
base_url="https://api.scx.ai/v1",
api_key="your-scx-api-key",
)
response = client.audio.speech.create(
model="tts-1",
voice="serene-assistant",
input="This is a streaming test. The audio will start playing before generation completes.",
)
# Stream to a file
with open("stream_output.mp3", "wb") as f:
for chunk in response.iter_bytes():
f.write(chunk)
Adjusting speed
You can adjust the speed of the generated audio by setting the speed parameter. Values range from 0.25 (slowest) to 4.0 (fastest), with 1.0 being the default normal speed.
# Slower speech (0.75x speed)
response = client.audio.speech.create(
model="tts-1",
voice="alice-bennett",
input="This will be spoken more slowly.",
speed=0.75,
)
# Faster speech (1.5x speed)
response = client.audio.speech.create(
model="tts-1",
voice="alice-bennett",
input="This will be spoken more quickly.",
speed=1.5,
)
# Slower speech (0.75x speed)
response = client.audio.speech.create(
model="tts-1",
voice="alice-bennett",
input="This will be spoken more slowly.",
speed=0.75,
)
# Faster speech (1.5x speed)
response = client.audio.speech.create(
model="tts-1",
voice="alice-bennett",
input="This will be spoken more quickly.",
speed=1.5,
)
Limitations
- Maximum input text length is 5000 characters per request.
- For
tts-1andtts-1-hd, split text into chunks and generate audio for each segment if you need to go beyond the per‑request limit. - For
scx-tts, the model auto‑splits long inputs at sentence boundaries internally and stitches the result into a single streamed response — you don't need to chunk client‑side under the 5000‑char cap. Sentences over ~400 characters fall back to clause‑level splits, and anything still longer is hard‑wrapped under the upstream's 600‑character per‑call cap. - Realtime custom voices must be enrolled first with
POST /v1/audio/voices; inline reference clips are rejected onWS /v1/realtime.