Speech

The Audio API provides a speech endpoint for converting text into lifelike spoken audio. Three models are available: scx-tts (the canonical SCX text-to-speech model with a configured default voice and reusable stored voices), and tts-1 / tts-1-hd (preset voice catalogs with six built-in accents). All three share the same wire format, so you can switch between them by changing the model field.

Quick start

The speech endpoint takes three key inputs: the model, the text to convert to audio, and the voice to use. A simple request looks like this:

from openai import OpenAI
from pathlib import Path

client = OpenAI(
    base_url="https://api.scx.ai/v1",
    api_key="your-scx-api-key",
)

response = client.audio.speech.create(
    model="tts-1",
    voice="ito",
    input="Hello! Welcome to SCX.ai text-to-speech.",
)

Path("output.mp3").write_bytes(response.content)

By default, the endpoint outputs an MP3 file of the spoken audio, but it can be configured to output other supported formats.

Choosing a model

ModelVoice mechanismWhen to use
scx-ttsConfigured default voice, or stored Qwen-backed voice IDs.Use the default voice for production TTS, or enroll a reusable speaker with POST /v1/audio/voices. Pick this when the preset voices don't match your brand or character.
tts-1Preset voice catalog — pick by name or UUID.Latency‑tuned. Right for chat/agent UX where speed matters and any of the six preset accents fit.
tts-1-hdSame preset catalog as tts-1.Quality‑tuned. Slightly slower, marginally cleaner output for offline / podcast / narration use cases.

The three models share the same request body, response stream, and error shape — you swap between them by changing only the model field.

Preset voices (tts-1, tts-1-hd)

Pick a preset by passing its name in the voice field. These apply to tts-1 and tts-1-hd only. For scx-tts, omit voice to use the configured default or pass the stored voice_id returned by POST /v1/audio/voices.

VoiceDescription
australian-samMale, Australian accent - natural and friendly
friendly-kiwiMale, New Zealand accent - casual and approachable
likeable-aussieFemale, Australian/NZ accent - likeable and pleasing
itoMale, American accent - warm and conversational
serene-assistantFemale, American accent - calm and professional
alice-bennettFemale, British accent - professional and articulate

Stored voices with scx-tts

scx-tts uses the configured default voice when voice is omitted. To use another reusable speaker, enroll a short reference clip once with POST /v1/audio/voices, then pass the returned voice_id in voice for HTTP speech or session.audio.output.voice for realtime TTS.

Realtime TTS does not accept inline voice_ref_wav_b64, ref_text, or x_vector_only_mode fields. If you need a custom realtime voice, create the stored voice first.

Default reference voice (no cloning)

If you just want to use scx-tts without supplying your own voice, omit voice entirely. The endpoint uses the configured default voice:

from openai import OpenAI
from pathlib import Path

client = OpenAI(
    base_url="https://api.scx.ai/v1",
    api_key="your-scx-api-key",
)

response = client.audio.speech.create(
    model="scx-tts",
    input="Hello! This is the SCX text-to-speech model.",
    response_format="wav",
)

Path("output.wav").write_bytes(response.content)

Creating a stored voice

Upload a clean reference clip and its exact transcript to POST /v1/audio/voices. SCX enrolls the speaker through Qwen's DashScope customization API and returns a reusable provider voice ID such as qwen-tts-vc-voice_....

import requests
from pathlib import Path

with open("reference-voice.wav", "rb") as f:
    voice_res = requests.post(
        "https://api.scx.ai/v1/audio/voices",
        headers={"Authorization": "Bearer your-scx-api-key"},
        files={"audio_sample": ("reference-voice.wav", f, "audio/wav")},
        data={
            "consent": "I own or have permission to use this voice.",
            "name": "David",
            "ref_text": "Exact transcript of the full reference clip.",
        },
    )

voice_res.raise_for_status()
voice_id = voice_res.json()["voice_id"]

response = requests.post(
    "https://api.scx.ai/v1/audio/speech",
    headers={"Authorization": "Bearer your-scx-api-key"},
    json={
        "model": "scx-tts",
        "voice": voice_id,
        "input": "This sentence will be spoken in the stored voice.",
        "response_format": "wav",
    },
)
response.raise_for_status()
Path("stored-voice.wav").write_bytes(response.content)

Reference clip tips

  • Clip length: use 5-15 seconds of clean, single-speaker speech.
  • Transcript: provide the exact full transcript in ref_text; partial or mismatched transcripts can produce unstable output.
  • Clean recordings beat long ones: a 6-second clip recorded in a quiet room with a decent mic is better than a 30-second clip with HVAC hum or a phone speaker.
  • Container format: WAV is recommended. Qwen also accepts MP3 and M4A/MP4 audio.

Streaming raw PCM for low‑latency playback

scx-tts supports the same streaming behaviour as the preset models. Combined with response_format: pcm, you get raw 16‑bit signed LE samples at 24 kHz mono — perfect for piping straight into a playback pipeline without any container parsing.

python
import requests, sounddevice as sd, numpy as np

resp = requests.post(
    "https://api.scx.ai/v1/audio/speech",
    headers={"Authorization": "Bearer your-scx-api-key"},
    json={
        "model": "scx-tts",
        "input": "Streaming PCM straight to the speaker.",
        "response_format": "pcm",
    },
    stream=True,
)
resp.raise_for_status()

# Raw 16-bit LE PCM at 24 kHz mono.
with sd.OutputStream(samplerate=24000, channels=1, dtype="int16") as out:
    for chunk in resp.iter_content(chunk_size=4096):
        if chunk:
            out.write(np.frombuffer(chunk, dtype="<i2"))

Supported output formats

The default response format is WAV. Format availability depends on the model: scx-tts produces wav and pcm (requesting mp3 returns invalid_request_error), while tts-1 and tts-1-hd additionally offer mp3.

FormatModelsDescription
wavallDefault. Lossless 16-bit PCM in a WAV container.
pcmallRaw 16-bit signed LE samples, mono, 24 kHz. Lowest latency — streams as it synthesises.
mp3tts-1, tts-1-hdLossy, smallest payload. Not available on scx-tts.

Request parameters

ParameterTypeDescriptionDefault
modelStringThe model to use: scx-tts, tts-1, or tts-1-hd.Required
inputStringThe text to generate audio for. Maximum length is 5000 characters.Required
voiceStringPreset voice name for tts-1 / tts-1-hd, or the stored voice_id returned by POST /v1/audio/voices for scx-tts. Omit for the scx-tts default.-
response_formatStringThe audio format: wav or pcm (mp3 additionally for tts-1 / tts-1-hd).wav
speedNumberThe speed of the generated audio. Range: 0.25 to 4.0. Treated as a model hint; scx-tts currently synthesises at its natural rate.1.0

Retired scx-tts inline fields

Inline voice cloning is no longer supported on any scx-tts surface. Sending voice_ref_wav_b64, ref_text, or x_vector_only_mode returns invalid_request_error. Enroll the reference clip once with POST /v1/audio/voices and pass the returned voice ID in voice instead. max_new_tokens is accepted for backwards compatibility but has no effect.

Streaming real-time audio

The Speech API supports real-time audio streaming using chunk transfer encoding. This means audio can begin playing before the full file is generated.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.scx.ai/v1",
    api_key="your-scx-api-key",
)

response = client.audio.speech.create(
    model="tts-1",
    voice="serene-assistant",
    input="This is a streaming test. The audio will start playing before generation completes.",
)

# Stream to a file
with open("stream_output.mp3", "wb") as f:
    for chunk in response.iter_bytes():
        f.write(chunk)

Adjusting speed

You can adjust the speed of the generated audio by setting the speed parameter. Values range from 0.25 (slowest) to 4.0 (fastest), with 1.0 being the default normal speed.

# Slower speech (0.75x speed)
response = client.audio.speech.create(
    model="tts-1",
    voice="alice-bennett",
    input="This will be spoken more slowly.",
    speed=0.75,
)

# Faster speech (1.5x speed)
response = client.audio.speech.create(
    model="tts-1",
    voice="alice-bennett",
    input="This will be spoken more quickly.",
    speed=1.5,
)

Limitations

  • Maximum input text length is 5000 characters per request.
  • For tts-1 and tts-1-hd, split text into chunks and generate audio for each segment if you need to go beyond the per‑request limit.
  • For scx-tts, the model auto‑splits long inputs at sentence boundaries internally and stitches the result into a single streamed response — you don't need to chunk client‑side under the 5000‑char cap. Sentences over ~400 characters fall back to clause‑level splits, and anything still longer is hard‑wrapped under the upstream's 600‑character per‑call cap.
  • Realtime custom voices must be enrolled first with POST /v1/audio/voices; inline reference clips are rejected on WS /v1/realtime.