AUDIO
Create Voice
Enroll a reusable stored voice for `scx-tts`. Upload a short reference clip once, then pass the returned `voice_id` to HTTP speech or realtime TTS.
Authorization
AuthorizationstringheaderrequiredBearer token — your API key. Example: Bearer sk-...
Stored Voices
Stored voices are the required production path for realtime voice cloning. You upload a clean reference clip once, we enroll it, and you receive a stable voice ID such as qwen-tts-vc-voice_....
Use that ID in:
| Surface | Field |
|---|---|
POST /v1/audio/speech | voice: "<returned voice_id>" |
WS /v1/realtime | session.audio.output.voice: "<returned voice_id>" |
A readable preferred name is recorded when enrolling, but clients must use the exact voice_id returned by this endpoint.
Reference Clip Requirements
- Use 5-15 seconds of clean, single-speaker speech.
- Provide the exact full transcript in
ref_text; partial or mismatched transcripts can produce unstable output. - WAV is recommended. MP3 and M4A/MP4 audio are also accepted.
- Do not send inline
voice_ref_wav_b64to realtime. Realtime uses stored voice IDs.
Request body
multipart/form-dataaudio_sampleanyrequiredReference audio file for the voice. Sent as multipart/form-data. Use 5-15 seconds of clean speech; WAV is recommended.
consentstringrequiredRequired consent/rights acknowledgement for creating and using this synthetic voice.
namestringrequiredHuman-readable label for the stored voice.
ref_textstringrequiredExact transcript of the full reference clip. Required so voice cloning remains stable.
speaker_descriptionstringOptional notes about the speaker/accent/style for operators and dashboards.
Response
voice_idstringrequiredReusable stored voice ID, e.g. `qwen-tts-vc-voice_...`.
namestringformatstringcreated_atnumberhas_ref_textbooleanduration_snumbersample_ratenumbersize_bytesnumberspeaker_descriptionstringRequest
import requests
with open("reference.wav", "rb") as f:
res = requests.post(
"https://api.scx.ai/v1/audio/voices",
headers={"Authorization": "Bearer your-scx-api-key"},
files={"audio_sample": ("reference.wav", f, "audio/wav")},
data={
"consent": "I own or have permission to use this voice.",
"name": "David",
"ref_text": "Exact transcript of the full reference clip.",
"speaker_description": "Clean Australian-accent default voice.",
},
)
res.raise_for_status()
voice_id = res.json()["voice_id"]
print(voice_id)import requests
with open("reference.wav", "rb") as f:
res = requests.post(
"https://api.scx.ai/v1/audio/voices",
headers={"Authorization": "Bearer your-scx-api-key"},
files={"audio_sample": ("reference.wav", f, "audio/wav")},
data={
"consent": "I own or have permission to use this voice.",
"name": "David",
"ref_text": "Exact transcript of the full reference clip.",
"speaker_description": "Clean Australian-accent default voice.",
},
)
res.raise_for_status()
voice_id = res.json()["voice_id"]
print(voice_id)