AUDIO

Create Voice

Enroll a reusable stored voice for `scx-tts`. Upload a short reference clip once, then pass the returned `voice_id` to HTTP speech or realtime TTS.

POST/v1/audio/voices

Authorization

Authorizationstringheaderrequired

Bearer token — your API key. Example: Bearer sk-...

Stored Voices

Stored voices are the required production path for realtime voice cloning. You upload a clean reference clip once, we enroll it, and you receive a stable voice ID such as qwen-tts-vc-voice_....

Use that ID in:

SurfaceField
POST /v1/audio/speechvoice: "<returned voice_id>"
WS /v1/realtimesession.audio.output.voice: "<returned voice_id>"

A readable preferred name is recorded when enrolling, but clients must use the exact voice_id returned by this endpoint.

Reference Clip Requirements

  • Use 5-15 seconds of clean, single-speaker speech.
  • Provide the exact full transcript in ref_text; partial or mismatched transcripts can produce unstable output.
  • WAV is recommended. MP3 and M4A/MP4 audio are also accepted.
  • Do not send inline voice_ref_wav_b64 to realtime. Realtime uses stored voice IDs.

Request body

multipart/form-data
audio_sampleanyrequired

Reference audio file for the voice. Sent as multipart/form-data. Use 5-15 seconds of clean speech; WAV is recommended.

consentstringrequired

Required consent/rights acknowledgement for creating and using this synthetic voice.

namestringrequired

Human-readable label for the stored voice.

ref_textstringrequired

Exact transcript of the full reference clip. Required so voice cloning remains stable.

speaker_descriptionstring

Optional notes about the speaker/accent/style for operators and dashboards.

Response

voice_idstringrequired

Reusable stored voice ID, e.g. `qwen-tts-vc-voice_...`.

namestring
formatstring
created_atnumber
has_ref_textboolean
duration_snumber
sample_ratenumber
size_bytesnumber
speaker_descriptionstring

Request

import requests

with open("reference.wav", "rb") as f:
    res = requests.post(
        "https://api.scx.ai/v1/audio/voices",
        headers={"Authorization": "Bearer your-scx-api-key"},
        files={"audio_sample": ("reference.wav", f, "audio/wav")},
        data={
            "consent": "I own or have permission to use this voice.",
            "name": "David",
            "ref_text": "Exact transcript of the full reference clip.",
            "speaker_description": "Clean Australian-accent default voice.",
        },
    )

res.raise_for_status()
voice_id = res.json()["voice_id"]
print(voice_id)