VIDEOS

Create a Video

Start an asynchronous video generation job. Returns a video object immediately; poll it until the status is completed, then download the MP4 from the content endpoint.

POST/v1/videos

Authorization

Authorizationstringheaderrequired

Bearer token — your API key. Example: Bearer sk-...

Lifecycle

Video generation is asynchronous. Creating a video returns a video object in the queued state; the job then moves through in_progress to either completed or failed. Short clips typically finish in well under a minute. Poll Retrieve video until the status is completed, then download the MP4 from Download video content — or subscribe a webhook endpoint to the video.completed / video.failed events and skip polling entirely.

Models

ModelDurationSizes (WIDTHxHEIGHT)AudioImage-to-videoReferences
seedance-1.0-pro-fast2–12s864x480, 736x544, 1248x704, 1280x720*, 704x1248, 720x1280*, 960x960, 1920x1088, 1920x1080*, 1088x1920, 1080x1920*—✓—
seedance-1.0-pro2–12s864x480, 736x544, 1248x704, 1280x720*, 704x1248, 720x1280*, 960x960, 1920x1088, 1920x1080*, 1088x1920, 1080x1920*—✓—
seedance-1.5-pro4–12s or "auto"864x496, 1280x720, 720x1280, 960x960, 1920x1080, 1080x1920✓✓—
seedance-2.0-mini4–15s or "auto"864x496, 1280x720, 720x1280, 960x960✓✓15 (9 img, 3 vid, 3 aud)
seedance-2.0-fast4–15s or "auto"864x496, 1280x720, 720x1280, 960x960✓✓15 (9 img, 3 vid, 3 aud)
seedance-2.04–15s or "auto"864x496, 1280x720, 720x1280, 960x960, 1920x1080, 1080x1920, 3840x2160✓✓15 (9 img, 3 vid, 3 aud)
seedance-2.54–30s or "auto"854x480, 1280x720, 720x1280, 960x960, 1920x1080, 1080x1920✓✓50 (30 img, 10 vid, 10 aud)
wan2.2-t2v-plus5–5s1920x1080, 1080x1920, 1440x1440, 1632x1248, 1248x1632, 832x480, 480x832, 624x624, 720x1280*———
wan2.2-i2v-plus5–5s832x480, 1920x1080, 720x1280*—✓—
wan2.2-i2v-flash5–5s832x480, 1280x720, 720x1280*—✓—
happyhorse-1.1-t2v3–15s1280x720, 720x1280, 1920x1080, 1080x1920———
happyhorse-1.1-i2v3–15s1280x720, 1920x1080, 720x1280*—✓—

Each cell lists every accepted size value; "auto" is additionally accepted on every model. Values marked * are convenience aliases that render at the model's nearest native shape (a 1280x720 request on the 1.0 series renders 1248x704) — the video object always reports the actual pixels once completed. All models render 24fps MP4 (H.264; 4k output is 10-bit HEVC).

Request format

Send multipart/form-data (what the OpenAI SDK does, and required for file uploads) or plain application/json — both are accepted. JSON native types are normalized: a numeric seconds: 5 becomes the string form "5".

Image-to-video

Pass input_reference to animate a still image as the clip's first frame. Three forms are accepted: a multipart file upload (OpenAI-SDK compatible), the id of a file uploaded via the Files API, or a public https image URL. Reference images are kept alongside the job.

Reference inputs (omni reference)

Models with reference support accept a content list of mixed reference assets — images, video clips and audio tracks — that guide the generation: character appearance and style from images, motion and camera work from videos, timbre and music from audio.

ModelImagesVideosAudioPer-clipTotal videoPer-audiotask_type
seedance-2.0-mini9332–15s15s2–15sprompt-driven
seedance-2.0-fast9332–15s15s2–15sprompt-driven
seedance-2.09332–15s15s2–15sprompt-driven
seedance-2.53010102–30s30s2–30s✓

Address assets in your prompt positionally per type, in content order — "Image 1", "Video 2", "Audio 1" ("the character from Image 1 dances with the camera moves of Video 1"). Each part carries either the file_id of a file uploaded via the Files API or a public https url; in multipart bodies, send content as a JSON string field. content and input_reference are mutually exclusive — omni references and first-frame image-to-video are distinct modes (to pin a reference image as the first or last frame in omni mode, say so in the prompt).

Accepted formats and caps: images as for input_reference (30 MB each); videos MP4/QuickTime, H.264/H.265 with AAC/MP3 audio, 24–60fps, 300–6000 px sides (200 MB each); audio WAV/MP3 (15 MB each). Requests whose only references are audio need a model with audio-only support (seedance-2.5).

On models with the task_type column checked, pass task_type to have the request validated for a specific intent — "reference" (new video guided by the assets), "edit" (change elements of a reference video) or "extend" (continue a reference video). edit/extend require at least one video reference and size: "auto" (the output follows the input clip); edit also requires seconds: "auto". On other reference models the task type is inferred from the prompt alone.

Jobs whose references include video bill the whole generation at the model's with-video token rate (the token count also grows with the input clips' duration). Reference assets are kept alongside the job, and a file referenced by a queued or running job cannot be deleted until the job finishes.

Billing and limits

Video is billed per video token — proportional to pixels × duration (roughly width × height × 24fps × seconds ÷ 1024). Your usage history itemizes each job into per-meter line items: the base video charge, an audio surcharge on models that generate sound (seedance-1.5-pro), and reference-image lines. When you create a job, its estimated maximum cost is reserved against your credit balance and the final charge settles from the actual rendered output; failed generations are never charged. Creating a job requires enough spendable credit to cover the reservation (402 otherwise), and each organization can run a limited number of concurrent video jobs (429 when at the limit — wait for a job to finish or delete a queued one).

Prompts are passed to the model verbatim. Note that the model service also interprets legacy inline directives of the form --parameter value inside prompt text; the strongly-typed request parameters above always take precedence, and billing always follows what was actually rendered.

Request body

multipart/form-data
modelstringrequired

The video model to use (e.g. "seedance-1.0-pro"). See the model table below for capabilities.

promptstringrequired

Text description of the video to generate (1–32,000 characters). Passed to the model verbatim.

secondsstring

Clip duration as an integer string within the model's supported range (see the model table), or "auto" on models that support it to let the model pick. Defaults to "4".

sizestring

Output size as WIDTHxHEIGHT from the model's supported list (see the model table), or "auto" — accepted on every model — to let the model match the input image / pick a shape. Defaults to "720x1280".

input_referencestring | any

Reference image for image-to-video: a multipart file upload, the ID of a file uploaded via the Files API (`file-…`), or an https image URL. JPEG/PNG/WebP/BMP/TIFF/GIF (HEIC/HEIF on newer models), up to 30 MB. The image becomes the first frame. Mutually exclusive with `content`.

contentobject[]

Omni-reference asset list (models with reference support only — see the model table). Mixed images, videos and audio; each part is a Files-API id or an https URL (upload binaries via the Files API first). Order matters: prompts address assets positionally per type. Mutually exclusive with `input_reference`.

task_type"auto" | "reference" | "edit" | "extend"

How the model should treat the reference assets (models with the `task_type` column checked only; elsewhere the task type is inferred from the prompt). "edit" and "extend" require at least one video reference and `size: "auto"`; "edit" also requires `seconds: "auto"`. Defaults to "auto".

autoreferenceeditextend
audioboolean

Generate a synchronized audio track on models that support it (defaults to true). Silent models ignore this. On seedance-1.5-pro, audio is billed as a surcharge on top of the base video rate — omit it (audio: false) to pay the base rate only.

Response

idstringrequired

The video ID (e.g. "video_abc123").

object"video"required

Object type — always "video".

modelstringrequired

The model used for the generation.

status"queued" | "in_progress" | "completed" | "failed"required

The current status of the generation job.

queuedin_progresscompletedfailed
progressintegerrequired

Estimated completion percentage. 0 while queued, an elapsed-time estimate (capped at 95) while in progress, 100 on completion.

created_atintegerrequired

Unix timestamp of when the job was created.

completed_atintegerrequired

Unix timestamp of when the job reached a terminal state.

expires_atintegerrequired

Unix timestamp of when the stored video expires. Currently always null — generated videos are retained until you delete them.

promptstringrequired

The prompt used for the generation.

sizestringrequired

Output size as WIDTHxHEIGHT. Shows the requested size while the job is pending, and the actual rendered pixels once completed (some models render close-but-not-exact sizes, e.g. 1248x704 for a 1280x720 request).

secondsstringrequired

Clip duration in seconds. Shows the requested value while pending and the actual rendered duration once completed.

audiobooleanrequired

Whether the video was generated with a synchronized audio track. Always false on models that only produce silent video.

errorobjectrequired

Error details when the job failed; null otherwise.

remixed_from_video_idstringrequired

Always null — remixing is not supported.

Request

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: "your-scx-api-key",
  baseURL: "https://api.scx.ai/v1",
});

const video = await client.videos.create({
  model: "seedance-1.0-pro",
  prompt: "A red kite surfs over turquoise waves at golden hour, cinematic",
  seconds: "5",
  size: "1280x720",
});

console.log(video.id, video.status); // video_… queued

Response

{
  "id": "video_9f6c2c4e-1c39-4d1a-b7bd-1c0d6d6a2f30",
  "object": "video",
  "model": "seedance-1.0-pro",
  "status": "queued",
  "progress": 0,
  "created_at": 1783652619,
  "completed_at": null,
  "expires_at": null,
  "prompt": "A red kite surfs over turquoise waves at golden hour, cinematic",
  "size": "1280x720",
  "seconds": "5",
  "audio": false,
  "error": null,
  "remixed_from_video_id": null
}