Endpoints & Rate Limits
Every request to SCX.ai goes to a single base URL over HTTPS — there's no per-region routing to configure.
https://api.scx.ai/v1
https://api.scx.ai/v1
Don't have an API key? Get yours from the API keys page.
API endpoints
All endpoints are relative to the base URL above.
| Resource | Path |
|---|---|
| Chat completions | /chat/completions |
| Completions | /completions |
| Embeddings | /embeddings |
| Moderations | /moderations |
| Messages | /messages |
| Models | /models |
| Responses | /responses |
| Audio (transcription/translation) | /audio/transcriptions, /audio/translations |
| Audio speech | /audio/speech |
| Audio voices | /audio/voices |
| Vector stores | /vector_stores |
| Connectors | /connectors |
| Files | /files |
| Batches | /batches |
| Videos | /videos |
| Images | /images/generations, /images/edits |
Request size limits
Each endpoint enforces a maximum request body size. A request over its endpoint's limit is rejected with a 413 before it's processed — retrying the same body can never succeed, so shrink the payload first.
| Endpoint | Max body size |
|---|---|
| Default (any endpoint not listed below) | 100 MB |
/audio/transcriptions, /audio/translations | 25 MB |
/videos | 40 MB |
/videos/extensions | 210 MB |
/images/generations | 1 MB |
/images/edits | 40 MB |
/files | 200 MB per file |
/batches input file | 200 MB |
{
"error": {
"message": "request body is too large for this endpoint's limit",
"type": "invalid_request_error"
}
}
{
"error": {
"message": "request body is too large for this endpoint's limit",
"type": "invalid_request_error"
}
}
Rate limits
Rate limits are enforced per organization and, where a request resolves to a specific model, per model — not per API key or per IP. Three limits apply together:
| Limit | Meaning |
|---|---|
| RPM | Requests per minute |
| TPM | Tokens per minute |
| TPD | Tokens per day |
Your actual RPM/TPM/TPD depend on your organization's plan and can change over time, so they aren't published as a fixed table here — check the usage page for your current limits, or read them straight off the response headers on any request:
| Header | Description |
|---|---|
X-RateLimit-Limit-RPM | Requests-per-minute limit |
X-RateLimit-Remaining-RPM | Requests remaining in the current minute |
X-RateLimit-Limit-TPM | Tokens-per-minute limit |
X-RateLimit-Remaining-TPM | Tokens remaining in the current minute |
X-RateLimit-Limit-TPD | Tokens-per-day limit |
X-RateLimit-Remaining-TPD | Tokens remaining in the current day |
A limit header reading unlimited means that dimension has no cap for your organization.
Exceeding any of the three returns a 429:
{
"error": {
"message": "Rate limit exceeded",
"type": "rate_limit_error"
}
}
{
"error": {
"message": "Rate limit exceeded",
"type": "rate_limit_error"
}
}
On a 429, back off and retry — the same request will succeed once the window resets.
Spend limits
Spend can be capped at three levels: an API key, a member of the organization, and the organization as a whole. A request is refused by whichever budget it reaches first, checked from the narrowest: the key, the key's budget for the requested model, the member, then the organization. Budgets count usage paid from your credit balance; they don't cap usage covered by a subscription or invoiced separately.
On a request that goes through, these headers report whichever of your own budgets has less left, your key's or your member budget:
| Header | Description |
|---|---|
X-Budget-Limit | Maximum spend for the current period |
X-Budget-Current-Spend | Spend so far in the current period |
X-Budget-Remaining | Remaining budget in the current period |
They are omitted when neither your key nor your membership has a budget, and they never report the organization's budget on a request that goes through.
A request over a budget returns a 429 whose code names the budget that refused it, with the headers describing that budget:
{
"error": {
"message": "Member budget exceeded",
"type": "quota_exceeded",
"code": "member_budget_exceeded"
}
}
{
"error": {
"message": "Member budget exceeded",
"type": "quota_exceeded",
"code": "member_budget_exceeded"
}
}
code | Refused by |
|---|---|
key_budget_exceeded | The API key's budget |
key_model_budget_exceeded | The API key's budget for the requested model |
member_budget_exceeded | Your member budget in the organization |
organization_budget_exceeded | The organization's budget |
Unlike a rate limit, retrying won't succeed until the budget's period resets or the budget is raised.
Errors and status codes
Every error response follows the same shape:
{
"error": {
"message": "human-readable description",
"type": "error_type"
}
}
{
"error": {
"message": "human-readable description",
"type": "error_type"
}
}
| Status | Type | Meaning |
|---|---|---|
| 400 | invalid_request_error | The request is malformed, or a parameter is missing or invalid. |
| 401 | authentication_error | The API key is missing or invalid. |
| 402 | insufficient_credit | Your organization has run out of credit. |
| 403 | permission_error | Your API key doesn't have permission for this resource. |
| 404 | not_found_error | The requested resource doesn't exist. |
| 413 | invalid_request_error | The request body is over this endpoint's size limit — see Request size limits. |
| 429 | rate_limit_error / quota_exceeded | A rate or spend limit was hit — see Rate limits and Spend limits. |
| 500 | api_error | Something went wrong on our end. |
| 529 | overloaded_error | The model is temporarily overloaded. |
400–404 mean the request itself needs to change before retrying — retrying the same request will fail the same way. 429 and 529 are transient: back off and retry, honoring the Retry-After header when one is present. 500 is also generally safe to retry.