Batches
The Batch API lets you send large volumes of requests as a single job, processed asynchronously within a 24-hour window. This is ideal for evaluations, bulk classification, embedding generation, and other workloads where you don't need an immediate response.
How batch processing works
- Prepare an input file — create a JSONL file where each line is a request.
- Upload the file — use the Files API with purpose
"batch". - Create a batch — reference the uploaded file and target endpoint.
- Check status — poll the batch until it completes.
- Download results — retrieve the output and error files.
Supported endpoints
Batches can target any of these endpoints:
/v1/chat/completions/v1/embeddings/v1/responses
All requests within a single batch must target the same endpoint.
Preparing the input file
The input file is a JSONL file (one JSON object per line). Each line has the following structure:
{"custom_id": "request-1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "Meta-Llama-3.3-70B-Instruct", "messages": [{"role": "user", "content": "What is 2+2?"}]}}
{"custom_id": "request-1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "Meta-Llama-3.3-70B-Instruct", "messages": [{"role": "user", "content": "What is 2+2?"}]}}
| Field | Required | Description |
|---|---|---|
custom_id | Yes | A unique identifier for this request within the batch. Used to match results. |
method | No | HTTP method. Must be "POST". Defaults to "POST" if omitted. |
url | No | Target endpoint. Must be a supported endpoint. Uses the batch endpoint if omitted. |
body | Yes | The request payload. Must include model. |
Each batch supports a maximum of 50,000 requests and the input file cannot exceed 200MB.
Building the input file
import json
requests = [
{
"custom_id": "request-1",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "Meta-Llama-3.3-70B-Instruct",
"messages": [{"role": "user", "content": "What is the capital of France?"}],
},
},
{
"custom_id": "request-2",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "Meta-Llama-3.3-70B-Instruct",
"messages": [{"role": "user", "content": "What is the capital of Germany?"}],
},
},
]
with open("batch_input.jsonl", "w") as f:
for req in requests:
f.write(json.dumps(req) + "\n")
import json
requests = [
{
"custom_id": "request-1",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "Meta-Llama-3.3-70B-Instruct",
"messages": [{"role": "user", "content": "What is the capital of France?"}],
},
},
{
"custom_id": "request-2",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "Meta-Llama-3.3-70B-Instruct",
"messages": [{"role": "user", "content": "What is the capital of Germany?"}],
},
},
]
with open("batch_input.jsonl", "w") as f:
for req in requests:
f.write(json.dumps(req) + "\n")
Uploading the input file
Upload the JSONL file with purpose "batch" using the Files API:
curl -X POST 'https://api.scx.ai/v1/files' \
-H 'Authorization: Bearer your-scx-api-key' \
-F 'purpose=batch' \
-F 'file=@batch_input.jsonl'
curl -X POST 'https://api.scx.ai/v1/files' \
-H 'Authorization: Bearer your-scx-api-key' \
-F 'purpose=batch' \
-F 'file=@batch_input.jsonl'
The response includes a file id (e.g. file-abc123...) that you will pass as input_file_id when creating the batch.
Creating a batch
from openai import OpenAI
client = OpenAI(
base_url="https://api.scx.ai/v1",
api_key="your-scx-api-key",
)
batch = client.batches.create(
input_file_id="file-a1b2c3d4-e5f6-7890-abcd-ef1234567890",
endpoint="/v1/chat/completions",
completion_window="24h",
metadata={"description": "capital cities evaluation"},
)
print(batch.id, batch.status) # batch_... validating
from openai import OpenAI
client = OpenAI(
base_url="https://api.scx.ai/v1",
api_key="your-scx-api-key",
)
batch = client.batches.create(
input_file_id="file-a1b2c3d4-e5f6-7890-abcd-ef1234567890",
endpoint="/v1/chat/completions",
completion_window="24h",
metadata={"description": "capital cities evaluation"},
)
print(batch.id, batch.status) # batch_... validating
| Parameter | Required | Description |
|---|---|---|
input_file_id | Yes | The file ID returned from the upload step. |
endpoint | Yes | One of /v1/chat/completions, /v1/embeddings, or /v1/responses. |
completion_window | No | Processing window. Currently only "24h" is supported. |
metadata | No | Key-value pairs for your own tracking. Maximum 16 pairs. |
Checking batch status
Retrieve a batch by ID to check its progress:
batch = client.batches.retrieve("batch_abc123def456")
print(batch.status)
print(batch.request_counts) # total, completed, failed
batch = client.batches.retrieve("batch_abc123def456")
print(batch.status)
print(batch.request_counts) # total, completed, failed
Batch status values
| Status | Description |
|---|---|
validating | The input file is being validated. |
in_progress | Requests are actively being processed. |
finalizing | All requests processed; output files are being generated. |
completed | Every request reached a verdict and the output files are ready. Note that this does not mean every request succeeded — check request_counts for the split, and error_file_id for the ones that failed. |
failed | Input file validation failed. No requests were processed. |
expired | The batch exceeded the 24-hour processing window. |
cancelling | A cancellation was requested. |
cancelled | The batch was cancelled. |
Listing batches
Retrieve all batches for your organization with optional pagination:
batches = client.batches.list(limit=10)
for b in batches.data:
print(b.id, b.status, b.request_counts.completed)
batches = client.batches.list(limit=10)
for b in batches.data:
print(b.id, b.status, b.request_counts.completed)
Use the after parameter with the last batch ID to paginate through results.
Downloading results
When a batch completes, output_file_id contains the results and error_file_id contains any failures. Download them using the Files API.
Output file format
Each line in the output file is a JSON object:
{
"id": "batch_req_abc123",
"custom_id": "request-1",
"response": {
"status_code": 200,
"request_id": "req_xyz789",
"body": { "id": "chatcmpl-...", "choices": [...] }
},
"error": null
}
{
"id": "batch_req_abc123",
"custom_id": "request-1",
"response": {
"status_code": 200,
"request_id": "req_xyz789",
"body": { "id": "chatcmpl-...", "choices": [...] }
},
"error": null
}
Error file format
Failed requests appear in the error file:
{
"id": "batch_req_def456",
"custom_id": "request-2",
"response": null,
"error": {
"code": "invalid_request_error",
"message": "model not found"
}
}
{
"id": "batch_req_def456",
"custom_id": "request-2",
"response": null,
"error": {
"code": "invalid_request_error",
"message": "model not found"
}
}
Check request_counts on the batch object to quickly see how many requests succeeded or failed without downloading the files.
Cancelling a batch
You can cancel a batch that is currently in_progress:
batch = client.batches.cancel("batch_abc123def456")
print(batch.status) # cancelling
batch = client.batches.cancel("batch_abc123def456")
print(batch.status) # cancelling
Only batches with status in_progress can be cancelled. Requests that have already completed will still be available in the output file after cancellation finishes.
End-to-end example
A complete example that creates a batch, polls until completion, and prints results:
import json
import time
from openai import OpenAI
client = OpenAI(
base_url="https://api.scx.ai/v1",
api_key="your-scx-api-key",
)
# 1. Build the input file
requests = [
{
"custom_id": f"req-{i}",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "Meta-Llama-3.3-70B-Instruct",
"messages": [{"role": "user", "content": f"What is {i} * {i}?"}],
},
}
for i in range(1, 6)
]
with open("batch_input.jsonl", "w") as f:
for req in requests:
f.write(json.dumps(req) + "\n")
# 2. Upload the file
uploaded = client.files.create(
file=open("batch_input.jsonl", "rb"),
purpose="batch",
)
# 3. Create the batch
batch = client.batches.create(
input_file_id=uploaded.id,
endpoint="/v1/chat/completions",
completion_window="24h",
)
print(f"Batch created: {batch.id}")
# 4. Poll until complete
while batch.status not in ("completed", "failed", "expired", "cancelled"):
time.sleep(10)
batch = client.batches.retrieve(batch.id)
print(f"Status: {batch.status} ({batch.request_counts.completed}/{batch.request_counts.total})")
# 5. Download and print results
if batch.output_file_id:
content = client.files.content(batch.output_file_id)
for line in content.text.strip().split("\n"):
result = json.loads(line)
answer = result["response"]["body"]["choices"][0]["message"]["content"]
print(f"{result['custom_id']}: {answer}")
import json
import time
from openai import OpenAI
client = OpenAI(
base_url="https://api.scx.ai/v1",
api_key="your-scx-api-key",
)
# 1. Build the input file
requests = [
{
"custom_id": f"req-{i}",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "Meta-Llama-3.3-70B-Instruct",
"messages": [{"role": "user", "content": f"What is {i} * {i}?"}],
},
}
for i in range(1, 6)
]
with open("batch_input.jsonl", "w") as f:
for req in requests:
f.write(json.dumps(req) + "\n")
# 2. Upload the file
uploaded = client.files.create(
file=open("batch_input.jsonl", "rb"),
purpose="batch",
)
# 3. Create the batch
batch = client.batches.create(
input_file_id=uploaded.id,
endpoint="/v1/chat/completions",
completion_window="24h",
)
print(f"Batch created: {batch.id}")
# 4. Poll until complete
while batch.status not in ("completed", "failed", "expired", "cancelled"):
time.sleep(10)
batch = client.batches.retrieve(batch.id)
print(f"Status: {batch.status} ({batch.request_counts.completed}/{batch.request_counts.total})")
# 5. Download and print results
if batch.output_file_id:
content = client.files.content(batch.output_file_id)
for line in content.text.strip().split("\n"):
result = json.loads(line)
answer = result["response"]["body"]["choices"][0]["message"]["content"]
print(f"{result['custom_id']}: {answer}")
Limitations
| Constraint | Limit |
|---|---|
| Completion window | 24 hours |
| Requests per batch | 50,000 |
| Input file size | 200 MB |
| Metadata key-value pairs | 16 |
For detailed parameter documentation, see the Batch API reference.