Batches

The Batch API lets you send large volumes of requests as a single job, processed asynchronously within a 24-hour window. This is ideal for evaluations, bulk classification, embedding generation, and other workloads where you don't need an immediate response.

How batch processing works

  1. Prepare an input file — create a JSONL file where each line is a request.
  2. Upload the file — use the Files API with purpose "batch".
  3. Create a batch — reference the uploaded file and target endpoint.
  4. Check status — poll the batch until it completes.
  5. Download results — retrieve the output and error files.

Supported endpoints

Batches can target any of these endpoints:

  • /v1/chat/completions
  • /v1/embeddings
  • /v1/responses

All requests within a single batch must target the same endpoint.

Preparing the input file

The input file is a JSONL file (one JSON object per line). Each line has the following structure:

json
{"custom_id": "request-1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "Meta-Llama-3.3-70B-Instruct", "messages": [{"role": "user", "content": "What is 2+2?"}]}}
FieldRequiredDescription
custom_idYesA unique identifier for this request within the batch. Used to match results.
methodNoHTTP method. Must be "POST". Defaults to "POST" if omitted.
urlNoTarget endpoint. Must be a supported endpoint. Uses the batch endpoint if omitted.
bodyYesThe request payload. Must include model.

Building the input file

import json

requests = [
    {
        "custom_id": "request-1",
        "method": "POST",
        "url": "/v1/chat/completions",
        "body": {
            "model": "Meta-Llama-3.3-70B-Instruct",
            "messages": [{"role": "user", "content": "What is the capital of France?"}],
        },
    },
    {
        "custom_id": "request-2",
        "method": "POST",
        "url": "/v1/chat/completions",
        "body": {
            "model": "Meta-Llama-3.3-70B-Instruct",
            "messages": [{"role": "user", "content": "What is the capital of Germany?"}],
        },
    },
]

with open("batch_input.jsonl", "w") as f:
    for req in requests:
        f.write(json.dumps(req) + "\n")

Uploading the input file

Upload the JSONL file with purpose "batch" using the Files API:

bash
curl -X POST 'https://api.scx.ai/v1/files' \
  -H 'Authorization: Bearer your-scx-api-key' \
  -F 'purpose=batch' \
  -F 'file=@batch_input.jsonl'

The response includes a file id (e.g. file-abc123...) that you will pass as input_file_id when creating the batch.

Creating a batch

from openai import OpenAI

client = OpenAI(
    base_url="https://api.scx.ai/v1",
    api_key="your-scx-api-key",
)

batch = client.batches.create(
    input_file_id="file-a1b2c3d4-e5f6-7890-abcd-ef1234567890",
    endpoint="/v1/chat/completions",
    completion_window="24h",
    metadata={"description": "capital cities evaluation"},
)

print(batch.id, batch.status)  # batch_... validating
ParameterRequiredDescription
input_file_idYesThe file ID returned from the upload step.
endpointYesOne of /v1/chat/completions, /v1/embeddings, or /v1/responses.
completion_windowNoProcessing window. Currently only "24h" is supported.
metadataNoKey-value pairs for your own tracking. Maximum 16 pairs.

Checking batch status

Retrieve a batch by ID to check its progress:

batch = client.batches.retrieve("batch_abc123def456")

print(batch.status)
print(batch.request_counts)  # total, completed, failed

Batch status values

StatusDescription
validatingThe input file is being validated.
in_progressRequests are actively being processed.
finalizingAll requests processed; output files are being generated.
completedEvery request reached a verdict and the output files are ready. Note that this does not mean every request succeeded — check request_counts for the split, and error_file_id for the ones that failed.
failedInput file validation failed. No requests were processed.
expiredThe batch exceeded the 24-hour processing window.
cancellingA cancellation was requested.
cancelledThe batch was cancelled.

Listing batches

Retrieve all batches for your organization with optional pagination:

batches = client.batches.list(limit=10)

for b in batches.data:
    print(b.id, b.status, b.request_counts.completed)

Use the after parameter with the last batch ID to paginate through results.

Downloading results

When a batch completes, output_file_id contains the results and error_file_id contains any failures. Download them using the Files API.

Output file format

Each line in the output file is a JSON object:

json
{
  "id": "batch_req_abc123",
  "custom_id": "request-1",
  "response": {
    "status_code": 200,
    "request_id": "req_xyz789",
    "body": { "id": "chatcmpl-...", "choices": [...] }
  },
  "error": null
}

Error file format

Failed requests appear in the error file:

json
{
  "id": "batch_req_def456",
  "custom_id": "request-2",
  "response": null,
  "error": {
    "code": "invalid_request_error",
    "message": "model not found"
  }
}

Cancelling a batch

You can cancel a batch that is currently in_progress:

batch = client.batches.cancel("batch_abc123def456")

print(batch.status)  # cancelling

End-to-end example

A complete example that creates a batch, polls until completion, and prints results:

import json
import time
from openai import OpenAI

client = OpenAI(
    base_url="https://api.scx.ai/v1",
    api_key="your-scx-api-key",
)

# 1. Build the input file
requests = [
    {
        "custom_id": f"req-{i}",
        "method": "POST",
        "url": "/v1/chat/completions",
        "body": {
            "model": "Meta-Llama-3.3-70B-Instruct",
            "messages": [{"role": "user", "content": f"What is {i} * {i}?"}],
        },
    }
    for i in range(1, 6)
]

with open("batch_input.jsonl", "w") as f:
    for req in requests:
        f.write(json.dumps(req) + "\n")

# 2. Upload the file
uploaded = client.files.create(
    file=open("batch_input.jsonl", "rb"),
    purpose="batch",
)

# 3. Create the batch
batch = client.batches.create(
    input_file_id=uploaded.id,
    endpoint="/v1/chat/completions",
    completion_window="24h",
)
print(f"Batch created: {batch.id}")

# 4. Poll until complete
while batch.status not in ("completed", "failed", "expired", "cancelled"):
    time.sleep(10)
    batch = client.batches.retrieve(batch.id)
    print(f"Status: {batch.status} ({batch.request_counts.completed}/{batch.request_counts.total})")

# 5. Download and print results
if batch.output_file_id:
    content = client.files.content(batch.output_file_id)
    for line in content.text.strip().split("\n"):
        result = json.loads(line)
        answer = result["response"]["body"]["choices"][0]["message"]["content"]
        print(f"{result['custom_id']}: {answer}")

Limitations

ConstraintLimit
Completion window24 hours
Requests per batch50,000
Input file size200 MB
Metadata key-value pairs16

For detailed parameter documentation, see the Batch API reference.