Learn how to use OpenAI’s Batch API to send asynchronous groups of requests with 50% lower costs, a separate pool of significantly higher rate limits, and a clear 24-hour turnaround time. The service is ideal for processing jobs that don’t require immediate responses. You can also explore the API reference directly here.
Overview
While some uses of the OpenAI Platform require you to send synchronous requests, there are many cases where requests do not need an immediate response or rate limits prevent you from executing a large number of queries quickly. Batch processing jobs are often helpful in use cases like:
- Running evaluations
- Classifying large datasets
- Embedding content repositories
- Queuing large offline video-render jobs
The Batch API offers a straightforward set of endpoints that allow you to collect a set of requests into a single file, kick off a batch processing job to execute these requests, query for the status of that batch while the underlying requests execute, and eventually retrieve the collected results when the batch is complete.
Compared to using standard endpoints directly, Batch API has:
- Better cost efficiency: 50% cost discount compared to synchronous APIs
- Higher rate limits: Substantially more headroom compared to the synchronous APIs
- Fast completion times: Each batch completes within 24 hours (and often more quickly)
Getting started
1. Prepare your batch file
Batches start with a .jsonl file where each line contains the details of an individual request to the API. For now, the available endpoints are:
/v1/responses(Responses API)/v1/chat/completions(Chat Completions API)/v1/embeddings(Embeddings API)/v1/completions(Completions API)/v1/moderations(Moderation guide)/v1/images/generations(Images API)/v1/images/edits(Images API)/v1/videos(Video generation guide)
For a given input file, the parameters in each line’s body field are the same as the parameters for the underlying endpoint. Each request must include a unique custom_id value, which you can use to reference results after completion. Here’s an example of an input file with 2 requests. Note that each input file can only include requests to a single model.
For video generation in Batch:
- Batch currently supports
POST /v1/videosonly. - Batch requests for videos must use JSON, not multipart.
- Upload assets ahead of time and pass supported asset references in the request body rather than using multipart uploads.
- Use
input_referencefor image-guided generations in Batch. In JSON requests, passinput_referenceas an object with eitherfile_idorimage_url. - Multipart
input_referenceuploads, including video reference inputs, aren’t supported in Batch. - Batch-generated videos are available for download for up to
24hours after the batch completes.
When targeting /v1/moderations, include an input field in every request body. Batch accepts plain-text inputs and content arrays with text or image inputs using omni-moderation-latest. The Batch worker rejects requests that set stream=true, matching the synchronous moderation endpoint.
{"custom_id": "request-1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "gpt-3.5-turbo-0125", "messages": [{"role": "system", "content": "You are a helpful assistant."},{"role": "user", "content": "Hello world!"}],"max_tokens": 1000}}
{"custom_id": "request-2", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "gpt-3.5-turbo-0125", "messages": [{"role": "system", "content": "You are an unhelpful assistant."},{"role": "user", "content": "Hello world!"}],"max_tokens": 1000}}
Moderation input examples
Text-only request:
{
"custom_id": "moderation-text-1",
"method": "POST",
"url": "/v1/moderations",
"body": {
"model": "omni-moderation-latest",
"input": "This is a harmless test sentence."
}
}
Request with text and image input: