Overview
The NanoGPT Batch API runs many independent requests asynchronously at a lower token price than the equivalent synchronous calls. It is a good fit for classification, summarization, evals, synthetic data, document processing, and image analysis where immediate results are not required. You can submit a batch in either of two ways:- File-backed API: upload JSONL, create a batch, poll its status, then download output and error files. This OpenAI-compatible workflow is best for large or reusable inputs.
- Inline API: send the requests in the batch creation body and receive the results when polling the completed batch. This is simpler for smaller jobs.
/v1/chat/completions or /v1/responses.
All rows in one batch must use the same endpoint and the same model. Batch requests are non-streaming and the only supported completion window is
24h.Authentication and base URLs
All requests require an API key:
Use
api.nano-gpt.com for uploads. The main website host can reject larger multipart requests before they reach the Batch API.
Supported row endpoints
Each request in a batch must target one of these endpoints:/v1/chat/completions/v1/responses
Model support
Chat Completions batches
/v1/chat/completions batches support selected OpenAI, Claude, Gemini, and additional models. Always submit the canonical NanoGPT model slug; upstream routing is handled automatically.
Supported canonical model slugs include:
deepseek/deepseek-v4-proz-ai/glm-5.2openai/gpt-oss-120bopenai/gpt-oss-20bthinkingmachines/inklingmoonshotai/kimi-k2.6moonshotai/kimi-k2.7-codemoonshotai/kimi-k3minimax/minimax-m2.7minimax/minimax-m3meta/muse-glimmer-30bnvidia/nemotron-3.5-lightningnvidia/nemotron-3-ultra-550b-a55bqwen-3.6-plusqwen3.7-plus
:thinking suffix is accepted where a thinking variant is configured. Claude thinking aliases are also accepted for compatible models; a numeric thinking budget must be lower than max_tokens.
For the live list, call GET /api/v1/models?detailed=true and select models whose supported_batch_endpoints include /v1/chat/completions.
Model availability changes. Validate a small batch before submitting a large job; an unsupported model returns an unsupported_model error.
Responses batches
/v1/responses batches currently support direct OpenAI models only. The openai/ prefix is accepted. gpt-5.2-pro and gpt-5.4-pro, including dated snapshots, are not available through the upstream Responses Batch API and are rejected.
Responses rows support function and custom tools, structured text output, and remote or data-URL image inputs. Provider-hosted tools and stateful Responses features are not supported in batches.
File-backed API
Endpoints
POST /filesGET /files/{file_id}GET /files/{file_id}/contentPOST /batchesGET /batchesGET /batches/{batch_id}POST /batches/{batch_id}/cancel
JSONL rules
Each non-empty line must be a JSON object with:- a unique, non-empty
custom_id method: "POST"urlset to/v1/chat/completionsor/v1/responses- a
bodyobject containing the model and endpoint-specific input
stream: true is rejected.
Chat Completions rows require a non-empty messages array and a positive max_tokens or max_completion_tokens. They support text and compatible image_url content. Image URLs may use HTTP, HTTPS, or base64 data URLs for PNG, JPEG, GIF, and WebP images.
Responses rows require a non-empty input and an integer max_output_tokens of at least 16.
Chat Completions example
Responses example
Upload the file
Create the batch
Setendpoint to the same endpoint used by every JSONL row:
Poll, cancel, and download
validating, in_progress, finalizing, completed, failed, expired, cancelling, and cancelled. A completed batch normally has an output_file_id; row failures may also produce an error_file_id.
Inline API
Inline batches accept up to 10,000 requests, 20 MiB of normalized input, and 250,000 aggregate requested output tokens. Use the file-backed API for larger inputs.Endpoints
POST /batchesGET /batchesGET /batches/{batch_id}POST /batches/{batch_id}/cancel
https://api.nano-gpt.com/api/beta base URL.
Create an inline batch
202 Accepted. Poll GET /api/beta/batches/{batch_id} and read results after completion. Cancel an active job with POST /api/beta/batches/{batch_id}/cancel.
The batch-level model is inherited by every request body. If an inline row omits its endpoint-specific output cap, NanoGPT applies a 4096-token default. File-backed rows must always include the output cap explicitly.
List batches
Use the authenticated list route for the workflow you want to inspect. Each route returns only batches visible to the supplied API key, ordered newest first.
Both routes support the same filters and bidirectional cursor pagination:
after and before cannot be combined. Cursors are batch IDs, not arbitrary tokens. Creation boundaries accept Unix timestamps in seconds or timezone-qualified ISO 8601/RFC 3339 timestamps with no more than millisecond precision.
Supported status values are validating, in_progress, finalizing, completed, failed, expired, cancelling, and cancelled.
Filter inline batches
Filter all batches
Response and pagination
Both endpoints return the standard list envelope:datais always ordered newest first.first_idandlast_idarenullfor an empty page.- To fetch older batches, send
after=<last_id>with the same filters. - To return to the adjacent newer page, send
before=<first_id>with the same filters. has_moreindicates that more records exist in the direction being paged.
id, object, endpoint, model, completion_window, status, created_at, finalized_at, request_counts, usage, error, and a results field that is always null.
The OpenAI-compatible list includes both inline and file-backed jobs. Its items retain the standard batch shape, including input_file_id, output_file_id, error_file_id, lifecycle timestamps, request_counts, metadata, usage, and errors where applicable.
Listing errors
Both list routes return:401for missing or invalid authentication.400whenafterandbeforeare both present.400for an unknown or inaccessible cursor. The inline-only route also rejects a cursor that identifies a non-inline batch. A cursor does not need to match the active filters; it only supplies the page boundary.400for an unsupportedstatusvalue.400for an invalid creation timestamp or a reversed creation-time range.
Responses Batch restrictions
Responses Batch is stateless and executes directly through the upstream batch service. NanoGPT forcesstore: false and rejects:
previous_response_id,conversation, andbackground: true- reusable
promptreferences input_file,item_reference, video inputs, andinput_image.file_id- provider-hosted tools; function and custom tools remain supported
- NanoGPT-only features such as Advisor, memory, scraping, retention overrides, provider or BYOK controls, caching controls, and billing overrides
text.format field is supported.
Billing
Batch jobs use NanoGPT account balance and do not use subscription included tokens. At creation, NanoGPT checks the balance against a conservative maximum-liability estimate based on the input and output caps. Completed usage is charged once after the batch reaches a terminal state. If there is no billable usage, no usage charge is created. Supported batch token usage is priced 50% below the equivalent synchronous request. Non-token charges, where supported, keep their normal rate. Use the live pricing page or pricing API as the source of truth.Common errors
- Unsupported endpoint: use
/v1/chat/completionsor/v1/responses, consistently across all rows. - Missing output cap: add
max_tokensormax_completion_tokensfor Chat Completions, ormax_output_tokens >= 16for Responses. - Mixed model: every row must use the same model.
- Streaming unsupported: remove
stream: true. - Unsupported Responses model: choose a direct OpenAI model supported by the upstream Batch API.
- Unsupported Responses field: remove stateful features, provider-hosted tools, file references, or NanoGPT-only extensions.