Skip to main content

Overview

The NanoGPT Batch API runs many independent requests asynchronously at a lower token price than the equivalent synchronous calls. It is a good fit for classification, summarization, evals, synthetic data, document processing, and image analysis where immediate results are not required. You can submit a batch in either of two ways:
  • File-backed API: upload JSONL, create a batch, poll its status, then download output and error files. This OpenAI-compatible workflow is best for large or reusable inputs.
  • Inline API: send the requests in the batch creation body and receive the results when polling the completed batch. This is simpler for smaller jobs.
Both workflows support rows targeting /v1/chat/completions or /v1/responses.
All rows in one batch must use the same endpoint and the same model. Batch requests are non-streaming and the only supported completion window is 24h.

Authentication and base URLs

All requests require an API key:
Use the dedicated API host for batch operations: Use api.nano-gpt.com for uploads. The main website host can reject larger multipart requests before they reach the Batch API.

Supported row endpoints

Each request in a batch must target one of these endpoints:
  • /v1/chat/completions
  • /v1/responses
Completions, embeddings, image generation, audio, video, transcription, TTS, moderation, and other endpoints cannot be used as batch rows.

Model support

Chat Completions batches

/v1/chat/completions batches support selected OpenAI, Claude, Gemini, and additional models. Always submit the canonical NanoGPT model slug; upstream routing is handled automatically. Supported canonical model slugs include:
  • deepseek/deepseek-v4-pro
  • z-ai/glm-5.2
  • openai/gpt-oss-120b
  • openai/gpt-oss-20b
  • thinkingmachines/inkling
  • moonshotai/kimi-k2.6
  • moonshotai/kimi-k2.7-code
  • moonshotai/kimi-k3
  • minimax/minimax-m2.7
  • minimax/minimax-m3
  • meta/muse-glimmer-30b
  • nvidia/nemotron-3.5-lightning
  • nvidia/nemotron-3-ultra-550b-a55b
  • qwen-3.6-plus
  • qwen3.7-plus
The :thinking suffix is accepted where a thinking variant is configured. Claude thinking aliases are also accepted for compatible models; a numeric thinking budget must be lower than max_tokens. For the live list, call GET /api/v1/models?detailed=true and select models whose supported_batch_endpoints include /v1/chat/completions. Model availability changes. Validate a small batch before submitting a large job; an unsupported model returns an unsupported_model error.

Responses batches

/v1/responses batches currently support direct OpenAI models only. The openai/ prefix is accepted. gpt-5.2-pro and gpt-5.4-pro, including dated snapshots, are not available through the upstream Responses Batch API and are rejected. Responses rows support function and custom tools, structured text output, and remote or data-URL image inputs. Provider-hosted tools and stateful Responses features are not supported in batches.

File-backed API

Endpoints

  • POST /files
  • GET /files/{file_id}
  • GET /files/{file_id}/content
  • POST /batches
  • GET /batches
  • GET /batches/{batch_id}
  • POST /batches/{batch_id}/cancel

JSONL rules

Each non-empty line must be a JSON object with:
  • a unique, non-empty custom_id
  • method: "POST"
  • url set to /v1/chat/completions or /v1/responses
  • a body object containing the model and endpoint-specific input
All rows must use the same endpoint and model. stream: true is rejected. Chat Completions rows require a non-empty messages array and a positive max_tokens or max_completion_tokens. They support text and compatible image_url content. Image URLs may use HTTP, HTTPS, or base64 data URLs for PNG, JPEG, GIF, and WebP images. Responses rows require a non-empty input and an integer max_output_tokens of at least 16.

Chat Completions example

Responses example

Upload the file

Create the batch

Set endpoint to the same endpoint used by every JSONL row:

Poll, cancel, and download

Batch statuses are validating, in_progress, finalizing, completed, failed, expired, cancelling, and cancelled. A completed batch normally has an output_file_id; row failures may also produce an error_file_id.

Inline API

Inline batches accept up to 10,000 requests, 20 MiB of normalized input, and 250,000 aggregate requested output tokens. Use the file-backed API for larger inputs.

Endpoints

  • POST /batches
  • GET /batches
  • GET /batches/{batch_id}
  • POST /batches/{batch_id}/cancel
All inline routes use the https://api.nano-gpt.com/api/beta base URL.

Create an inline batch

Creation returns 202 Accepted. Poll GET /api/beta/batches/{batch_id} and read results after completion. Cancel an active job with POST /api/beta/batches/{batch_id}/cancel. The batch-level model is inherited by every request body. If an inline row omits its endpoint-specific output cap, NanoGPT applies a 4096-token default. File-backed rows must always include the output cap explicitly.

List batches

Use the authenticated list route for the workflow you want to inspect. Each route returns only batches visible to the supplied API key, ordered newest first. Both routes support the same filters and bidirectional cursor pagination: after and before cannot be combined. Cursors are batch IDs, not arbitrary tokens. Creation boundaries accept Unix timestamps in seconds or timezone-qualified ISO 8601/RFC 3339 timestamps with no more than millisecond precision. Supported status values are validating, in_progress, finalizing, completed, failed, expired, cancelling, and cancelled.

Filter inline batches

You can also filter by a strict creation-time range:

Filter all batches

Response and pagination

Both endpoints return the standard list envelope:
  • data is always ordered newest first.
  • first_id and last_id are null for an empty page.
  • To fetch older batches, send after=<last_id> with the same filters.
  • To return to the adjacent newer page, send before=<first_id> with the same filters.
  • has_more indicates that more records exist in the direction being paged.
For example:
Inline list responses contain metadata only. results is always null, including for completed batches. Retrieve GET /api/beta/batches/{batch_id} to obtain a completed inline batch’s results.
Inline list items include id, object, endpoint, model, completion_window, status, created_at, finalized_at, request_counts, usage, error, and a results field that is always null. The OpenAI-compatible list includes both inline and file-backed jobs. Its items retain the standard batch shape, including input_file_id, output_file_id, error_file_id, lifecycle timestamps, request_counts, metadata, usage, and errors where applicable.

Listing errors

Both list routes return:
  • 401 for missing or invalid authentication.
  • 400 when after and before are both present.
  • 400 for an unknown or inaccessible cursor. The inline-only route also rejects a cursor that identifies a non-inline batch. A cursor does not need to match the active filters; it only supplies the page boundary.
  • 400 for an unsupported status value.
  • 400 for an invalid creation timestamp or a reversed creation-time range.

Responses Batch restrictions

Responses Batch is stateless and executes directly through the upstream batch service. NanoGPT forces store: false and rejects:
  • previous_response_id, conversation, and background: true
  • reusable prompt references
  • input_file, item_reference, video inputs, and input_image.file_id
  • provider-hosted tools; function and custom tools remain supported
  • NanoGPT-only features such as Advisor, memory, scraping, retention overrides, provider or BYOK controls, caching controls, and billing overrides
Remote HTTP(S) image URLs and image data URLs are accepted. Structured output through the Responses text.format field is supported.

Billing

Batch jobs use NanoGPT account balance and do not use subscription included tokens. At creation, NanoGPT checks the balance against a conservative maximum-liability estimate based on the input and output caps. Completed usage is charged once after the batch reaches a terminal state. If there is no billable usage, no usage charge is created. Supported batch token usage is priced 50% below the equivalent synchronous request. Non-token charges, where supported, keep their normal rate. Use the live pricing page or pricing API as the source of truth.

Common errors

  • Unsupported endpoint: use /v1/chat/completions or /v1/responses, consistently across all rows.
  • Missing output cap: add max_tokens or max_completion_tokens for Chat Completions, or max_output_tokens >= 16 for Responses.
  • Mixed model: every row must use the same model.
  • Streaming unsupported: remove stream: true.
  • Unsupported Responses model: choose a direct OpenAI model supported by the upstream Batch API.
  • Unsupported Responses field: remove stateful features, provider-hosted tools, file references, or NanoGPT-only extensions.