These examples use the website API host, which currently lists their recommended models. The direct API host can have a different catalog during rollouts. For larger requests or longer runtimes, use
https://api.nano-gpt.com/api/v1 with a model from that host’s catalog. See API Hosts for limits and availability checks.Overview
The NanoGPT API offers multiple ways to generate text, including OpenAI-compatible endpoints and our legacy options. This guide covers all available text generation methods. If you are using a TEE-backed model (e.g., prefixed withTEE/), you can also verify the enclave attestation and signatures for your chat completions. See the TEE Model Verification guide for more details.
For authenticated API-key requests, you can opt in to a paid input safety preflight by sending the moderation header. See Inline Moderation for supported text routes, model selection, billing behavior, and error codes.
Provider Selection
Provider selection is available for supported open-source models.X-Provider explicitly selects a provider for the request and is always billed pay-as-you-go at the selected provider’s price, including provider-selection markup. For subscription users, sending X-Provider bypasses subscription coverage for that request; X-Billing-Mode: paygo is only needed when forcing pay-as-you-go without an explicit provider or when saved provider preferences should apply to subscription-included traffic. See Provider Selection and Pay-As-You-Go Billing Override.
For one-off routing preferences, append a suffix to eligible model IDs:
:fast/:speedfor fastest estimated completion:cheap/:price/:floorfor cheapest provider:throughputfor highest TPS:latencyfor lowest TTFT:toolsfor tools-capable routing
OpenAI Compatible Endpoints
Chat Completions (v1/chat/completions)
This endpoint mimics OpenAI’s chat completions API: For high-volume offline workloads where latency is not important, use the Batch API to upload JSONL chat completion requests and process them asynchronously.Responses (v1/responses)
Use the OpenAI Responses-compatible endpoint for stateful threading (previous_response_id), background processing, and Responses-style streaming events. See the dedicated docs at /api-reference/endpoint/responses.
Direct Web Search (api/web)
UsePOST /api/web when you need direct search control instead of chat orchestration:
- Explicit
querypayload control - Linkup output types:
searchResults,sourcedAnswer,structured - Date and domain filters (
fromDate,toDate,includeDomains,excludeDomains)
Text Completions (v1/completions)
This endpoint mimics OpenAI’s legacy text completions API:POST /api/v1/completions is best effort. Performance and compatibility may be less consistent than POST /api/v1/chat/completions because some upstream providers do not support the legacy completions API.Legacy Text Completions
For the older, non-OpenAI compatible endpoint:Caching (Implicit and Explicit Controls)
For the full guide (supported models, thresholds, pricing, and usage fields), see Prompt Caching. NanoGPT automatically applies implicit caching on providers/models that support it (including OpenAI, Gemini, and many open-source provider/model routes), so most requests do not need caching flags. Set top-levelcaching: true or append :caching / :cache / :cached to the model when you want NanoGPT to route the request to any available provider that supports prompt/input caching. This is capability-based routing: you do not need to choose a provider. If no cache-capable provider is available for the model, the request fails rather than silently using a non-caching provider.
Use explicit prompt-caching controls (prompt_caching, promptCaching, and body-level cache_control alias, plus inline cache_control) when you need Claude-specific cache boundaries, TTL selection, or prompt_caching.stickyProvider consistency control. Top-level caching: true does not add Anthropic-style cache_control markers or configure cache TTLs.
Cache-Capable Provider Routing
caching: true prefers cache-capable providers and enables soft provider affinity: after a successful matching request, later matching requests from the same API key or session try to reuse that provider. Explicit provider or routing preferences can take precedence. Affinity improves the chance of a cache hit; it does not guarantee one.
stickyprovider defaults to false. Setting it to false does not disable this automatic affinity. Set it to true when preserving the provider matters more than failover:
stickyProvider is accepted as a camelCase alias for stickyprovider. The nested prompt_caching.stickyProvider option controls the same strict behavior; an explicit top-level value wins. Strict mode disables automatic fallback, so a provider failure can return an error instead of switching providers.
Equivalent model suffix:
caching: true routes, NanoGPT filters to available, non-excluded, cache-capable providers and prefers a recorded provider when still usable. Otherwise it tries eligible saved provider preferences, then chooses using prices and recent aggregate cache-read observations. Saved restrictions remain binding. A price routing preference ranks candidates again for each request and does not promise cache affinity. See prompt caching for cache eligibility, affinity lifetime, and provider-specific behavior.
The prompt_caching / promptCaching helper accepts these options:
Explicit cache_control markers
cache_controlbelongs to individual content blocks (system,user, tool definitions, etc.). Each marker caches the entire prefix up to and including that block.- Supported explicit TTLs are
5mand1h(Claude flows). Omitttlto use the default5mwindow. anthropic-beta: prompt-caching-2024-07-31is supported for compatibility and required for Anthropic-native Claude caching flows.- For implicit-caching providers, no explicit
cache_controlmarkers are required. - Check
usage.prompt_tokens_details.cached_tokensin NanoGPT’s response to confirm what was billed at the discounted rate.
Using the prompt_caching helper
If you prefer not to duplicate cache_control entries manually, NanoGPT accepts a helper object that tags the leading prefix for you.
cut_after_message_index is zero-based and points at the last message in the static prefix. NanoGPT will attach a cache_control block with your TTL to each message up to that index before forwarding the request upstream. If you omit cut_after_message_index, NanoGPT will select a cache boundary automatically; set it explicitly if you need full control. If you need different cache durations or non-contiguous breakpoints, fall back to explicit cache_control markers in your messages array.
Explicit Prompt Cache Consistency
NanoGPT automatically fails over to backup services when the primary service is temporarily unavailable. While this ensures high availability, it can break your prompt cache because each backend service maintains its own separate cache. If cache consistency is more important than availability for your use case, you can enable thestickyProvider option:
stickyProvider: false(default) — If the primary service fails, NanoGPT automatically retries with a backup service. Your request succeeds, but the cache may be lost (you’ll pay full price for that request and need to rebuild the cache).stickyProvider: true— If the primary service fails, NanoGPT returns a 503 error instead of failing over. An existing cache may still be usable after recovery if it has not expired; strict mode does not extend its TTL.
stickyProvider: true:
- You have very large cached contexts where cache misses are expensive
- You prefer to retry failed requests yourself rather than pay for cache rebuilds
- Cost predictability is more important than request success rate
stickyProvider: false (default):
- You prefer requests to always succeed when possible
- Occasional cache misses are acceptable
- You’re using shorter contexts where cache rebuilds are inexpensive
Chat Completions with Web Search
Enable real-time web access for any model by appending special suffixes:Web Search Options
:online- Standard search with 10 results ($0.006 per request). On GPT-5+ and o-series models it uses native search instead, billed per search; see Native search.:online/linkup-deep- Deep iterative search ($0.06 per request)
:online/sofya, :online/exa-instant, :online/exa-deep-reasoning, :online/brave, and :online/valyu-web-deep, see Model Suffixes.
Web search dramatically improves factuality - Gemini 3 Flash Preview with web access shows a 10x improvement in accuracy, making it twice as accurate as non-web baselines.
For direct /api/web usage with structured output, domain/date filters, and explicit query control, see Direct Web Search API.