> ## Documentation Index
> Fetch the complete documentation index at: https://docs.nano-gpt.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Provider Selection

> Choose the upstream provider for supported open-source models

# Provider Selection

Provider selection chooses the upstream provider for a supported model. It does not change the model ID. For example, to use Kimi K2.6 through Novita, keep the model as `moonshotai/kimi-k2.6`, or another supported alias for that model, and send the provider separately with `X-Provider`.

Provider selection is optional; if you do nothing, the platform picks the provider. `X-Provider`, a body `provider` string, or a body `provider` object can select or constrain providers for a request. Explicit provider selection is always pay-as-you-go and is charged at the selected provider's price, including provider-selection markup. For subscription users, sending an explicit provider selection bypasses subscription coverage for that request.

`X-Billing-Mode: paygo` is only needed when forcing pay-as-you-go billing without an explicit provider, or when saved provider preferences should apply to subscription-included traffic. See [Pay-As-You-Go Billing Override](/api-reference/miscellaneous/billing-override).

This page intentionally documents only public provider IDs returned by the provider-discovery endpoints. Internal routing/provider names are never part of the public API contract.

## Choose Your Routing Goal

| Goal | Use | What to expect |
| - | - | - |
| Lowest estimated cost for each request | `provider.sort: "price"` | Recalculates the ranking each request; may switch providers and reduce cache reuse. |
| Stay with one provider | `provider.only: ["provider-id"]` and `allow_fallbacks: false` | Fails if that provider cannot serve the request. |
| Reuse a prompt cache when possible | `caching: true` | Prefers a usable provider from previous matching requests, with fallback allowed by default. |

See [examples and precedence](#price-caps-and-cache-reuse) before combining these controls. Cache contents are provider-local; changing providers does not transfer a warm cache.

## When Provider Selection Applies

* Provider selection only applies when a model reports `supportsProviderSelection: true`.
* Use `GET /api/models/:canonicalId/providers` to discover eligible models and provider IDs.
* `supportsProviderSelection` is exposed on the model-specific provider endpoint. Do not rely on `/api/v1/models` or `/api/v1/models?detailed=true` to determine provider-selection support.
* If a model does not support provider selection, the request ignores provider preferences.
* You cannot currently force a provider and have that same request count as subscription-included usage.

## Discover Providers and Pricing

Use this endpoint to list available providers and the pricing you will pay when selecting one.

```
GET /api/models/:canonicalId/providers
```

Provider rows may also include `distillationPolicy`, which indicates whether that specific hosted provider route is allowed for output-based model training or distillation under NanoGPT's recorded provider-terms and model-license rules. See [Distillation Policy](/api-reference/miscellaneous/distillation-policy).

This endpoint is not under `/api/v1`. If your base URL is `https://api.nano-gpt.com/api/v1`, do not append this path to that base URL. Use:

```http theme={null}
GET https://api.nano-gpt.com/api/models/moonshotai%2Fkimi-k2.6/providers
```

URL-encode `/` in the model ID as `%2F`, as shown above. The unencoded slash is treated as a path separator:

```http theme={null}
GET https://api.nano-gpt.com/api/models/moonshotai/kimi-k2.6/providers
```

Here `:canonicalId` means the model ID or one of its supported aliases.

Example response:

```json theme={null}
{
  "canonicalId": "moonshotai/kimi-k2.6",
  "displayName": "Kimi K2.6",
  "supportsProviderSelection": true,
  "defaultPrice": {
    "inputPer1kTokens": 0.0005,
    "outputPer1kTokens": 0.0026
  },
  "providers": [
    {
      "provider": "moonshot",
      "available": true
    },
    {
      "provider": "novita",
      "available": true,
      "distillationPolicy": {
        "status": "allowed",
        "label": "License permits distillation",
        "basis": "permissive-open-weights",
        "sourceUrl": "https://example.com/license-or-terms",
        "note": "Provider terms do not record an output-training restriction, and the model-level policy allows distillation."
      }
    },
    {
      "provider": "cloudflare",
      "available": true
    },
    {
      "provider": "baseten",
      "available": true
    }
  ]
}
```

Notes:

* `defaultPrice` is used when you do not select a provider.
* `providers[].pricing` includes provider-selection markup. Fields ending in `Per1kTokens` are **USD per 1,000 tokens**; multiply by 1,000 to compare with `max_price`, which is **USD per 1 million tokens**.
* When present, `providers[].effectivePricing.effective_price` describes the effective rates used for routing, including current discounts and markup. Prices and availability can change between discovery and the request.
* Use the exact value from `providers[].provider` in the `X-Provider` header.
* Provider-specific `distillationPolicy` can differ from the model-level policy. Explicit provider restrictions override model-level allowances.
* If a model is unsupported, the response includes `supportsProviderSelection: false`.

## Per-Request Provider Override

For a single request, send the provider ID in the `X-Provider` header.

```bash theme={null}
curl https://api.nano-gpt.com/api/v1/chat/completions \
  -H "Authorization: Bearer $NANOGPT_API_KEY" \
  -H "Content-Type: application/json" \
  -H "X-Provider: novita" \
  -d '{
    "model": "kimi-k2.6",
    "messages": [
      { "role": "user", "content": "Hello" }
    ]
  }'
```

Notes:

* `X-Provider` is case-insensitive (`x-provider` also works).
* The provider ID must be one of the values returned by the providers endpoint.
* The body `provider` field also accepts the same provider ID as a string on `POST /api/v1/chat/completions`, `POST /api/v1/completions`, and `POST /api/talk-to-gpt`.
* Recommended: use `X-Provider` or body `provider` for explicit provider overrides. The API also accepts a trailing provider suffix for user-selectable providers, such as `model-id:cerebras`, but request fields are clearer and easier to validate against provider-discovery responses.
* To avoid Moonshot for Kimi K2.6, select any available non-Moonshot provider ID returned by the provider-discovery endpoint, such as `novita`, `cloudflare`, `baseten`, or `inceptron`.
* If the request would otherwise be subscription-covered, explicit provider selection still bypasses subscription coverage and bills the request as pay-as-you-go. You do not need to also send `billing_mode: "paygo"` or `X-Billing-Mode: paygo`.

## Provider Routing Object

`POST /api/v1/chat/completions`, `POST /api/v1/completions`, and `POST /api/talk-to-gpt` also accept a structured `provider` object for per-request routing controls.

```json theme={null}
{
  "model": "model-id",
  "provider": {
    "order": ["provider-a", "provider-b"],
    "only": ["provider-b"],
    "ignore": ["provider-c"],
    "sort": "throughput",
    "quantizations": ["fp8", "fp16"],
    "min_quantization": "fp8",
    "max_price": {
      "prompt": 0.5,
      "completion": 2
    },
    "allow_fallbacks": false,
    "require_parameters": true
  },
  "messages": [
    { "role": "user", "content": "Hello" }
  ]
}
```

Provider entries can use public NanoGPT provider IDs or accepted display names/aliases. Use `GET /api/models/:canonicalId/providers` to discover public provider IDs for a model. Do not rely on internal provider names that are not returned by public provider-discovery responses.

| Field | Type | Behavior |
| - | - | - |
| `provider.order` | string\[] | Soft preference order. NanoGPT tries these providers first when possible, but may route elsewhere. Unknown providers are ignored. |
| `provider.only` | string\[] | Hard provider pin. NanoGPT restricts routing to these providers and fails instead of routing outside the list. Unknown providers return `400` with `error.code: "provider_unknown_provider"`. |
| `provider.ignore` | string\[] | Excludes providers from routing. Unknown providers are ignored. |
| `provider.sort` | string | Routing preference. Supported values are `speed`, `throughput`, `latency`, `price`, `auto`, `none`, and `default`. `auto`, `none`, and `default` suppress stored/default routing preferences without requesting a sort. |
| `provider.quantizations` | string\[] | Hard allowlist of model weight quantization levels. Supported values are `int4`, `fp4`, `fp6`, `int8`, `fp8`, `fp16`, `bf16`, `fp32`, and `unknown`. |
| `provider.min_quantization` | string | Minimum precision floor. Supported values are `int4`, `fp4`, `fp6`, `int8`, `fp8`, `fp16`, `bf16`, and `fp32`. `unknown` is not allowed because it cannot be compared to a precision floor. |
| `provider.max_price` | object | Optional `prompt` and/or `completion` caps on provider input/output rates, including selection markup, in USD per 1 million tokens. See [price-cap limits](#price-cap-limits). |
| `provider.allow_fallbacks` | boolean | When `false`, disables cross-provider fallback for this request. |
| `provider.require_parameters` | boolean | When `true`, requires the selected provider to support the requested parameters on routes that support parameter-level capability checks. |

Examples:

```json theme={null}
{
  "provider": {
    "order": ["provider-a", "provider-b"]
  }
}
```

```json theme={null}
{
  "provider": {
    "only": ["provider-b"]
  }
}
```

```json theme={null}
{
  "provider": {
    "ignore": ["provider-c"]
  }
}
```

```json theme={null}
{
  "provider": {
    "quantizations": ["fp8"]
  }
}
```

```json theme={null}
{
  "provider": {
    "min_quantization": "fp8"
  }
}
```

```json theme={null}
{
  "provider": {
    "max_price": {
      "prompt": 0.5,
      "completion": 2
    }
  }
}
```

```json theme={null}
{
  "provider": {
    "require_parameters": true
  }
}
```

```json theme={null}
{
  "provider": {
    "order": ["provider-b"],
    "allow_fallbacks": false
  }
}
```

### Routing Semantics

* `order` is soft. Use `only` when the request must fail instead of using any provider outside the list.
* `only` is hard. It restricts routing to the listed providers and disables internal fallback outside the pinned set.
* `ignore` and `max_price` can resolve to a concrete provider. For direct provider routing, NanoGPT cannot represent "any provider except X" as a multi-provider pool. If filters like `ignore` or `max_price` are provided without `order` or `only`, NanoGPT resolves to one cheapest concrete provider that satisfies those filters. That is treated as explicit provider selection for billing and routing because the request changed the provider candidate set.
* If `max_price` is provided without `order` or an explicit `sort`, NanoGPT treats it as a request for price-aware routing among providers that satisfy the cap. `only` does not prevent this inference; an `order` does.
* Quantization filters are hard routing constraints. If no available provider satisfies the requested quantization, provider list, price, fallback, caching, and parameter constraints, the request may fail with a provider-availability error.
* Per-request controls can change routing order, but saved hard restrictions still apply. Saved provider allowlists are intersected with request allowlists, exclusions remain excluded, and a saved fallback restriction is preserved. A request cannot broaden a saved allowlist by sending `order`, `only`, or `allow_fallbacks: true`.

### Quantization Filters

Use quantization filters when you need to avoid lower-precision provider routes for a model.

`provider.quantizations` is an exact allowlist. It must be a non-empty array:

```json theme={null}
{
  "model": "deepseek/deepseek-v3-0324",
  "messages": [
    { "role": "user", "content": "Hello" }
  ],
  "provider": {
    "quantizations": ["fp8"]
  }
}
```

When an exact quantization filter is supplied, providers without matching known metadata are excluded. To allow providers whose quantization metadata is not known, include `unknown` explicitly:

```json theme={null}
{
  "provider": {
    "quantizations": ["fp8", "unknown"]
  }
}
```

`provider.min_quantization` sets a minimum precision floor:

```json theme={null}
{
  "model": "deepseek/deepseek-v3-0324",
  "messages": [
    { "role": "user", "content": "Hello" }
  ],
  "provider": {
    "min_quantization": "fp8"
  }
}
```

Minimum quantization is ordered by bit width. For example, `min_quantization: "fp8"` allows `int8`, `fp8`, `fp16`, `bf16`, and `fp32`. Integer and floating-point formats with the same bit width are treated as equivalent for minimum filtering.

You may combine exact and minimum filters. NanoGPT uses the overlap:

```json theme={null}
{
  "provider": {
    "quantizations": ["fp8", "fp16"],
    "min_quantization": "fp8"
  }
}
```

If the filters do not overlap, the request is rejected as invalid. For example, `quantizations: ["fp4"]` with `min_quantization: "fp8"` is invalid because `fp4` is below the requested minimum.

Use NanoGPT's canonical quantization values in API requests. Provider metadata may use different internal labels, but request parameters must use the supported canonical values listed above.

### Validation

Invalid object shapes return a structured `400` error:

* `provider.order`, `provider.only`, and `provider.ignore` must be arrays of strings.
* `provider.quantizations` must be a non-empty array of supported quantization strings.
* `provider.min_quantization` must be a supported comparable quantization string. `unknown` is only valid in `provider.quantizations`.
* If `provider.quantizations` and `provider.min_quantization` have no overlap, the request is invalid.
* Unknown providers in `only` return `provider_unknown_provider`.
* Unknown providers in `order` and `ignore` are ignored.
* `provider.max_price.prompt` and `provider.max_price.completion` must be non-negative numbers.
* `provider.allow_fallbacks` and `provider.require_parameters` must be booleans.

Example error:

```json theme={null}
{
  "error": {
    "message": "Unknown or unavailable provider id in provider.only: not-a-provider",
    "type": "invalid_request_error",
    "param": "provider.only",
    "code": "provider_unknown_provider"
  }
}
```

## Consecutive Prompts and Prompt Caching

When you explicitly select a provider with `X-Provider` or body `provider`, NanoGPT routes that request to the selected provider. For consecutive prompts, using the same provider helps keep routing stable, which is the setup you want for provider-side prompt-cache reuse when that provider and model support caching.

For capability-based routing, use top-level `caching: true` or the `:caching` model suffix instead of choosing a specific provider. NanoGPT will route to any available provider that supports prompt/input caching for the requested provider-selection model, and later matching requests prefer the same eligible provider through soft affinity.

General automatic routing may use cache-affinity routing for eligible cache-capable providers, but it does not require a cache-capable provider and may still choose different upstream providers when request shape, provider availability, or routing constraints change. For cache-sensitive workflows, either consistently send the same `X-Provider` value or use `caching: true` / `:caching` when any cache-capable provider is acceptable.

## Cache-Capable Provider Routing

Set `caching: true` when you want NanoGPT to route a chat completion request to any available provider that supports prompt/input caching. This is capability-based provider selection, not explicit provider selection and not prompt-cache annotation. It does not add Anthropic-style `cache_control` markers or configure cache TTLs.

```json theme={null}
{
  "model": "model-id",
  "caching": true,
  "messages": [
    { "role": "user", "content": "Hello" }
  ]
}
```

If no usable cache-capable provider exists for the model, the request fails instead of falling back to a non-cache-capable provider.

`caching: true` prefers the provider from your previous matching request (soft affinity). `stickyProvider` defaults to `false`; set it to `true` to disable cross-provider fallback. Neither setting guarantees a cache hit.

To prevent cross-provider fallback when preserving the cache route matters more than availability:

```json theme={null}
{
  "model": "model-id",
  "caching": true,
  "stickyprovider": true,
  "messages": [
    { "role": "user", "content": "Hello" }
  ]
}
```

Top-level `stickyProvider` is accepted as a camelCase alias for `stickyprovider`. Setting it to `true` disables cross-provider fallback for the caching request. Leaving it `false` allows fallback; it does **not** disable the soft affinity established by explicit caching.

Affinity expires after about one hour without a successful matching request. This is a routing preference lifetime, separate from the provider's cache TTL. Availability, exclusions, price caps, and request compatibility still apply.

The model suffixes `:caching`, `:cache`, and `:cached` request the same cache-capable provider routing:

```json theme={null}
{
  "model": "moonshotai/kimi-k2.6:thinking:caching",
  "messages": [
    { "role": "user", "content": "Hello" }
  ]
}
```

Routing order for `caching: true`:

1. Filter to providers that are available, not excluded by preferences, and marked as prompt-caching capable.
2. Prefer the previously recorded provider for the same cache-relevant request shape when still usable.
3. Try eligible per-model and global provider preferences.
4. Otherwise use provider prices and recent aggregate cache-read observations to choose a route. A provider with better observed cache reuse may be preferred within a bounded price range.

The model page's cache-hit figures are historical aggregates across requests, not a prediction of your hit rate. Cache reuse also depends on prompt-prefix stability, provider requirements, and cache TTL.

## Price Caps and Cache Reuse

Use public provider IDs from the discovery endpoint and choose caps suitable for that model. The prices below are illustrative.

### Lowest estimated cost with caps

```json theme={null}
{
  "model": "model-id",
  "provider": {
    "sort": "price",
    "max_price": { "prompt": 0.5, "completion": 2 }
  },
  "messages": [{ "role": "user", "content": "Hello" }]
}
```

The ranking uses estimated input tokens and the output budget when available, with defaults of 3,000 input and 1,000 output tokens for missing estimates. Known cache-write costs may be included, but chat routing does not assume a future cache hit or subtract a predicted cache-read discount. This is an estimate, not a guarantee of the cheapest final bill.

### One provider with caps

```json theme={null}
{
  "model": "model-id",
  "provider": {
    "only": ["provider-id"],
    "allow_fallbacks": false,
    "max_price": { "prompt": 0.5, "completion": 2 }
  },
  "messages": [{ "role": "user", "content": "Hello" }]
}
```

Price ranking is inferred here, but `only` leaves a single candidate. This request fails if the provider is unavailable, incompatible, or over the cap. Keeping the provider fixed helps cache reuse, but does not guarantee a cache hit.

### Cache reuse with automatic provider choice

```json theme={null}
{
  "model": "model-id",
  "caching": true,
  "messages": [{ "role": "user", "content": "Hello" }]
}
```

This prefers the previous eligible cache provider. Use the [prompt-caching helper](/api-reference/miscellaneous/prompt-caching) when the provider also needs explicit cache markers or a TTL.

### Combining controls

| Controls | Routing behavior |
| - | - |
| `caching: true` alone | Cache-capable selection with soft affinity. |
| `caching: true` with `stickyProvider: true` | Cache-capable selection with cross-provider fallback disabled. |
| `max_price` without `order` or an explicit sort override | Infers price ranking, which takes precedence over cache-affinity selection. |
| `caching: true` with `sort: "price"` | Price ranking takes precedence; this does not mean “cheapest once, then stay there.” |
| `caching: true` with `order` and `max_price`, without a routing sort | `order` prevents inferred price ranking; caching selection still applies the cap and saved restrictions. |

An ordered provider list is evaluated for each request. It is not a sticky failover session: failure on one request does not permanently advance the list to the next provider. In caching mode, an eligible provider recorded by affinity can take priority over the order.

### Price-cap limits

For requests with a verified provider-selection price, `max_price` restricts retries to routes whose billed base rate NanoGPT can verify against your cap. If the selected provider fails and no such route is available, the request fails instead of being billed at an unverified rate. On models without provider selection, `max_price` only filters routes; billing uses the model's price. An explicitly selected gateway without a verifiable provider price rejects the cap.

`max_price` compares effective provider input/output rates **after provider-selection markup**, so a provider with a raw input rate of $0.06/M and 5% markup requires a prompt cap of at least $0.063/M. The selection markup is already included in provider-discovery prices; do not add it again.

These are token-rate filters, not a total request budget. They do not separately cap cache-read prices, cache-write prices, search fees, or other add-ons. Local comparisons use the provider's base input/output tier, so inspect higher-context tiers separately. Unknown prices cannot be verified locally; gateway routes also pass a price constraint upstream, while direct fallback routes with unknown prices are excluded. For a predictable choice, select a provider with published pricing and disable fallback.

## Per-Request Routing Preference

For provider-selection-capable models, you can append a routing preference suffix to the model ID:

* `:speed` / `:fast`: best estimated completion time using TTFT plus TPS.
* `:throughput`: highest tokens per second.
* `:latency`: lowest time to first token.
* `:price` / `:cheap`: lowest estimated cost for the request’s input and expected output, including selection markup.
* `:floor`: alias for `:price`.
* `:tools`: route to a tools-capable provider path when the model supports tools provider selection.

Example:

```json theme={null}
{
  "model": "zai-org/glm-5:fast",
  "messages": [{ "role": "user", "content": "Hello" }]
}
```

Notes:

* Routing preference suffixes are billed like explicit provider-selection requests.
* They cannot be combined with `X-Provider`, body `provider`, or provider model suffixes.
* `:tools` cannot be combined with routing preference suffixes or `caching: true`.
* If the model does not support provider selection, the request returns an invalid request error for these suffixes.

See [Model Suffixes](/api-reference/miscellaneous/model-suffixes) for the complete suffix reference and composition rules.

## Persistent Provider Preferences

These endpoints let a user save provider preferences in their session metadata.

```
GET /api/user/provider-preferences
PATCH /api/user/provider-preferences
DELETE /api/user/provider-preferences
```

These endpoints require a logged-in web session. API-key-only requests should use per-request provider selection with `X-Provider`, unless preferences were already saved on the associated user session.

Saved provider preferences apply to pay-as-you-go traffic. For subscription-included traffic, send `X-Billing-Mode: paygo` or `billing_mode: "paygo"` when you want saved provider preferences to apply.

Example `GET` response (placeholders):

```json theme={null}
{
  "preferredProviders": ["provider-a", "provider-b"],
  "excludedProviders": ["provider-c"],
  "enableFallback": true,
  "modelOverrides": {
    "model-id": {
      "preferredProviders": ["provider-b"],
      "enableFallback": false
    }
  },
  "availableProviders": ["provider-a", "provider-b", "provider-c"]
}
```

Example `PATCH` payload:

```json theme={null}
{
  "preferredProviders": ["provider-a", "provider-b"],
  "excludedProviders": ["provider-c"],
  "enableFallback": false,
  "modelOverrides": {
    "model-id": {
      "preferredProviders": ["provider-b"],
      "enableFallback": true
    }
  }
}
```

Field details:

* `preferredProviders`: ordered list of allowed providers; the system tries each in order.
* `excludedProviders`: providers that should never be used.
* `enableFallback`: when `true` (default), permits retries among eligible providers. A configured `preferredProviders` list remains a hard allowlist; fallback does not escape it to the platform default.
* `modelOverrides`: optional per-model overrides for the fields above.
* `availableProviders`: full set of provider IDs available to your account for the model.

## Resolution Order

When `caching: true` is set without a routing preference, use the routing order in [Cache-Capable Provider Routing](#cache-capable-provider-routing). Otherwise, when multiple selections exist, the system resolves providers in this order:

1. Per-request explicit provider selection (`X-Provider`, body `provider` string/object, or provider suffix)
2. Per-request routing preference suffix (`:fast`, `:cheap`, etc.)
3. Per-model `preferredProviders`
4. Global `preferredProviders`
5. Platform default when no binding provider allowlist or request constraint selects a route

## Billing and Markup

* If you explicitly select or constrain providers with `X-Provider` or body `provider`, billing uses the resolved provider-specific price plus the applicable provider-selection markup (normally 5%) and the request is treated as pay-as-you-go.
* If saved preferences select a provider on pay-as-you-go traffic, billing uses the provider-specific price plus the applicable provider-selection markup (normally 5%).
* If you do not select a provider, billing uses the model's default price.

## Which Provider Served My Request?

The [Usage page](https://nano-gpt.com/usage) shows a label such as **Via Cerebras** under the model name, and a Provider row in request details, when public attribution is available. This can include explicit selections, supported BYOK routes, or a public provider used as a fallback. Some automatic requests have no provider label.

Chat responses and the [request billing endpoint](/api-reference/endpoint/request-billing) do not expose the upstream provider. Do not rely on an undocumented response header or infer a provider from timing or price.

## Error Behavior

* `400` with `error.code: "provider_max_price_unverifiable"` means an explicitly selected gateway has no verifiable provider price for the model. Choose a provider with published pricing, or remove the cap to use normal gateway billing.

* If `enableFallback` is `false` and no preferred provider is available, `/api/v1/chat/completions` returns `400` with `error.code: "no_fallback_available"`.

* Invalid body `provider` objects return `400` with `error.type: "invalid_request_error"`.

* `PATCH /api/user/provider-preferences` validates provider IDs and returns:
  * `422 INVALID_INPUT` for malformed payloads.
  * `400 INVALID_EXCLUSIONS` if exclusions would leave a model with no usable provider.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.