These examples use the website API host, which currently lists their recommended models. The direct API host can have a different catalog during rollouts. For larger requests or longer runtimes, use
https://api.nano-gpt.com/api/v1 with a model from that host’s catalog. See API Hosts for limits and availability checks.OpenCode Integration
Use NanoGPT with OpenCode, the open source AI coding agent for your terminal.Requirements
- OpenCode installed (
curl -fsSL https://opencode.ai/install | bash) - A NanoGPT account with API access
One-Click Setup (Mac/Linux)
Run this command in your terminal:- Open your browser to authenticate with NanoGPT (or let you paste an API key)
- Configure OpenCode to use NanoGPT as your AI provider
- Add starter model definitions while preserving existing model selections, filters, and connection options
opencode.json; if you use opencode.jsonc, follow the manual setup below and merge the settings into your existing file.
Setup replaces NanoGPT authentication with your chosen login or pasted key. It removes provider-level options.apiKey and Authorization overrides in provider and model headers from opencode.json so the new credentials in auth.json take effect, while preserving unrelated headers and custom model settings.
Manual Setup
If you prefer to configure manually:1. Create the auth file
Add your NanoGPT API key to~/.local/share/opencode/auth.json. If the file already exists, merge the nano-gpt entry into it and keep your other authentication entries:
2. Create the config file
Create~/.config/opencode/opencode.json, or merge these settings into your existing config:
models block is optional for models already in that catalogue. Use it to add a missing model or override model settings, as described in Adding More Models.
OpenCode also supports opencode.jsonc for configs with comments. The examples here use strict JSON.
Usage
After setup, start OpenCode:Switching Models
Use the/models command inside OpenCode to switch between models, or edit the model field in your opencode.json:
Available Models
Examples of NanoGPT models are listed below. OpenCode’s built-in catalogue and your custom definitions determine which models appear in/models; the list is not limited to entries in your models block.
Adding More Models
Themodels block adds model definitions or overrides settings for existing models. It does not hide models that you have not listed.
If a model already appears in OpenCode and you do not need custom settings, you can omit its entry. Keep entries for missing models or custom names, options, variants, and context/output limits.
To add or override a model, merge its definition into provider.nano-gpt.models:
GET https://nano-gpt.com/api/v1/models. OpenCode’s catalogue can lag behind the API, so a model missing from OpenCode may need a manual definition. A model appearing in the API list does not necessarily support the tool calls needed by OpenCode.
Limiting the Model List
To show only selected NanoGPT models, addwhitelist under provider.nano-gpt. Merge this example into your existing config, keeping your connection settings and any model overrides:
nano-gpt/ prefix inside whitelist, blacklist, and models. The top-level model field uses the full nano-gpt/model-id format.
Alternatively, blacklist hides specific models while keeping the rest. If both filters are set, a model must be in whitelist and must not be in blacklist. Keep the selected default model inside your allowed list, then restart OpenCode after saving.
Claude Prompt Caching
OpenCode adds explicitcache_control markers to Claude requests automatically. Its markers omit the TTL and use the five-minute default unless you request a different TTL. You do not need a promptCaching block or an anthropic-beta header just to use OpenCode’s default Claude caching.
A cache marker requests caching; it does not guarantee a cache hit. The prompt must meet the selected model’s minimum cacheable size, and later requests must reuse the same prefix before the cache expires.
On NanoGPT’s Chat Completions endpoint, supported explicit cache controls use pay-as-you-go billing, including markers OpenCode adds automatically. They bypass subscription coverage for that request. See subscription billing for prompt caching.
One-Hour Caching
To request one-hour caching for Opus 5.5, merge this configuration into your existingopencode.json. Keep your NanoGPT authentication in auth.json:
options.promptCaching as the top-level promptCaching field in the API request. NanoGPT applies the requested TTL to the cache boundaries, including OpenCode’s existing markers. No additional anthropic-beta header is required.
One-hour caching has a higher cache-write price than five-minute caching on standard Claude cache pricing. Choose it when you expect to reuse a large prefix after pauses longer than five minutes. See cache pricing.
The example leaves stickyProvider at its default, false, so NanoGPT can fail over when a service is unavailable. Adding "stickyProvider": true inside promptCaching disables failover and can return a 503 instead. It is optional and does not enable caching by itself.
Which Model Settings Should I Keep?
Removing the entire
options block also removes your caching and request-level privacy settings. Removing variants only removes your custom effort choices.
Opus 5.5 Reasoning Effort
Opus 5.5 uses adaptive thinking automatically. You do not needreasoning: true or custom variants to enable its thinking. Supported effort values are low, medium, high, xhigh, and max; none cannot disable this model’s thinking.
If your OpenCode version does not offer the choices you want, merge these variants into provider.nano-gpt.models["anthropic/claude-opus-5.5"], alongside the existing options:
reasoningEffort to NanoGPT’s reasoning_effort request field. See reasoning controls for the distinction between generating thinking and displaying it.
The snippet includes reasoning: true so OpenCode can expose reasoning choices even with missing or stale catalogue metadata. You can omit it when the catalogue already describes the model correctly.
Checking Cache Hits
Look forcache_creation_input_tokens on an eligible write and cache_read_input_tokens > 0 on a subsequent request that reuses the prefix. A successful answer alone does not establish a cache hit or the requested TTL. Raw API responses can also include x_nanogpt_pricing.cacheTTL for the billing TTL.
If your OpenCode version does not display these fields, use the API cache verification examples. A short prompt can return zero cached tokens even with a correct configuration.
The request format above was checked with OpenCode 1.14.48 and @ai-sdk/openai-compatible. Other SDKs or older client versions can map options differently.
Troubleshooting
OpenCode not finding the config? Useopencode.json or opencode.jsonc, without a leading dot:
- Global config:
~/.config/opencode/opencode.json, or$XDG_CONFIG_HOME/opencode/opencode.jsonwhenXDG_CONFIG_HOMEis set. - Project config:
opencode.jsonin your project.
OPENCODE_CONFIG can add another source. See OpenCode’s config documentation for the full precedence rules.
More models appear than expected?
Check that whitelist is inside provider.nano-gpt, not inside models. Defining a models block does not filter the catalogue. Also check project configs and opencode.jsonc for settings that override your global filters.
The installer cannot merge your config?
Install Python 3 when updating existing OpenCode files. The installer reads strict JSON only; if your config contains comments, use manual setup in opencode.jsonc. The installer stops before authentication or file updates if existing auth or JSON config files cannot be read or parsed.
Authentication errors?
Check that your API key is correctly set in ~/.local/share/opencode/auth.json (as an object with type and key) and that the key name matches the provider ID (nano-gpt).
Claude caching is not being used?
If you configured a TTL override, check that promptCaching is inside the selected model’s options, that enabled is true, and that the TTL is exactly "5m" or "1h". Keep reusable instructions and tool definitions at the beginning of the prompt. Changes to that prefix, insufficient prompt length, expiration, or failover can prevent a cache hit.
A request returns “temporarily unavailable”?
If you set promptCaching.stickyProvider: true, a 503 can mean failover was prevented to preserve the cache. Privacy and provider restrictions can also reduce the available routes. Keep intentional privacy requirements; a service error alone does not establish that those requirements caused the failure. Retry, and send support the new request ID, model ID, and approximate time if the error continues.