Overview
Some models generate a separate reasoning stream (sometimes called thinking) in addition to the final answer content. NanoGPT exposes this in an OpenAI-compatible way for Chat Completions:- Streaming (SSE):
choices[0].delta.reasoning(or legacychoices[0].delta.reasoning_content) - Non-streaming:
choices[0].message.reasoning(or legacychoices[0].message.reasoning_content)
:thinking is model-specific and only works when that exact ID (or a documented alias) exists.
-thinking is a legacy alias pattern for some model families only, not universal.
Do not assume -thinking works for arbitrary model IDs. Always check GET /api/v1/models for exact valid IDs.
Endpoint Variants (Chat Completions)
NanoGPT provides three base paths for chat completions. All accept the same request format and model names, but differ in how reasoning content is delivered:
This is the same behavior documented under Reasoning Streams on the Chat Completion page.
Controlling Reasoning Output
Hide Reasoning
To strip reasoning from the response (both streaming and non-streaming), send:reasoning.exclude controls output visibility. It is not the same as disabling reasoning compute.
Or append the model suffix:
:reasoning-exclude
Reasoning Effort
reasoning_effort (or reasoning.effort) controls reasoning depth and also acts as an explicit reasoning-mode signal.
An accepted value other than none requests reasoning/thinking behavior. Available levels and defaults depend on the model.
Use none to request disabled reasoning only on models that support turning it off. Always-thinking models keep their reasoning requirements; Claude Opus 5.5 rejects attempts to disable thinking.
reasoning_effort is authoritative for Chat Completions request shaping.
NanoGPT accepts none, minimal, low, medium, high, xhigh, and max. Not every model supports every level; some routes map an accepted level to the nearest supported effort. max is a separate value from xhigh where the model supports both.
Claude Opus 5.5
anthropic/claude-opus-5.5 uses adaptive thinking by default. You do not need a :thinking suffix or an explicit enable flag. Its native effort values are low, medium, high, xhigh, and max; NanoGPT defaults to high when no effort or thinking budget is supplied.
reasoning_effort: "none" or thinking.enabled: false return a reasoning_required validation error. To hide returned reasoning while preserving thinking, use reasoning.exclude: true instead.
Adaptive thinking also requires compatible tool choice. Use tool_choice: "auto" or "none"; forcing a tool with "required" or a named tool is unsupported for Opus 5.5.
In OpenCode, model-level reasoning: true describes a capability to the client. A variant’s reasoningEffort sets the API effort; these settings serve different purposes.
Change Effort Within a Conversation
For GPT-6 Astra and Claude Fable 5.1, use inline reasoning effort updates to change effort while keeping the request-level baseline and earlier history unchanged. Updates can precede a user turn or appear at the end of a request to control its generated answer. Chat Completions accepts emptysystem or developer update messages; Responses and Messages have equivalent forms.
Legacy Field Name Compatibility (reasoning_content)
If a client expects reasoning_content instead of reasoning, you can:
- Use
/api/v1legacy/chat/completions, or - Set
reasoning.delta_field: "reasoning_content", or - Use the shorthands
reasoning_delta_field/reasoning_content_compat.
How Reasoning Appears in Responses
Streaming (SSE)
Reasoning deltas appear before or alongside content deltas:Non-Streaming
The final message may include a separate reasoning field:Cost Notes
Reasoning tokens are billed as output tokens. If you enable higher reasoning effort (or use thinking variants), expect highercompletion_tokens and higher cost.
See Also
- Chat Completions: reasoning controls and endpoint variants (Chat Completion)
- Streaming protocol details across endpoints (Streaming Protocol)