openrouter/chat-completions, served by Comfy Router from Openrouter.
Request setup
Create a key in your Comfy workspace and export it asCOMFY_API_KEY. For Python, run pip install comfy-sdk. For TypeScript, run npm install @comfyorg/sdk. cURL uses raw HTTP.
Model ID: openrouter/chat-completions
Endpoint: POST https://api.comfy.org/v2/models/openrouter/chat-completions
This model has no runnable request example. Build the body from the input documentation below, then use it with the Router quickstart.
Schema
Input
object
Enable automatic prompt caching. When set at the top level, the system automatically applies cache breakpoints to the last cacheable block in the request. Currently supported for Anthropic Claude models.
string
Possible values:
5m, 1hstring
required
Possible values:
ephemeralobject
Debug options for inspecting request transformations (streaming only)
boolean
If true, includes the transformed upstream request body in a debug chunk at the start of the stream. Only works with streaming mode.
number
Frequency penalty (-2.0 to 2.0)Format:
doubleobject | string | number | object[]
object
Token logit bias adjustments
boolean
Return log probabilities
integer
Maximum tokens in completion
integer
Maximum tokens (deprecated, use max_completion_tokens). Note: some providers enforce a minimum of 16.
object[]
required
List of messages for the conversation
object
Key-value pairs for additional object information (max 16 pairs, 64 char keys, 512 char values)
`text`, `image`, `audio`[]
Output modalities for the response. Supported values are “text”, “image”, and “audio”.
string
Model to use for completion
string[]
Models to use for completion
boolean
Whether to enable parallel function calling during tool use. When true, the model may generate multiple tool calls in a single response.
object[]
Plugins you want to enable for this request, including their settings.
number
Presence penalty (-2.0 to 2.0)Format:
doubleobject
When multiple model providers are available, optionally indicate your routing preference.
boolean
Whether to allow backup providers to serve requests
- true: (default) when the primary provider (or your custom providers in “order”) is unavailable, use the next best provider.
- false: use only the primary/custom provider, and return the upstream error if it’s unavailable.
`deny`, `allow`
Data collection setting. If no available model provider meets the requirement, your request will return an error.
- allow: (default) allow providers which store user data non-transiently and may train on it
- deny: use only providers which do not collect user data.
boolean
Whether to restrict routing to only models that allow text distillation. When true, only models where the author has allowed distillation will be used.
`AkashML`, `AI21`, `AionLabs`, `Alibaba`, `Ambient`, `Baidu`, `Amazon Bedrock`, `Amazon Nova`, `Anthropic`, `Arcee AI`, `AtlasCloud`, `Avian`, `Azure`, `BaseTen`, `BytePlus`, `Black Forest Labs`, `Cerebras`, `Chutes`, `Cirrascale`, `Clarifai`, `Cloudflare`, `Cohere`, `Crucible`, `Crusoe`, `DeepInfra`, `DeepSeek`, `DekaLLM`, `Featherless`, `Fireworks`, `Friendli`, `GMICloud`, `Google`, `Google AI Studio`, `Groq`, `Hyperbolic`, `Inception`, `Inceptron`, `InferenceNet`, `Ionstream`, `Infermatic`, `Io Net`, `Inflection`, `Liquid`, `Mara`, `Mancer 2`, `Minimax`, `ModelRun`, `Mistral`, `Modular`, `Moonshot AI`, `Morph`, `NCompass`, `Nebius`, `Nex AGI`, `NextBit`, `Novita`, `Nvidia`, `OpenAI`, `OpenInference`, `Parasail`, `Poolside`, `Perceptron`, `Perplexity`, `Phala`, `Recraft`, `Reka`, `Relace`, `SambaNova`, `Seed`, `SiliconFlow`, `Sourceful`, `StepFun`, `Stealth`, `StreamLake`, `Switchpoint`, `Together`, `Upstage`, `Venice`, `WandB`, `Xiaomi`, `xAI`, `Z.AI`, `FakeProvider` | string[]
List of provider slugs to ignore. If provided, this list is merged with your account-wide ignored provider settings for this request.
object
The object specifying the maximum price you want to pay for this request. USD price per million tokens, for prompt and completion.
string
Price per million prompt tokens
string
Price per million prompt tokens
string
Price per million prompt tokens
string
Price per million prompt tokens
string
Price per million prompt tokens
`AkashML`, `AI21`, `AionLabs`, `Alibaba`, `Ambient`, `Baidu`, `Amazon Bedrock`, `Amazon Nova`, `Anthropic`, `Arcee AI`, `AtlasCloud`, `Avian`, `Azure`, `BaseTen`, `BytePlus`, `Black Forest Labs`, `Cerebras`, `Chutes`, `Cirrascale`, `Clarifai`, `Cloudflare`, `Cohere`, `Crucible`, `Crusoe`, `DeepInfra`, `DeepSeek`, `DekaLLM`, `Featherless`, `Fireworks`, `Friendli`, `GMICloud`, `Google`, `Google AI Studio`, `Groq`, `Hyperbolic`, `Inception`, `Inceptron`, `InferenceNet`, `Ionstream`, `Infermatic`, `Io Net`, `Inflection`, `Liquid`, `Mara`, `Mancer 2`, `Minimax`, `ModelRun`, `Mistral`, `Modular`, `Moonshot AI`, `Morph`, `NCompass`, `Nebius`, `Nex AGI`, `NextBit`, `Novita`, `Nvidia`, `OpenAI`, `OpenInference`, `Parasail`, `Poolside`, `Perceptron`, `Perplexity`, `Phala`, `Recraft`, `Reka`, `Relace`, `SambaNova`, `Seed`, `SiliconFlow`, `Sourceful`, `StepFun`, `Stealth`, `StreamLake`, `Switchpoint`, `Together`, `Upstage`, `Venice`, `WandB`, `Xiaomi`, `xAI`, `Z.AI`, `FakeProvider` | string[]
List of provider slugs to allow. If provided, this list is merged with your account-wide allowed provider settings for this request.
`AkashML`, `AI21`, `AionLabs`, `Alibaba`, `Ambient`, `Baidu`, `Amazon Bedrock`, `Amazon Nova`, `Anthropic`, `Arcee AI`, `AtlasCloud`, `Avian`, `Azure`, `BaseTen`, `BytePlus`, `Black Forest Labs`, `Cerebras`, `Chutes`, `Cirrascale`, `Clarifai`, `Cloudflare`, `Cohere`, `Crucible`, `Crusoe`, `DeepInfra`, `DeepSeek`, `DekaLLM`, `Featherless`, `Fireworks`, `Friendli`, `GMICloud`, `Google`, `Google AI Studio`, `Groq`, `Hyperbolic`, `Inception`, `Inceptron`, `InferenceNet`, `Ionstream`, `Infermatic`, `Io Net`, `Inflection`, `Liquid`, `Mara`, `Mancer 2`, `Minimax`, `ModelRun`, `Mistral`, `Modular`, `Moonshot AI`, `Morph`, `NCompass`, `Nebius`, `Nex AGI`, `NextBit`, `Novita`, `Nvidia`, `OpenAI`, `OpenInference`, `Parasail`, `Poolside`, `Perceptron`, `Perplexity`, `Phala`, `Recraft`, `Reka`, `Relace`, `SambaNova`, `Seed`, `SiliconFlow`, `Sourceful`, `StepFun`, `Stealth`, `StreamLake`, `Switchpoint`, `Together`, `Upstage`, `Venice`, `WandB`, `Xiaomi`, `xAI`, `Z.AI`, `FakeProvider` | string[]
An ordered list of provider slugs. The router will attempt to use the first provider in the subset of this list that supports your requested model, and fall back to the next if it is unavailable. If no providers are available, the request will fail with an error message.
number | object
Preferred maximum latency (in seconds). Can be a number (applies to p50) or an object with percentile-specific cutoffs. Endpoints above the threshold(s) may still be used, but are deprioritized in routing. When using fallback models, this may cause a fallback model to be used instead of the primary model if it meets the threshold.
number | object
Preferred minimum throughput (in tokens per second). Can be a number (applies to p50) or an object with percentile-specific cutoffs. Endpoints below the threshold(s) may still be used, but are deprioritized in routing. When using fallback models, this may cause a fallback model to be used instead of the primary model if it meets the threshold.
`int4`, `int8`, `fp4`, `fp6`, `fp8`, `fp16`, `bf16`, `fp32`, `unknown`[]
A list of quantization levels to filter the provider by.
boolean
Whether to filter providers to only those that support the parameters you’ve provided. If this setting is omitted or set to false, then providers will receive only the parameters they support, and ignore the rest.
`price`, `throughput`, `latency`, `exacto` | object
The sorting strategy to use for this request, if “order” is not specified. When set, no load balancing is performed.
boolean
Whether to restrict routing to only ZDR (Zero Data Retention) endpoints. When true, only endpoints that do not retain prompts will be used.
object
Configuration options for reasoning models
`xhigh`, `high`, `medium`, `low`, `minimal`, `none`
Constrains effort on reasoning for reasoning models
string
Possible values:
auto, concise, detailedobject
Response format configuration
object
Any type
integer
Random seed for deterministic outputs
`auto`, `default`, `flex`, `priority`, `scale`
The service tier to use for processing this request.
string
A unique identifier for grouping related requests (e.g., a conversation or agent workflow) for observability. If provided in both the request body and the x-session-id header, the body value takes precedence. Maximum of 256 characters.
string | string[] | object
Stop sequences (up to 4)
object[]
Stop conditions for the server-tool agent loop. Any condition firing halts the loop (OR logic). When set, this overrides
max_tool_calls.boolean
default:"false"
Enable streaming response
object
Streaming configuration options
boolean
Deprecated: This field has no effect. Full usage details are always included.
number
Sampling temperature (0-2)Format:
double`none` | `auto` | `required` | object
Tool choice configuration
object[]
Available tools for function calling
integer
Number of top log probabilities to return (0-20)
number
Nucleus sampling parameter (0-1)Format:
doubleobject
Metadata for observability and tracing. Known keys (trace_id, trace_name, span_name, generation_name, parent_span_id) have special handling. Additional keys are passed through as custom metadata to configured broadcast destinations.
string
string
string
string
string
string
Unique user identifier
GET /v2/models/openrouter/chat-completions/openapi.json, the same document it validates a call against before the request reaches the provider.
Output
object[]
required
List of completion choices
string
required
Possible values:
tool_calls, stop, length, content_filter, errorinteger
required
Choice index
object
Log probabilities for the completion
object[]
required
Log probabilities for content tokens
integer[]
required
UTF-8 bytes of the token
number
required
Log probability of the tokenFormat:
doublestring
required
The token
object[]
required
Top alternative tokens with probabilities
integer[]
required
number
required
Format:
doublestring
required
object[]
Log probabilities for refusal tokens
integer[]
required
UTF-8 bytes of the token
number
required
Log probability of the tokenFormat:
doublestring
required
The token
object[]
required
Top alternative tokens with probabilities
integer[]
required
number
required
Format:
doublestring
required
object
required
Assistant message for requests and responses
object
Audio output data or reference
string
Base64 encoded audio data
integer
Audio expiration timestamp
string
Audio output identifier
string
Audio transcript
string | object[] | object
Assistant message content
object[]
Generated images from image generation models
object
required
string
required
URL or base64-encoded data of the generated image
string
Optional name for the assistant
string
Reasoning output
object[]
Reasoning details for extended thinking models
string
Refusal message if content was refused
object[]
Tool calls made by the assistant
object
required
string
required
Function arguments as JSON string
string
required
Function name to call
string
required
Tool call identifier
string
required
Possible values:
functioninteger
required
Unix timestamp of creation
string
required
Unique completion identifier
string
required
Model used for completion
string
required
Possible values:
chat.completionobject
integer
required
object[]
string
required
string
required
integer
required
object
required
object[]
required
string
required
string
required
boolean
required
integer
required
boolean
required
object
number
Format:
doublenumber
Format:
doublestring
object[]
number
Format:
doubleobject
string
string
string
required
string
string
required
Categorical kind of a pipeline stage. Multiple plugins can share a type (e.g. all guardrail-level plugins emit
guardrail); the name field disambiguates which plugin emitted it.Possible values: guardrail, plugin, server_tools, response_healing, context_compressionstring
required
string
required
string
required
Possible values:
direct, auto, free, latest, alias, fallback, pareto, bodybuilder, fusionstring
required
string
The service tier used by the upstream provider for this request
string
required
System fingerprint
object
Token usage statistics
integer
required
Number of tokens in the completion
object
Detailed completion token usage
number
Cost of the completionFormat:
doubleobject
Breakdown of upstream inference costs
number
required
Format:
doublenumber
Format:
doublenumber
required
Format:
doubleboolean
Whether a request was made using a Bring Your Own Key configuration
integer
required
Number of tokens in the prompt
object
Detailed prompt token usage
integer
required
Total number of tokens
Examples
Output
Before you ship
The SDKs create anIdempotency-Key and reuse it for automatic retries. For manual retries, reuse the original key. Router can hold the connection for up to 10 minutes.
When a request fails, Router sends an X-Comfy-Error-Type response header explaining why. A 422 means Router rejected the input before calling the provider, and a 413 means the request body was larger than Router accepts. Download generated assets promptly because result URLs can expire.
Any size limit named in a field description above is the provider’s own bound on that field, quoted from the provider’s specification. Router applies a separate cap to the whole request body, which base64-encoded media counts against: see request body size.
Headers
Authentication, idempotency, request IDs, error buckets, retry pacing, spend limits.
Using the Router API
Model discovery, validation errors, retries, and billing.
Limitations
What Router does not do today, and what to use instead.