Skip to main content
If you already track spend with LiteLLM Proxy, you can add Comfy Router as one more pass-through endpoint. Your callers keep their LiteLLM keys, and your Comfy API key stays on the proxy. Router reports what a synchronous run cost, in US dollars, in the x-litellm-response-cost response header, for the models that report a Comfy credit price per run (see Runs that don’t report a cost). LiteLLM reads that header on every pass-through response and records it as the request’s spend, so you don’t maintain a price table for those models.
LiteLLM’s built-in fal_ai/ provider covers image generation only. To route a general workload through LiteLLM, use a pass-through endpoint as shown here, not a native provider.

Configure the pass-through

1

Create a Comfy API key

Create a key in your Comfy workspace and make it available to the LiteLLM Proxy process:
2

Add Router to config.yaml

config.yaml
  • include_subpath: true forwards /comfy/v2/models/{provider}/{model} to https://api.comfy.org/v2/models/{provider}/{model}, so one pair of entries covers every Router model.
  • Keep path and target scoped to /v2/models. Pointing the pass-through at the API root would let any caller with a LiteLLM key use your Comfy key on every other Comfy API route, outside LiteLLM’s spend tracking.
  • methods splits the two entries on the same path. It needs a LiteLLM Proxy release whose pass-through endpoints support methods; this page was checked against v1.103.0.
  • X-API-Key carries your comfyui- key. Router also accepts the same key as Authorization: Bearer comfyui-....
  • Leave forward_headers off. When it’s on, LiteLLM forwards every incoming header to Router, including the caller’s LiteLLM key. With it off, callers pass an Idempotency-Key as x-pass-Idempotency-Key: LiteLLM strips the x-pass- prefix and forwards it. A plain Idempotency-Key header doesn’t reach Router, so a retry would run and bill a second generation.
  • Set timeout above Router’s run deadline, so Router can return its own 504 and request ID before LiteLLM gives up. The deadline is 10 minutes by default, the same as LiteLLM’s default pass-through timeout of 600 seconds.
3

Call Router through the proxy

Callers authenticate with their LiteLLM key. The body is the model’s own input, exactly as in the Comfy Router reference:

How spend is recorded

A synchronous run that reports a cost carries these response headers:
X-Comfy-Credits-Used is the run’s price in Comfy credits, rounded to two decimal places. X-Litellm-Response-Cost is the same price converted to USD from the unrounded credit amount, so dividing the displayed credits by the conversion rate can differ from it in the last digits. Header names are case-insensitive, so LiteLLM reads it as x-litellm-response-cost. LiteLLM records it exactly as sent and uses it instead of your cost_per_request. x-litellm-total-tokens is always 0, because media runs aren’t metered in tokens. As with X-Comfy-Credits-Used, the header reports a price rather than a settled charge. Reconcile against workspace billing, not against LiteLLM’s totals.

Runs that don’t report a cost

Some runs don’t carry x-litellm-response-cost. It’s absent in the same cases as X-Comfy-Credits-Used:
  • a run billed to your own provider key rather than to your Comfy credits
  • a model that isn’t yet on Router’s list of models that report a per-run credit price. That list is conservative, and some providers aren’t on it at all
  • a charge that didn’t reach Comfy billing
  • an error response
It’s also absent when Router replays a stored result for a retry with the same Idempotency-Key (sent through the proxy as x-pass-Idempotency-Key). That retry isn’t charged again, so the replay doesn’t report the run’s exact cost a second time. When the header is absent, LiteLLM falls back to the endpoint’s cost_per_request. That applies to every case above. Error responses, runs billed to your own provider key, and idempotent replays are therefore recorded at your estimate even though Comfy didn’t charge your credits for them. Set it to a sensible estimate of your typical run rather than 0, so a run that doesn’t report a cost doesn’t appear free in your spend reports, and expect those rows to push LiteLLM’s totals above your Comfy bill. A run that was priced and genuinely cost nothing reports x-litellm-response-cost: 0, which LiteLLM records as zero.

Queued delivery

With queued delivery, one generation reaches LiteLLM as several requests, and each is its own spend row: The submit response’s status_url, response_url and cancel_url are absolute https://api.comfy.org/... URLs, and LiteLLM doesn’t rewrite response bodies. Call them through the proxy by replacing https://api.comfy.org with your proxy’s /comfy prefix, for example http://localhost:4000/comfy/v2/models/.... Sent straight to Comfy, your LiteLLM key is rejected. Queue responses don’t carry x-litellm-response-cost yet, so a queued generation is recorded at your estimate, not at its exact price. Splitting the endpoint by method books that estimate once, on the submit, instead of once per poll. For exact per-run spend in LiteLLM, use the synchronous route.

Notes

  • LiteLLM records spend but doesn’t change billing. Comfy bills your workspace exactly as it would without the proxy.
  • Full request and response schemas for every model are in the Comfy Router reference. Deadlines, retries, and other limits are in Capabilities and limits.