Errors
Every refusal uses one envelope, whatever produced it.
{
"error": {
"code": "NOT_FOUND",
"type": "TargetNotServingError",
"message": "the resource you targeted is not serving this model",
"status": 404,
"path": "/v1/chat/completions"
}
}Switch on error.type. It is the stable value. status and code are coarser — three different refusals share 404 — and message is prose we may reword.
Ten types reach you, in four families: a key that was not accepted, a request that cannot be carried, a model that would not run where you sent it, and the limits your plan carries. Every one of them is covered below.
error.type | Status | Retry |
|---|---|---|
Unauthorized | 401 | No |
UnsupportedShapeError | 400 | No |
InferenceUnavailableError | 503 | Yes, with backoff |
ModelNotInPlanError | 403 | No |
DedicatedOnlyModelError | 404 | No |
TargetNotServingError | 404 | No |
InferenceUpsellError | 409 | No |
ContextWindowExceededError | 413 | No |
InferenceRateLimitedError | 429 | Yes, after Retry-After |
InferenceConcurrencyLimitedError | 429 | Yes, when one finishes |
The key was not accepted
Unauthorized
401 · not retryable.
The bearer is missing, malformed, revoked, expired, or simply not ours. All of those answer identically, so a 401 tells you the key was not accepted and nothing more. Check the key and the Authorization header.
The request itself
UnsupportedShapeError
400 · not retryable · fix the request.
The request carries a field the API shape you called cannot hold. It is refused rather than answered with the field silently dropped — a request answered with one of its instructions discarded was answered as though it had said something else. Send the field on the shape that supports it, or leave it out; the surface sets out what each shape carries.
Where the request runs
These four come from the router behind the gateway and reach you unchanged, so the type you receive is the one the router chose. Each is about which model runs where — not about your key, and not about numbers your plan carries.
InferenceUnavailableError
503 · retryable · exponential backoff, starting around a second.
The model exists in the catalog and your plan covers it, but no node is answering for it right now. This is the one genuinely transient refusal the router returns.
ModelNotInPlanError
403 · not retryable.
A shared node is serving the model and would answer — your plan is what stands in the way. Retrying will never help. This is also where an org that holds a valid key but no inference entitlement lands: if you are debugging access, the status tells you which half failed — 401 is the key, 403 is the plan.
DedicatedOnlyModelError
404 · not retryable.
The model is in no shared pool and never will be. The remedy is to install it on a node you own, after which it is reachable on that stack's hostname.
This is the most misread of the four, because 404 reads as "no such model". The model exists; the shared pool is simply not where you can reach it. GET /v1/models on the shared host lists only what the pool serves, so a model you can call on your stack's hostname and cannot see on the shared one is this case.
TargetNotServingError
404 · not retryable.
You addressed a specific stack — see Routing — and it did not answer. Two different situations produce this, and they are identical on the wire: either the stack is yours and is not serving the model you asked for, or the stack is not yours at all. We do not distinguish them, deliberately: a distinguishable refusal would let anyone discover which stack identifiers exist under other customers by watching which one they got.
So check both: that the hostname is the stack you meant, and that the stack is serving the model you named. A pin never falls back to the shared pool, so a model your stack is not running fails here rather than being answered by someone else's hardware.
Your plan's limits
These four are settled by the gateway before routing, and all four name the tier in the message. Telling the two 429s apart is the point of switching on the type rather than the status: one asks you to wait, the other to run fewer requests at once. The limits themselves come with your plan — see Plans and tiers.
No X-RateLimit-* headers are returned. Retry-After on the two 429s is the only timing signal you get, so a client that wants to stay under a ceiling has to track its own rate.
InferenceUpsellError
409 · not retryable · the allowance refills when the period rolls.
You have used the allowance your tier carries — monthly tokens for a paying organisation, the free turn count for an anonymous visitor. There is no Retry-After, because the refill is a period away rather than a window away, and the message routes you instead: an anonymous caller is asked to sign in, a paying one is named the next rung, and a Platinum one is asked to contact us.
ContextWindowExceededError
413 · not retryable.
The request needs more context than your tier allows. This is a property of the request, not of what you have spent: the same request will be too large tomorrow. Neither retrying nor waiting helps — send less context, or move to a tier with a wider window.
InferenceRateLimitedError
429 · retryable · wait the seconds in Retry-After, then retry.
You sent more requests in a minute than the tier carries. The refusal answers with Retry-After in seconds, the remainder of the window — that long, and no longer.
InferenceConcurrencyLimitedError
429 · retryable · retry once one of your in-flight requests finishes.
You have as many requests in flight at once as the tier carries. Waiting out a window does not help — a slot frees when one of your own requests finishes, which is why Retry-After is a second rather than a deadline. Reduce your parallelism instead of backing off. Some models count more than one request against this allowance, and the message says so when they do.
When a stream dies partway
A stream that ends before it is finished — the connection dropped, the request failed, or you cancelled it — closes without a terminal frame, and whatever already arrived has been delivered. Usage is recorded when a stream completes, so one that dies partway records none: you are not billed for a turn that never finished.