Routing
Every inference request has to land on a node. If you own dedicated hardware, you usually care which one. This page is how that is decided, and how to decide it yourself.
How a request is resolved
Given an authenticated request naming a model, in order:
- Dedicated capacity you own. If your organisation owns a node currently serving that model, it answers. Always. Owning capacity for a model is never overridden by shared capacity for the same one.
- The shared pool, if a pooled node serves the model and your plan includes it.
- Nothing. The request is refused. It is never queued and never silently substituted.
That is the whole algorithm. There is no load balancing across the two, no preference you can express for the pool over your own hardware, and no third tier.
Addressing one stack
Each stack has its own hostname, and a request to it is a request to that stack and nothing else:
https://<stack-eid>.api.stackyak.aiThe label is your stack's identifier, exactly as shown on the stack's page in the product —
something like stk-amber-moose-9m2npq4rstv0. Use it as the base URL and everything else about
the request is unchanged:
curl https://stk-amber-moose-9m2npq4rstv0.api.stackyak.ai/v1/chat/completions \
-H "Authorization: Bearer yki_your_key_here" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-30b-a3b",
"messages": [{"role": "user", "content": "hello"}]
}'There is no routing header. The hostname is the entire mechanism, which is what makes it usable from a stock client library: set the base URL and you are pinned, with nothing to thread through per-request options.
The hostname also decides which API the endpoint speaks. A dedicated stack serves OpenAI by default and can be switched to Anthropic on its Credentials tab; the shared host is always OpenAI. See the surface for what each one answers.
The shared host, https://api.stackyak.ai, is unpinned and resolves by the algorithm above.
Your key does not change. The same key works on both hosts — the hostname says which stack, the key says whose, and both have to agree. A stack that is not yours is not reachable with your key, and your key does not reach a stack it does not own.
A pin never falls back
If the stack you addressed cannot serve the request, it fails. It is not retried against the shared pool.
This is the point of pinning, and it is worth being deliberate about: send everything to your stack's hostname and you are never quietly served by hardware you did not choose — but you will see failures where the shared host would have answered. Address the shared host when you want the fallback.
When a stack refuses
A stack that cannot answer returns TargetNotServingError with status 404. See
Errors for the envelope.
Two situations produce that refusal, and they are identical on the wire: the stack is yours and is not serving the model you named, or the stack is not yours at all. The message is byte-identical. We do not distinguish them on purpose — a distinguishable refusal would let anyone map which stack identifiers exist under other customers by watching which answer they got. That would be worth someone's time, so it is not available.
When you hit it, check both halves: the hostname is the stack you meant, and that stack is
serving the model you named. GET /v1/models against the stack's hostname lists exactly what it
can answer.
Anonymous traffic never reaches dedicated capacity
Unauthenticated and anonymous requests resolve against the shared pool only. Dedicated capacity belongs to an organisation, and a caller who has not proved one has no dedicated capacity to reach.
A stack's hostname therefore requires a key on every route, including GET /v1/models, which is
open on the shared host. Without one it answers Unauthorized — and it answers that whether or
not the stack in the hostname exists, so the gate cannot be used to test identifiers either.
Split capacity, and what it costs you
An organisation that owns dedicated capacity for some models and uses the pool for others has its traffic split between them, request by request, with no signal in the response. That surprises people, in two ways.
Latency differs. Your node and a pooled node are different hardware under different load. Two identical requests for two different models can differ substantially, and neither is wrong.
They are paid for differently. A dedicated node is a monthly commitment: you are paying for it whether or not you send it traffic. Pool usage draws on your plan's allowance. So moving traffic between them changes your bill even when the request count does not.
Knowing where a request ran
Address the stack. A request to a stack's hostname is served by that stack or refused, so there is nothing left to determine afterwards — which is why the answer to "where did this run?" is a decision you make before sending, not a field you read after.
On the shared host that certainty is not available, and it is not meant to be: the pool exists so you do not have to care. If you need to care, that is the signal to pin.