Shared or dedicated
There are two ways to run a model here: on our shared compute, or on a node that is yours. This is the first decision to make, because it sets what you pay, which models you can run, and what happens when your traffic stops.
| Shared | Dedicated | |
|---|---|---|
| What you provision | Nothing | A node, before it can serve |
| What you pay for | Your subscription | The node itself, per GPU, monthly |
| Which models | The shared set, by tier | Anything we carry, plus your own |
| Isolation | Pooled with other customers | Yours alone |
| If you stop sending traffic | Change or cancel your plan | You pay for the rest of the term |
The two are not exclusive. Plenty of setups run most traffic on shared compute and keep a node for the one model that needs it.
Shared compute
You send requests to our API, and they are served on hardware we run and share between customers. There is nothing to provision: sign up, get an API key, and start calling. Your subscription sets how much you can do, and the models you can reach are the ones we keep loaded.
This is the fastest way to start building, and a good place to develop against. When you move to production, we recommend going dedicated: the capacity is yours, and how it performs stops depending on who else is busy.
Dedicated compute
A GPU node that belongs to you for the term you commit to. You choose what runs on it, including models that are not on shared compute and models we do not carry in the catalog at all. No other customer's traffic touches it.
You pay monthly per GPU for the length of the commitment. A node is provisioned before it serves, and it only takes a few moments before it is answering your requests.
This is the right choice when isolation matters to you, when your traffic is steady enough to keep a node busy, when you need a model we do not pool, and when you are ready to run in production.
What shared compute is running can change
GET /v1/models returns what is loaded right now, not everything we could load. Treat it as the current answer rather than a fixed list, and check it rather than assuming a model is there. The API covers the endpoint.
Where to go next
Shared: the API sets out the endpoints, and authentication gets you a key.
Dedicated: compute pricing has the rates, and support can help you size a node before you commit to one.
Not sure which you want? Ask the Yak will talk through your workload and say which one fits.