What StackYak is
Running a model in production normally means assembling three things from three vendors: somewhere to serve the model, GPUs to serve it on, and a network path to reach it from your own systems. Each is bought separately, operated separately, and billed separately.
We sell the three as one unit. You order it once, it comes up configured, and it arrives on one bill.
A stack
That unit is a stack: a model, the compute it runs on, and the network path that reaches it. A stack has its own endpoint, and it is the thing you order, modify, and shut down.
Your first decision
Everything else follows from one choice: run on our shared compute, or on hardware that is yours.
Shared compute needs nothing provisioned and is covered by a subscription. Dedicated compute is a node of your own, which you can run any model on, including models we do not carry. Shared or dedicated sets out the trade.
What we run
NVIDIA and AMD accelerators, in a number of locations across the United States. We add locations and capacity regularly, so what you can order today is not the ceiling.
On the model side we carry a curated catalog rather than mirroring everything published on Hugging Face, so that what we serve is already loaded and behaves predictably. The model catalog explains how it is organised, and stackyak.ai/models is the current list.
Who this is for
You are running open-weight models and would rather not assemble and operate the stack yourself. You want isolation, predictable performance, and one vendor to call. You want the bill to make sense.
Who this is not for
You want a general-purpose cloud. We do model serving, the hardware under it, and the network to it — not databases, object storage, or the rest of a hyperscaler's catalogue.
You want every model on Hugging Face on shared compute. We are unapologetic about carrying a small, curated set: these are the models we think do the job best, and keeping the set tight is what lets us keep them loaded and performing. You are the judge of your own use case, though, and dedicated compute is yours to do as you like with — install whatever you want on it.
You want capacity brokered across every neocloud. We run our own fleet with suppliers we have real relationships with, rather than reselling whoever is cheapest this week. That is what gives you stability, scalability, and a platform that stays consistent as you grow.
Where to go next
Plans and tiers covers what a subscription buys. Pricing covers the three ways money moves.