The model catalog
The catalog is the set of models you can call on our shared compute, or install on hardware you own. Browse it at stackyak.ai/models.
Why it is curated
We carry a chosen set rather than mirroring everything published on Hugging Face. The set is picked to cover most of what people build with: general chat, vision, coding, a fast small model, and a safety classifier, across a range of sizes.
That choice has a cost and a benefit, and both are worth knowing. On shared compute you are limited to the set we carry. In exchange, everything in it is already loaded, so a request does not wait for a model to start, and your bill does not depend on which one you picked.
The limit belongs to the shared pool rather than to the platform. On dedicated compute you can run a model we do not carry at all, which is covered further down.
If a model you need is missing, tell us. Open a support ticket naming the model and what you want it for. Requests are how the set grows, and the ones that come with a use case behind them carry the most weight.
Which models your plan can call
Every model on shared compute carries a minimum tier: the lowest plan that may call it. If your plan is at or above it, the model is yours to use.
A model with no minimum tier is not on shared compute, at any price. Those are install-only: you can still run them, on a node you own. Calling one on the shared endpoint is refused rather than served, and the refusal says so.
Moving up a tier does not necessarily add models. Neighbouring tiers can carry exactly the same set and differ only in throughput, so check the card rather than assuming. Plans and tiers covers what a tier does change.
Covered is not the same as running
Two separate things have to be true before a model answers you: your plan covers it, and it is currently loaded.
Models on the shared set are kept loaded, so one your plan covers normally answers without waiting for a start-up. A model that is listed but not loaded will not answer, and the product shows you which state it is in. If you get a refusal you did not expect, Errors tells you which of the two you hit.
Reading a model card
Each card carries:
- Model ID — the value you put in the
modelfield of a request. - Family and vendor — who built it. A family groups a vendor's related models.
- Context — how much a single request may carry. Your tier caps this as well, so the lower of the two applies.
- Parameters — model size.
- Capabilities — what it can do: chat, tool calling, vision, long context, speed.
- Licence — the vendor's terms, which govern what you may do with the output.
- Links — the vendor's own model page and announcement, so you can check our summary against the source.
Input, Output, and Speed are our estimates, not vendor figures. We set them across the catalog by model size, to help you plan capacity on your own hardware. They are blank for models on the shared set, because shared inference is covered by your subscription rather than billed per token.
Any benchmark scores shown are the vendor's published numbers. They are not our measurements, and we do not present them as such.
Preview means unproven, not unavailable
Most cards are marked preview. That means the specification was read off the vendor's published card, and we have not yet run the model on our own hardware to confirm it.
A preview model is available and callable. What is not promised is that the figures on the card match what you will see. A model is only marked available once it has started on the runtime version we pin and answered a real request.
If you are planning production capacity, this is the distinction to pay attention to.
Running a model on your own hardware
Some models need one extra step before a node can serve them, and the card says which:
- A gated repo — the vendor requires their terms to be accepted once before the weights can be downloaded.
- Remote code — the model ships its own code that the runtime has to be allowed to execute.
On shared compute this is already handled. On a node of your own, it is a step you will do as part of installing the model.
Running a model we do not carry
The catalog is not the limit of what you can run. A model we do not list can still be served on dedicated compute, including your own fine-tune or weights you hold privately. The shared pool is the fixed set described above; your own hardware is not.
There is no self-serve path for this yet. Open a support ticket with the model you want to run and the hardware you have or are planning to buy, and we will take it from there.