Saturn Cloud is now available as a vCluster Stack, so operators can launch a managed inference service for every customer in one deployment.
Most neoclouds want to sell more than raw GPU hours. Per-token inference earns more from the same fleet, but offering it to enterprise customers runs into a hard requirement. Every customer needs their own isolated environment, and standing that up by hand for each one is a custom integration you don’t want to repeat.
We partnered with vCluster Labs to remove that. Saturn Cloud is now available as a vCluster Stack, a single, reusable, pre-validated template that installs the full Saturn Cloud platform inside an isolated tenant cluster. An operator defines the environment once, then deploys it per customer. Each customer gets their own isolated cluster and per-token endpoints on the operator’s own GPUs.
This follows our integrations with NVIDIA DSX OS and NVIDIA Run:ai, which give operators on NVIDIA infrastructure the same managed platform without building it themselves.
How it works
Saturn Cloud sells the tokens. It turns GPUs into per-token inference customers can call, with per-token metering and billing, fine-tuning, and a self-service portal under the operator’s own brand. Inference is the first Stack, and the same platform also runs dedicated GPU capacity, managed fine-tuning, and distributed training on the same fleet.
vCluster isolates each customer. Every customer gets their own tenant cluster with a virtualized control plane. With Private Nodes, each tenant also runs on dedicated nodes, with network, storage, and compute isolated per tenant.
Stacks makes it repeatable. The Saturn Cloud Stack describes what to install, in what order, and what has to be healthy before the next step. Each deployment is its own managed environment, with its own health status, logs, and clean removal.
Onboarding a customer
Say an AI cloud wants to offer enterprise customers private inference. The operator creates the customer’s tenant cluster on Private Nodes, running on GPU servers they already own. They deploy the Saturn Cloud Stack into it, and the customer can start calling per-token endpoints. Onboarding the next customer is the same automated process again, consistent every time.
The point is that a new customer is a deployment, not a project. Inference sells alongside GPU capacity and Kubernetes, Slurm, and Ray clusters on the same fleet, and every new customer is one more run of a template the operator already trusts.
Available now
The first Saturn Cloud Stack, built by vCluster engineers from Saturn Cloud’s Kubernetes installation, is available today, and operators can adapt it to their own environments.
You can read more about Stacks in vCluster Platform 4.12, or talk to an engineer about standing up a managed inference offering on your own fleet.



