← Back to Blog

Turn a GPU Fleet Into an Inference Business with NVIDIA Run:ai

Saturn Cloud now integrates with NVIDIA Run:ai. Run:ai handles GPU orchestration and scheduling across the fleet. Saturn Cloud adds the commercial layer (multi-tenancy, metering, and billing) that lets operators sell that capacity as per-token inference under their own brand.

Turn a GPU Fleet Into an Inference Business with NVIDIA Run:ai

A GPU fleet rented by the hour earns a flat return on hardware you have already financed and powered. Sold as inference, by the token, across many customers, the same fleet can earn considerably more. The catch is that capturing that revenue means running a multi-tenant inference business, and that is a real build most operators would rather not take on.

We integrated Saturn Cloud with NVIDIA Run:ai so they don’t have to. Operators get the orchestration layer from NVIDIA and the commercial layer from Saturn Cloud, and the combination turns raw capacity into per-token products they run under their own brand.

The gap between owning GPUs and selling inference

For a neocloud, how capacity is sold decides how much it earns. Renting GPUs by the hour puts a ceiling on revenue. Selling the same GPUs by the token, across many tenants, moves the cost-per-token economics in the operator’s favor and puts idle capacity to work.

Getting there is the hard part. It means isolating every customer on shared hardware, metering what each one uses, billing them for it, and keeping their data inside the operator’s security boundary. That is the tenant plane, and no open-source serving component ships it. It is the layer a token factory adds on top of GPU infrastructure, and it is what Saturn Cloud provides.

What the integration does

NVIDIA Run:ai maximizes GPU utilization across the fleet through NVIDIA KAI Scheduler, with gang scheduling for NVIDIA Grove-managed distributed workloads and policy-driven governance. Saturn Cloud runs on top as the commercial layer between the fleet and its customers, so operators can sell inference without building that layer themselves.

As Sebastian Metti, Founder of Saturn Cloud, put it, “A neocloud can raise revenue per megawatt without adding a single GPU, just by changing how the capacity is sold. Saturn Cloud gives operators the full commercial platform to do it, handling serving, multi-tenancy, and billing, all running under their own brand on NVIDIA Run:ai.”

Several products from the same fleet

Beyond per-token serving, operators get more than one product to sell from the same infrastructure, without building any of it in-house.

They can offer dedicated GPU capacity to customers who bring their own stack. They can offer per-token model-as-a-service to customers who just want an API endpoint. And they can offer self-service development environments for teams that build and fine-tune their own models. Saturn Cloud also handles model onboarding, shared and dedicated isolation tiers for regulated customers, and the identity, access, and governance controls enterprise buyers require.

How the serving stack works

Underneath, the serving layer runs on NVIDIA Dynamo, which handles distributed inference with disaggregated prefill and decode and drives vLLM, SGLang, or NVIDIA TensorRT-LLM. NVIDIA Grove orchestrates the multi-node workloads, and NVIDIA KAI Scheduler places them with GPU awareness, allocates fractional GPUs, and enforces quotas across tenants. NVSentinel and NVIDIA Fleet Intelligence handle fleet-wide health monitoring and fault remediation, so degraded hardware is pulled from placement automatically.

The result is a production inference path an operator can stand up on NVIDIA infrastructure without assembling the stack themselves, with standard inference endpoints their tenants can call.

Part of a wider NVIDIA stack

NVIDIA Run:ai is one of several NVIDIA technologies Saturn Cloud builds on. It follows our integration of the NVIDIA DSX AI Factory Platform and extends our work across the NVIDIA inference stack. For telcos, sovereign AI clouds, and enterprise operators, the shape is the same. NVIDIA provides the orchestration and serving, and Saturn Cloud does the production engineering that turns it into a service an operator can run and sell.

“Cloud providers are looking for new ways to turn GPU infrastructure into differentiated AI services,” said Omri Geller, NVIDIA VP, DSX OS Platform Software. “Together, NVIDIA Run:ai and Saturn Cloud help operators improve GPU utilization while delivering scalable, multi-tenant inference services under their own brand.”

Available now

The NVIDIA Run:ai-integrated Saturn Cloud platform is available now. If you operate a GPU cloud and want to sell inference on the fleet you already run, talk to an engineer.

Keep reading

Related articles

Turn a GPU Fleet Into an Inference Business with NVIDIA Run:ai
Sep 15, 2026

The Cost of Keeping a Model Catalog Current

Turn a GPU Fleet Into an Inference Business with NVIDIA Run:ai
Aug 12, 2026

Integrating NVIDIA DSX OS Into the Saturn Cloud Token Factory

Turn a GPU Fleet Into an Inference Business with NVIDIA Run:ai
Aug 11, 2026

Running Production AI on Your Own GPUs, with Rafay and Saturn Cloud