Inference Attach

What Is Inference Attach?

Inference attach, also called software attach, is the degree to which an operator layers managed inference and software services on top of raw GPU capacity, rather than renting bare metal alone. It is the difference between selling access to hardware and selling tokens.

A bare metal rental transfers the GPU and the problem. The customer brings their own serving stack, their own scaling, and their own reliability engineering, and the operator is paid for hours of access. An attached inference product sells the output instead: an endpoint with a rate card, a latency target, and a bill measured in tokens.

How the Market Prices It

The market increasingly prices this directly. Operators that attach inference show it in:

  • Higher gross margins, because the revenue is tied to output rather than to a commodity hourly rate.
  • Larger remaining performance obligations, because token contracts are longer and stickier than hourly capacity.
  • Stronger returns on power, which is the constraint that actually binds. See revenue per megawatt.

The emerging investor test is that valuation dispersion should be set by software attachment and fully burdened returns per megawatt, not by how many megawatts an operator controls. Capacity is increasingly assumed. What separates operators is what runs on top of it.

What Attaching Requires

Attaching inference means running the layers a bare metal renter does not: an inference serving stack kept current across engine releases, and a tenant plane that meters, isolates, and bills. Those are the two pieces that convert GPU capacity into a product with a per-token price.

Try Saturn Cloud today

Start for free. On a team? Contact Us!