Cost per Token

What Is Cost per Token?

Cost per token is the fully loaded cost of producing one unit of model output, usually quoted per million tokens. It rolls up hardware amortization, power, networking, and the software stack into a single number, which makes it the base unit of inference economics.

For an operator selling inference, it is the denominator in every margin calculation. The price the market will pay per million tokens is largely set externally. Cost per token is the side of that equation an operator can actually move.

What Sets It

The figure is set by more than hardware. The same GPU produces very different cost per token depending on:

  • Utilization. Idle accelerators still draw power and still amortize. Low utilization inflates cost per token directly.
  • Batch strategy. Continuous batching keeps the GPU fed under concurrent traffic instead of stalling behind long requests.
  • Quantization. Lower precision raises throughput per GPU, within the limits of what the silicon supports.
  • Serving engine version. Engines ship kernel and scheduler improvements continuously, and the gains land without a hardware change.

Software gains alone can move cost per token several fold on identical silicon. That is why it is a moving target rather than a fixed property of a chip, and why a configuration that was efficient last quarter may not be today.

Why It Matters

For operators, cost per token is the number that decides whether a fleet earns a margin at the price the market will pay. It sits underneath the broader question of token economics and connects directly to tokens per watt, since power is the largest recurring input to the cost.

Try Saturn Cloud today

Start for free. On a team? Contact Us!