Yield Theory

Calculator

AI Inference Cost per Million Tokens Calculator

Convert accelerator cost, utilization, throughput, and overhead into token economics.
Cost per million tokens
$1.74
Price for target margin
$4.34
Effective tokens per hour
13,824,000

Results are workload-specific. Prompt length, output length, batching, latency targets, model architecture, precision, cache reuse, and memory bandwidth can dominate realized throughput.

Peak benchmark throughput is not the same as billable production output. This calculator applies utilization and overhead to tokens per second, compares effective output with hourly accelerator cost, and estimates cost and required price per million tokens at a target gross margin.

Formula reviewed 2026-08-021 primary sourceMethodology and disclosures →

How this ai inference cost per million tokens calculator works

The model converts benchmark or measured throughput into hourly tokens, reduces it for utilization and non-billable overhead, and divides accelerator-hour cost by effective output. Target price is cost divided by one minus margin.

Formula

Cost per million tokens = hourly accelerator cost ÷ effective tokens per hour × 1,000,000

This simplified model excludes networking, CPU, storage, engineering, failed requests, reserved-capacity terms, customer acquisition, and model-development costs unless embedded in the hourly input.

Before you use the result

Assumptions

  • Throughput is measured for the relevant model and service-level target.
  • Utilization and overhead do not double-count the same idle capacity.
  • The entered hourly cost captures the intended hardware or cloud cost scope.

Quick start

  1. 1. Use measured production throughput where possible.
  2. 2. Separate powered-on utilization from billable utilization.
  3. 3. Stress latency, concurrency, model size, precision, energy, and pricing assumptions.

Frequently asked questions

Why is benchmark throughput insufficient?

Production workloads face latency targets, varying sequence lengths, batching limits, downtime, and non-billable activity.

What is utilization in this model?

It is the share of theoretical time represented by productive workload before applying the separate overhead assumption.

Does price for margin include all company costs?

Only costs included in the accelerator-hour input. Sales, research, networking, facilities, and support may require additional margin.

Token cost is one layer of the AI return stack.

Yield Theory connects inference economics with accelerators, HBM, networking, power, capex, pricing, and public-company exposure.

$15/mo or $150/yr · cancel future renewals anytime · sources included