Calculator
AI Inference Cost per Million Tokens Calculator
Results are workload-specific. Prompt length, output length, batching, latency targets, model architecture, precision, cache reuse, and memory bandwidth can dominate realized throughput.
Peak benchmark throughput is not the same as billable production output. This calculator applies utilization and overhead to tokens per second, compares effective output with hourly accelerator cost, and estimates cost and required price per million tokens at a target gross margin.
How this ai inference cost per million tokens calculator works
The model converts benchmark or measured throughput into hourly tokens, reduces it for utilization and non-billable overhead, and divides accelerator-hour cost by effective output. Target price is cost divided by one minus margin.
Formula
Cost per million tokens = hourly accelerator cost ÷ effective tokens per hour × 1,000,000
This simplified model excludes networking, CPU, storage, engineering, failed requests, reserved-capacity terms, customer acquisition, and model-development costs unless embedded in the hourly input.
Primary specifications
Before you use the result
Assumptions
- • Throughput is measured for the relevant model and service-level target.
- • Utilization and overhead do not double-count the same idle capacity.
- • The entered hourly cost captures the intended hardware or cloud cost scope.
Quick start
- 1. Use measured production throughput where possible.
- 2. Separate powered-on utilization from billable utilization.
- 3. Stress latency, concurrency, model size, precision, energy, and pricing assumptions.
Frequently asked questions
Why is benchmark throughput insufficient?
Production workloads face latency targets, varying sequence lengths, batching limits, downtime, and non-billable activity.
What is utilization in this model?
It is the share of theoretical time represented by productive workload before applying the separate overhead assumption.
Does price for margin include all company costs?
Only costs included in the accelerator-hour input. Sales, research, networking, facilities, and support may require additional margin.
More calculators
Token cost is one layer of the AI return stack.
Yield Theory connects inference economics with accelerators, HBM, networking, power, capex, pricing, and public-company exposure.
$15/mo or $150/yr · cancel future renewals anytime · sources included