Inference Economics
A live view into the real costs of getting AI models to perform real tasks over time, adjusted for intelligence levels and live market prices.
Reach out to @ojasvi_goel if you want to chat inference economics :)
Data from Artificial Analysis, OpenRouter, and Ramp. OpenRouter update time loading…
Model-family activity and economics
Share of OpenRouter tokens
Seven-day average; named families plus all other models.
Average cost per equivalent task for model providers
Seven-day average, weighted by estimated task-equivalent volume.
Estimated task-equivalent volume by model family
Seven-day average; observed tokens converted into Artificial Analysis standardized-task equivalents.
Estimated listed revenue by model family
Seven-day average of daily OpenRouter token volume repriced at each model’s cheapest listed endpoint.
Task equivalents are estimated, not observed user tasks. Revenue is an indicative listed-floor estimate, not observed billing, OpenRouter net revenue, or provider receipts. Both assume observed tokens follow each model’s Artificial Analysis standardized-task mix; hover a family to see covered-token coverage.
Rentable-GPU inference spread
Test a transparent GPU-hour scenario against inference token revenue.
Revenue capacity vs. utilization
Maximum token revenue per GPU-hour; observed rental asks are horizontal hurdles.
Gross spread vs. H100 rental price
How quickly the economics compress at different achieved utilization levels.
Historical inference revenue capacity
Measured throughput repriced at the historical OpenRouter Llama 3.3 70B listed floor.
Gross revenue capacity—not profit. Assumes all generated tokens sell at the listed floor and excludes CPU, RAM, storage, bandwidth, orchestration, failures, fees, support, and reserve capacity.
Cost per task over time
Custom
Current model comparison
| Model and sources | AA score | Listed cost per task | Effective cost per task |
|---|
How to read this chart
Each line shows what a standardized Artificial Analysis task would have cost using either the cheapest or average model around a given threshold. Data is based on listed market offers, with cost per task estimated using Artificial Analysis benchmarks of token usage per task and OpenRouter listed costs per token.
Intelligence level vs. cost for a standardized task
Compares the average Artificial Analysis Intelligence Index score of models used by OpenRouter users vs. the cost per task.
Capability and cost use separate stacked panels with one shared date axis.
Inspect the underlying token price
Use this to understand one model's listed or realized endpoint floor.
Input and output price
Advertised endpoint price floor.