LLM Inference Cost & Serving Calculators

What it costs to actually run a language model in production: per-token API pricing split across input and output, cost per user per month, prompt-cache savings, throughput in tokens per second, and the point at which self-hosting beats paying per token.

5 calculators in this category

More ai, llm & machine learning engineering categories

Related categories in other subjects

Where this work overlaps other trades and disciplines.