About the AI API Cost Calculator
This AI API cost calculator estimates what a large language model feature will cost to run. Enter the average number of input (prompt) tokens and output (completion) tokens per request, how many requests you expect per day, and the per-million-token prices from your provider’s pricing page. It returns the cost per request, per 1,000 requests, per day, per month, per year and per active user.
It is built for developers sizing a new AI feature, product managers pricing an AI plan, and founders checking gross margin on an AI-heavy product. Prices are entered by you rather than hard-coded, because model prices change frequently and differ between providers and tiers — copy the current input, output and cached-input prices for the model you plan to use.
You can also model prompt caching (repeated system prompts or documents billed at a lower cached-input price) and batch or committed-use discounts. Output tokens usually cost several times more than input tokens, so shortening responses often saves more than trimming prompts.
With the default inputs, the monthly api cost is $3,600.00. Change any value above to recalculate instantly.
How to use the ai api cost calculator
- 1Estimate average input and output tokens per request from logs or a tokenizer.
- 2Enter expected requests per day and active days per month.
- 3Copy the current input, output and cached-input prices from your provider’s pricing page.
- 4Add a cache hit share and any batch discount if you use them.
- 5Enter monthly active users to see the cost per user and compare it with your price.
Formula and method
LLM APIs bill input (prompt) and output (completion) tokens separately, quoted as a price per million tokens. The cost of one request is the number of each kind of token multiplied by its price, divided by one million. When part of the prompt is read from a prompt cache, that share is billed at the lower cached-input price instead.
Any batch or negotiated discount is applied to the whole request. Daily cost multiplies by requests per day, monthly cost by active days per month, and annual cost is 12 months. Cost per user divides the monthly bill by monthly active users. The estimate excludes cache-write surcharges, fine-tuning, embeddings, tool-call overhead and taxes.
- P_in, P_out
- Price per million input and output tokens
- P_cache
- Price per million cached input tokens
- c
- Share of input tokens served from cache
Worked examples
10,000 requests a day at $3 / $15 per million tokens
Each request uses 1,500 input tokens (1,500 × $3 ÷ 1M = $0.0045) and 500 output tokens (500 × $15 ÷ 1M = $0.0075), so it costs $0.012. At 10,000 requests a day that is $120 daily and $3,600 over a 30-day month, or $3.60 per active user.
Same workload with 60% of input cached
Caching 900 of the 1,500 input tokens at $0.30 per million instead of $3 cuts the input cost per request from $0.0045 to $0.00207. The request now costs $0.00957 and the monthly bill falls to $2,871.
Small model with a 50% batch discount
At $0.15 and $0.60 per million tokens a request costs $0.00024, halved to $0.00012 by batch processing. Even 50,000 requests a day — 1.5 billion tokens a month — costs only about $180.
Frequently asked questions
How are LLM API costs calculated?+
Providers charge separately for input tokens (your prompt and context) and output tokens (the model’s reply), each at a price per million tokens. Multiply each token count by its price, divide by one million, and add the two for the cost of a request.
How many tokens is a word?+
For English text a token is roughly four characters, or about three-quarters of a word, so 1,000 tokens is around 750 words. Code, non-English languages and unusual formatting often use more tokens per word.
Why are output tokens more expensive than input tokens?+
Generating text requires the model to produce tokens one at a time, which uses more compute than reading a prompt. Output prices are commonly several times higher than input prices, so long responses dominate many bills.
How does prompt caching reduce cost?+
When the same prefix — a long system prompt, instructions or reference document — is sent repeatedly, providers that support caching bill those repeated tokens at a discounted cached-input rate. Some charge a small premium to write the cache, which this estimate does not include.
How can I lower my AI API bill?+
Use the smallest model that meets quality needs, cap and shorten outputs, trim prompts and retrieved context, cache repeated prefixes, use batch processing for non-urgent jobs, and cache final answers for repeated questions.