BTC$76,709-0.68%ETH$2,477-1.82%SOL$99.81-1.84%XRP$1.34-1.74%XAU$4,342-0.68%XAG$64.27-0.98%S&P 500$7,657+0.86%Nasdaq 100$29,368+0.91%DAX$25,569+0.82%NVDA$218-0.03%AAPL$333+1.75%MSFT$495+0.65%TSLA$365+0.52%TSM$433+1.22%ASML$1,701+0.64%COIN$175+1.73%MOOD61GreedBTC$76,709-0.68%ETH$2,477-1.82%SOL$99.81-1.84%XRP$1.34-1.74%XAU$4,342-0.68%XAG$64.27-0.98%S&P 500$7,657+0.86%Nasdaq 100$29,368+0.91%DAX$25,569+0.82%NVDA$218-0.03%AAPL$333+1.75%MSFT$495+0.65%TSLA$365+0.52%TSM$433+1.22%ASML$1,701+0.64%COIN$175+1.73%MOOD61Greed
All prices
inotok

Tokens per month to money

AI Inference Cost

What an AI feature costs per month at scale: requests, tokens in and out, model tier, cache hit rate, and the comparison with renting a GPU for a self-hosted model.

What this tool does

Language models are priced per million tokens, with input and output at different rates and cached input usually cheaper. The calculator takes requests per month, average input and output length, a model tier with editable prices, and a cache share, and returns the monthly bill. Beside it, a self-hosting estimate from GPU hourly cost, utilisation and throughput shows where the break-even lies.

How to use it

  • Enter requests per month and the average input and output tokens per request.
  • Pick a model tier or type your own prices per million tokens, and set the share of input served from cache.
  • Compare the monthly API cost with the self-hosting estimate and read the break-even in requests.

Good to know

  • Prices in the presets are typical ranges for 2026, not the price list of any provider; check the current list before budgeting.
  • Self-hosting cost depends heavily on utilisation. A GPU that runs at twenty percent load costs the same as one at eighty.
  • Latency, rate limits and evaluation effort are costs too and are not in the number.

Read next

Reads that use this tool

Tools

Other tools