Tokens In AI

What Are Tokens in AI and Why They Cost You Money

October 3, 2026

A token is the unit AI models read and write in roughly four characters of English, or about three-quarters of a word and it’s the unit you’re billed in, because providers charge for every token going in and every token coming out. A thousand tokens is around 750 words. Every message you send, every reply you get, and every document you upload is converted into tokens before anything else happens.

That’s the definition. The part that actually costs people money is less obvious: tokens don’t only bill you on an API invoice. They quietly govern what you get from a flat-rate subscription too, and the way conversations accumulate them means a long chat costs far more than a short one not proportionally more, but dramatically more.

What a token actually is

Models don’t read words. During training, a tokenizer learns to break text into frequently occurring chunks, and that vocabulary becomes the model’s alphabet.

Common words are usually a single token. “The”, “and”, “computer” one each. Less common words get split: “tokenization” might be three pieces. Punctuation and spaces count. Numbers often split digit by digit, which is part of why models fumble arithmetic.

Tokens in AI

Two consequences follow immediately. Every provider trains its own tokenizer, so identical text can be 100 tokens on one model and 130 on another the rates aren’t directly comparable between providers. And the four-characters rule only holds for English prose. More on that below, because it matters more than most people expect.

The four ways tokens cost you

1. Directly, on an API bill. If you’re building anything, you pay per token, quoted per million. Input and output are billed at different rates.

2. Indirectly, through subscription limits. On a $20-a-month plan you’re not shown a token counter, but it’s there. Message caps, the point where a conversation stops accepting new messages, and how much of a document the assistant will read are all token budgets in disguise. When a chat tells you it’s too long, you’ve hit a token ceiling.

3. Through wasted spend. Rambling output, over-long system instructions and re-sent context that nobody needed.

4. Through time. More tokens means slower responses. On a long conversation the delay before the first word appears grows noticeably.

Input and output are not priced the same

This is the single most useful fact about AI pricing, and it’s counterintuitive.

Output tokens typically cost several times more than input tokens on the same model commonly three to ten times, depending on the provider. Reading is cheap; generating is expensive, because each output token requires a full pass through the model while input can be processed in parallel.

The practical implication reverses most people’s instincts. A long, detailed prompt is cheap. A model that answers at length is expensive. If you’re trying to cut costs, “be concise” in your instructions saves more than trimming your prompt.

The arithmetic

The formula is simple:

cost = (input tokens ÷ 1,000,000 × input rate) + (output tokens ÷ 1,000,000 × output rate)

A worked example, using illustrative rates of $3 per million input and $15 per million output:

  • You send a 2,000-token prompt: 0.002 × $3 = $0.006
  • The model writes a 800-token reply: 0.0008 × $15 = $0.012
  • Total: $0.018, under two cents.

That looks trivial, and for one call it is. Run it 10,000 times a month and it’s $180. Run it inside a product where each user action triggers three calls and you’re at $540. Token cost is a volume problem wearing a small price tag.

I’d add one caution about estimating: don’t count tokens from your own word count. Use the provider’s tokenizer or read the token counts returned in the API response, because your estimate will be wrong in a direction you can’t predict.

Why long conversations get expensive fast

Here’s the mechanism almost nobody explains, and it’s the biggest hidden cost in everyday use.

Language models have no memory. Every time you send a message, the entire conversation is sent again from the beginning. Turn one sends your first message. Turn ten sends all nine previous exchanges plus the new one.

So the tokens processed across a conversation don’t grow in a straight line they grow with the square of its length. A twenty-message conversation doesn’t cost twice a ten-message one. It costs roughly four times as much.

For subscribers this shows up as hitting limits faster than you expected in a long thread. For anyone building, it’s why chat products get expensive and why prompt caching where providers charge a fraction of the normal rate for a repeated prefix exists at all.

The fix in both cases is the same: start a fresh conversation when you change topic. Carrying twenty turns of unrelated history into a new question buys you nothing and costs you on every subsequent message.

The things that quietly multiply your token count

Non-English text. Tokenizers are trained predominantly on English, so English gets the efficient encoding. Urdu, Arabic, Hindi, Thai, Japanese and Chinese frequently consume two to three times as many tokens for the same meaning, occasionally more. Same question, same model, multiple times the cost a real and under-discussed inequality in how these tools are priced.

Images. Uploading a screenshot doesn’t cost one token. Vision models convert an image into a token equivalent based on its resolution, often running to hundreds or over a thousand. Crop screenshots before uploading; it genuinely matters.

Documents. A 20-page PDF can be 15,000 tokens or more. Uploading it puts all of that into every subsequent turn of that conversation.

Code and structured data. Indentation, brackets and JSON punctuation all tokenize. Minified or trimmed code costs noticeably less than the same logic pretty-printed.

Reasoning tokens. Models that “think” before answering generate internal tokens you never see, and on most providers you’re billed for them at the output rate. A short visible answer can carry a large invisible cost.

System prompts and tool definitions. In any application, the hidden instructions and the descriptions of available tools are re-sent with every single call.

How to spend less

If you’re paying a subscription:

  • Start new conversations often rather than running one endless thread.
  • Upload only the pages of a document you need.
  • Crop images.
  • Ask for the length you want. “In three sentences” costs a fraction of an unbounded answer.

If you’re paying per token:

  • Choose a smaller model for routine tasks. Price spreads across model tiers are enormous, and most workloads don’t need the flagship.
  • Set a maximum output length on every call. It’s the single most effective cap.
  • Use prompt caching for any prefix that repeats.
  • Batch where the provider offers a discounted asynchronous tier.
  • Log token counts per request from day one. You cannot reduce what you aren’t measuring, and the response object already tells you.