# What Is an AI Token? Count Tokens, Context, and Cost

> What is an AI token? AI tokens explained: how models split text, why code and language change counts, how to estimate a bill and count Claude tokens.

What is an AI token? It's a piece of text—sometimes a whole word, sometimes part of one, or even punctuation—that a model processes as a unit. A space can belong to a token too. The sentence “AI reads text” isn't necessarily three tokens.

**The short answer:** An AI token is a model's working unit for text, not a word, character, or universal measurement shared by every provider. Your prompt and the model's reply are divided into tokens; providers use those counts for context limits and often for billing. “1,000 tokens ≈ 750 English words” is a useful first estimate, but code, JSON, language, model, caching, and reasoning can change the count. Use the counter for the model you will call when the number matters.

## What is an AI token?

An LLM token is a sequence of characters from a model's vocabulary. A common word can be one token; an uncommon or long word may split into several. OpenAI uses “tokenization” as an example that can split into “token” and “ization” ([OpenAI's token guide](https://developers.openai.com/api/docs/concepts)). Different tokenizers split the same text differently.

The tokenizer is the practical answer to “what is a token in AI?” It converts text into the pieces the model sees. Those pieces don't carry meaning by themselves, and there's no single token count for a paragraph independent of the model.

## How many words are 1,000 tokens?

For ordinary English prose, start with about 750 words per 1,000 tokens, then treat the result as a budget estimate. OpenAI's own rule of thumb is one token per four characters or 0.75 words of English; that gets you close for plain writing, not an exact count ([OpenAI token guide](https://developers.openai.com/api/docs/concepts)).

I measured a 3,338-word Markdown article with links at 1.44 tokens per word on `o200k_base`, or about 696 words per 1,000 tokens. In the English Universal Declaration of Human Rights (UDHR), the same tokenizer measured 1.26 tokens per word, or about 791 words per 1,000 tokens. Formatting and vocabulary moved the estimate by roughly 14% across these two samples.

| Text sample                      | Tokens per word | Characters per token | What to budget for           |
| -------------------------------- | --------------: | -------------------: | ---------------------------- |
| English UDHR                     |            1.26 |                 5.58 | Clean prose, few links       |
| English blog Markdown with links |            1.44 |                 4.47 | Headings, links, punctuation |
| TypeScript file                  |            2.14 |                 3.59 | Syntax and identifiers       |
| Rust source file                 |            2.19 |                 3.75 | Syntax and identifiers       |
| JSON catalog excerpt             |            4.41 |                 2.36 | Keys, quotes, braces, values |

These are `o200k_base` counts, not a universal conversion. In the 20,000-character JSON excerpt, the same tokenizer counted 8,470 tokens. Dividing 20,000 by four would predict 5,000, an underestimate of 3,470 tokens. If you're budgeting for source code, tool schemas, or structured data, count a representative sample instead of multiplying characters by 0.25.

## Do language and model change the token count?

Yes: even translations of the same document can produce different counts, and the size of that difference depends on the tokenizer. I NFC-normalized the NLTK UDHR2 text files, then counted each language with the two OpenAI encodings. The Vietnamese corpus text is stored decomposed; without normalization, its count is artificially inflated.

| UDHR language      | `o200k_base` tokens vs English | `cl100k_base` tokens vs English |
| ------------------ | -----------------------------: | ------------------------------: |
| English            |                          1.00× |                           1.00× |
| Spanish            |                          1.22× |                           1.44× |
| Hindi              |                          1.58× |                           5.21× |
| Japanese           |                          1.72× |                           2.30× |
| Simplified Chinese |                          1.18× |                           1.68× |

Those are measurements of particular models' tokenizers, not a statement about the difficulty or quality of a language. A different tokenizer can reverse the comparison: the same Chinese UDHR text measured 0.87× English on Qwen3.5-9B and 0.85× on DeepSeek-V3.1 in the same run. A count from one model isn't a safe substitute for another model's count.

To reproduce the comparison, normalize the same sample to NFC, pass it through each model's own tokenizer without adding special tokens, and divide each language's count by English's count for that tokenizer. I used tiktoken 0.14.0 for `o200k_base` and `cl100k_base`, the tokenizer files from the Qwen3.5-9B, DeepSeek-V3.1 and GLM-4.6 Hugging Face repos, and Unsloth's public re-uploads of Llama 3.1 8B and Gemma 3 4B. The pre-run maps GPT-5, GPT-4.1, and GPT-4o to `o200k_base`, and GPT-4 and GPT-3.5 Turbo to `cl100k_base`; don't extend that map to other models by guesswork.

## How do tokens turn into an AI bill?

Multiply each usage category by that model's price for the category, then divide by one million. Input, cached input, cache writes, output, and reasoning may have different prices or billing labels. Rates below are GPT-6.1 Sol standard prices checked October 4, 2026: $2 per million uncached input tokens, $0.10 per million cached input, $2.50 per million cache writes, and $10 per million output tokens ([GPT-6.1 Sol pricing](https://developers.openai.com/api/docs/models/gpt-6.1-sol)).

Here's an illustrative agent turn, not a bill from an API request. Assume its 170,000 input tokens include 20,000 uncached, 120,000 cache reads, and 30,000 cache writes. The model then generates 3,000 visible answer tokens and 12,000 reasoning tokens:

| Category       | Calculation                 |         Cost |
| -------------- | --------------------------- | -----------: |
| Uncached input | 20,000 × $2 / 1,000,000     |     $0.04000 |
| Cached input   | 120,000 × $0.10 / 1,000,000 |     $0.01200 |
| Cache write    | 30,000 × $2.50 / 1,000,000  |     $0.07500 |
| Visible answer | 3,000 × $10 / 1,000,000     |     $0.03000 |
| Reasoning      | 12,000 × $10 / 1,000,000    |     $0.12000 |
| **Total**      |                             | **$0.27700** |

Reasoning tokens are included in the 15,000 generated output tokens for pricing; they're not a fifth price tier in this example. They also occupy context space. OpenAI's usage example reports `output_tokens: 1186` with `reasoning_tokens: 1024`, and its docs say reasoning is billed as output ([OpenAI reasoning docs](https://developers.openai.com/api/docs/guides/reasoning)). The calculation uses 170,000 input tokens, under GPT-6.1 Sol's 272K input threshold for its standard short-context prices.

A bill can look different even when the prompt text doesn't. A cached prefix may shift usage from ordinary input to the cheaper cached-input category; the initial cache write has its own rate. Read the provider's usage object and pricing table together. For a deeper treatment of repeated-prefix savings, see [prompt caching explained](/blog/prompt-caching-explained); for per-model token prices, see the [model price catalog](/models).

## How do you count Claude tokens exactly?

For Claude API requests, use Anthropic's model-specific `count_tokens` endpoint as your Claude token counter before sending the message, then use the response usage for what the request actually consumed. Anthropic explicitly describes the preflight count as an estimate that can differ slightly from actual usage, so calling it “exact” overstates what the endpoint promises ([Anthropic token-counting docs](https://platform.claude.com/docs/en/build-with-claude/token-counting)).

**Prerequisites:** a Python environment with Anthropic's SDK and API credentials; any OS supported by that SDK. This is the documentation's example with the `import` line added, not run here; the docs don't pin an SDK version. Copy the request shape, then replace the sample system prompt and message with your actual request:

```python
import anthropic

client = anthropic.Anthropic()

response = client.messages.count_tokens(
    model="claude-opus-5-5",
    system="You are a scientist",
    messages=[{"role": "user", "content": "Hello, Claude"}],
)

print(response.json())
```

**Expected output:** `{"input_tokens": 14}` for the documented example. The call doesn't create a message; Anthropic says counting is free, with separate rate limits of 5,000 requests per minute for Start, 10,000 for Build, and 20,000 for Scale.

The error people hit in practice is a 400 when the request carries server tools or uploaded files. An [August 2026 PydanticAI issue](https://github.com/pydantic/pydantic-ai/issues/7867) reports Anthropic's `count_tokens` returning `File sources are not supported in the token counting endpoint` when a code-execution tool carries uploaded files, and the docs say requests with server tools other than the advisor tool return an error. Keep preflight counts to supported client tools and content (Anthropic says to send PDFs or images as base64); when the real request depends on server tools, MCP, or provider-side files, read its returned `usage` instead.

Anthropic's docs also flag a Claude tokenizer change. From Claude 4.7 on (and in Claude Mythos Preview), identical text comes out at "approximately 30 percent more tokens" than on earlier Claude models, and the real increase varies with your content. A count taken on an older model is stale for these: recount with the model ID you'll call ([Anthropic's token-counting docs](https://platform.claude.com/docs/en/build-with-claude/token-counting)).

For the other major APIs: OpenAI gives you `tiktoken` for model-aware local text counts ([OpenAI tiktoken](https://github.com/openai/tiktoken)); Gemini's SDK exposes `client.models.count_tokens(model=..., contents=...)` ([Gemini token guide](https://ai.google.dev/gemini-api/docs/tokens)). Prefer each provider's counter over a third-party “Claude tokenizer” estimate.

## When does token count affect the context window?

It matters whenever you need to fit the whole request and response within a model's limit. The context window is counted in tokens, not words; the prompt and generated output share that space, and reasoning can use some of the output allowance before a visible answer appears ([OpenAI token guide](https://developers.openai.com/api/docs/concepts), [OpenAI reasoning docs](https://developers.openai.com/api/docs/guides/reasoning)).

Use this decision rule: estimate with words only to choose a rough size; count with the intended model before sending a large prompt; leave room for the answer and reasoning; then compare the provider's returned usage with your estimate. For an agent, count the history and tool definitions it resends, not just the latest message. The [context-window guide](/blog/what-is-context-rot) explains why the same conversation can grow even when your newest prompt is short.

## Where do token estimates break?

A local count is useful only when the tokenizer matches the model and the input you count matches what the provider receives. Chat wrappers, tool schemas, images, PDFs, and hidden or automatically added prompt material can make “count the pasted text” a poor proxy. Anthropic's endpoint supports many structured inputs, but documents specific unsupported server tools and warns that its returned value can differ slightly from actual input usage.

| If you need…                              | Use…                                                       | Treat the result as…                                           |
| ----------------------------------------- | ---------------------------------------------------------- | -------------------------------------------------------------- |
| A quick budget for ordinary English prose | ~750 words per 1,000 tokens                                | Estimate only                                                  |
| A local count for an OpenAI model         | Its mapped `tiktoken` encoding                             | Exact for that encoded text; not a full API request accounting |
| A Claude API preflight                    | Anthropic `messages.count_tokens` with the target model ID | Provider's estimate before sending                             |
| The count that was actually billed        | The response usage fields                                  | Post-request record                                            |

Don't infer a closed model's tokenizer from a similar brand or a nearby model. For a Claude API preflight use Anthropic; for an actual hosted call inspect its usage. Octomind is ours: it records provider-reported input, output, cache-read, and cache-write tokens and tracks cost per request and per session. That's usage accounting after the provider responds, not a pre-count for Claude.

[How much do AI agent tokens cost?](/blog/tokenmaxxing-ai-agent-costs) breaks down the repeated-read problem behind long agent sessions.

**[Get Octomind](https://octomind.run)** — see provider-reported token usage and cost per request in an open-source agent runtime.

## FAQ

### What does it mean to buy AI tokens?

In an API bill, tokens are units of processed input and generated output, and the provider charges for usage according to its pricing rules. A product that sells “tokens” as credits may stretch the AI token meaning to cover prepaid balance; check whether its credit maps to model tokens, dollars, or another unit before comparing prices.

### Is one AI token one word?

No. A token can be a whole word, part of a word, punctuation, or characters that include a space. The tokenizer and the text decide the count.

### Does 1 million tokens mean 1 million words?

No. A rough English estimate is about 750,000 words, but our measured English Markdown sample was closer to 696,000 words per million tokens. Code, JSON, language, and model can move the count substantially, so measure the content you plan to send.

### Is Anthropic's Claude token count exact?

Anthropic's preflight endpoint returns an estimate and says actual input usage may differ slightly. It's the right model-specific Anthropic token counter to use before a Claude API call; the response usage is the record after the call.
