Token counter

Count plain-text tokens with o200k_base or cl100k_base. See token boundaries, copy IDs and download JSON with every token.

On your device · No account
Privacy details

Token counter processes your input in this browser. Inputs and results are excluded from analytics. Recent tools save tool names, visit counts and last-visit times. Saved tools store only their names. Draft saving is optional and stays on this device; delete saved text with the draft controls. Loading text from a URL is optional and contacts that server; its URL and your IP address are visible to the server. An optional “Use in” action keeps a handoff in this tab. The next tool removes it on arrival and ignores it if more than a minute has passed.

Copy link includes the inputs you choose to share in the address. Anyone with that link can reopen them. Files and generated results are excluded. Password and key tools share settings only.

Counts plain text using the selected encoding. Chat roles, tools, images and message wrappers add tokens outside this count.

Your token count

See token boundaries and IDs

␠ is a space, ↵ a line break, and ⇥ a tab. A token can split a Unicode character; � in a piece does not change the original bytes. Hover a piece for its ID and bytes.

Reset reloads the tool with its starting values and releases open files and device controls. Saved drafts stay on this device until you delete them.

Links include your inputs and settings. Anyone with the link can see them, so leave out private text. Files and generated results are left out; random draws run again.

Keyboard shortcuts

Updated

What this token count includes

Choose the exact encoding required by your application. The tokenizer splits your text into UTF-8 byte pieces and merges them using the selected vocabulary. Whitespace, capitalization and punctuation can change the tokens even when the words look similar.

The count covers only the pasted text. It excludes chat roles, system wrappers, tools, images, audio and other request overhead. Encoding names are shown directly so a changing model name does not silently select the wrong vocabulary.

Open the token details to inspect boundaries and IDs. Exported token IDs reconstruct the UTF-8 bytes of the text; the tool checks that match before enabling export. Special-token-looking strings are counted as ordinary text.

Next: Character counter

A worked example

“hello world” is 2 tokens in both encodings, but the IDs differ: cl100k_base gives [15339, 1917] and o200k_base gives [24912, 2375].

Next: Word counter

Before you use the result

This is not a billed API-usage count or a context-window guarantee. It supports the two named base encodings only. Tokens may split a Unicode character; an individual preview piece can show a replacement symbol even when the combined bytes are exact. Isolated invalid UTF-16 surrogates normalize to the replacement character.

Cite this page

ToolOctopus. “Token counter.” Updated 2026-09-20. https://tooloctopus.com/token-counter.

Add an access date if your instructions require one.

Citation formatting uses citeproc-js by Frank Bennett and Citation Style Language styles. Licence and source code.

Find a tool

Search by task. Use ↓ and ↑ to choose, Enter to open.