Private Token Counting for OpenAI & Llama 3
A token is the unit an LLM uses to process text. Before you send a prompt, count its tokens to check the context-window budget, estimate supported input costs, and understand where the tokenizer splits code, punctuation, and non-English text.
Token Counter runs in your browser. The text you paste stays out of shareable URLs and is not sent to the tokenization service. Choose the tokenizer that matches the model you plan to use, then paste or type the raw text you want to measure.
Choose the Matching Tokenizer
Different models can split the same text differently. This tool intentionally lists only tokenizers it can run accurately on-device.
- GPT-3.5 Turbo & GPT-4 Turbo: use the
cl100k_baseencoding. - GPT-4o & o1 Preview: use the
o200k_baseencoding. - Llama 3: uses the published Llama 3 tokenizer, with hosting-specific costs left to your provider.
The token heatmap shows the boundaries used by the selected tokenizer. For very large inputs, it displays the first 2,000 token boundaries while preserving the exact total token count.
Understand What the Cost Estimate Covers
For supported OpenAI models, the estimate uses the displayed input-token price. It is a planning aid, not a billing quote. Output tokens, cached tokens, batch discounts, tool calls, images, and chat-message formatting can change an API invoice.
Use the estimate to compare prompt alternatives before you send a request. Confirm final pricing and model limits in your provider’s documentation before production use.
A Practical Prompt-Review Workflow
Start with the exact model tokenizer, paste the raw prompt, check the context-window percentage, and review the colored boundaries where wording or formatting looks unexpectedly expensive. If you are counting a complete chat request, remember to account for system instructions and message-format overhead separately.
The tool is free to use without an account. For more details on local processing and analytics, read the Privacy Policy.