Token Counter: How Counting Works and the Fastest Way to Do It
A token counter answers two questions before you send anything: does this prompt fit inside the context window, and what is this call going to cost you.
Counting happens here in the page. When you want to run a model instead of measuring text, synexa.ai is free to start and bills per run.
What the number is actually measuring
A token is the unit a model reads in, and it is neither a word nor a character. Both of the things you care about are denominated in it: the context window that decides whether a request is even accepted, and the bill that arrives at the end of the month. A counter estimates how much text a model will process before you send anything, which is why it belongs at the start of a prompt-building session rather than after the first error. The better tools report tokens alongside word count, character count with spaces, and character count without, because those four numbers together tell you whether a prompt is genuinely long or just verbose. Only the first one costs money, but the gap between them is usually where the trimming is.
Local tokenisers versus uploading your text
The distinction that matters most when the text is a customer contract or an unreleased spec: does the tool run the tokenizer in your browser, or does it post the content to a server first. Good ones say so plainly, state that your text stays in the browser, and render a live visualisation of the pieces as the local tokenizer reads through your input. That visualisation is worth watching once even if you never need it again, because seeing where a long identifier or an unusual name splits into four pieces explains more about token cost than any article will. If a tool does not state where the counting happens, assume the text leaves your machine and behave accordingly.
The cost estimate is where it gets subtle
One count feeds a comparison across providers, and the decent calculators cover major APIs from three vendors, with roughly ten models listed for each. Three details decide whether that number is useful. Input, cached input and output are billed at different rates, so a single blended figure is close to meaningless for a chat workload with a large fixed system prompt. Long-context tiers can apply automatically once a request crosses a threshold, which is why estimates jump without the prompt changing much. And the standard estimate covers paid text rates only, excluding cache writes and storage, batch and priority service tiers, tool calls and media charges. In one snapshot the lowest listed input rate was 0.1 dollars per million tokens, which tells you how wide the spread between tiers is.
Prose, code and JSON do not behave alike
Counting a blog post and counting a payload are different exercises. Structured data uses tokens differently from prose: punctuation-heavy JSON, long snake-case identifiers, base64 fragments and repeated key names all produce far more tokens per visible character than English sentences do. The practical consequence is that a request which looks small in a code editor can be the expensive half of your bill. Paste the real payload rather than a trimmed example, check what the visualisation does to your key names, and consider whether the model needs the whole structure or only three fields from it. The same applies to logs and transcripts, which compress badly and are usually pasted in full out of habit rather than need.
Counting is planning, not the job
Once a prompt fits and the estimate is acceptable, the counter has done everything it can do. What it never tells you is whether the output is any good, which is the only question that decides whether the spend was worth it. It also does not help at all for the non-text half of most projects. If what you need is an image, a video clip or generated audio, token maths does not apply: those models bill per run, not per token. Synexa hosts that side through one REST endpoint and a Python SDK, so a single job costs a known amount and you can measure quality before committing to a pipeline. Count the prompt, then go and run something.
What a good counter gives you
Four numbers, not one
Tokens, words, characters with spaces, and characters without them. The gap between the token count and the word count is where most of the trimming opportunity quietly hides.
Local by default
The better tools run the tokenizer in your browser and say so, which means a contract or an unreleased spec never leaves your machine just to be measured.
Cost across providers
One count compared against major APIs from three vendors, with input, cached input and output priced separately rather than blended into a single misleading figure.
Estimates, with exclusions
Standard paid text rates only. Cache writes and storage, batch and priority tiers, tool calls and media charges sit outside the number, so real invoices can run higher.
Measuring text against running a model
| Feature | Token counter tools | Synexa |
|---|---|---|
| What it answers | Will this prompt fit, and what will it cost | What does the model actually produce |
| Unit | Tokens, priced separately for input, cache and output | One run, billed per job |
| Where it runs | In the browser, on your own text | Hosted, through one REST endpoint or the Python SDK |
| Covers media | No, text only | Image, video and audio models |
| Accuracy | Close, and model-family dependent | Exact, because a run either happened or it did not |
| When to use it | Before writing the request | When you need the output, not the estimate |
Counting properly in four moves
- Paste the real thing
Use the actual prompt, document or payload rather than a shortened example. Structured data and long identifiers count very differently from the prose you tested with. - Read tokens against words
Compare the token number with the word and character counts. A wide gap means punctuation, identifiers or formatting are doing the spending, not your instructions. - Split input from output
Price the fixed system prompt, the variable user input and the expected response separately. Cached input is billed differently again, which changes chat maths entirely. - Trim, then re-count
Cut repeated instructions, drop fields the model never uses, or move to a cheaper tier. Measure again rather than assuming the edit helped.
FAQ
How accurate is an online token counter?
Close enough for planning, not exact. Tokenisation differs between model families, so a browser tool gives you an estimate rather than the number the provider will bill against. The official tokenizer page did not load during this research, so treat any third-party count as approximate and leave headroom in a context-window calculation.
Does a token counter send my text anywhere?
It depends on the tool, and the good ones state it plainly. Browser-based calculators run the tokenizer locally and keep the text on your device, with a live visualisation of how it splits. If a tool does not say where counting happens, assume the content is uploaded and do not paste anything confidential.
Why is my JSON so much more expensive than my prose?
Structured data tokenises badly. Punctuation, repeated key names, long identifiers and encoded blobs all produce many more tokens per visible character than ordinary sentences. Paste the real payload into a counter, then ask whether the model needs the whole object or only a few fields from it.
What does the cost estimate leave out?
Standard calculators quote paid text rates and exclude cache writes and storage, batch and priority service tiers, tool calls and media charges. Long-context pricing tiers may also apply automatically once a request passes a threshold, which is a common reason an estimate jumps unexpectedly.
How many words is a million tokens?
Only roughly answerable, which is why conversion guides exist as a separate page on most calculator sites. The ratio shifts with language, formatting and how much code or punctuation is in the text, so convert with a counter on a real sample rather than applying a rule of thumb.
I need images or video, not text. Does any of this apply?
No. Those models bill per run rather than per token, so counting does not help you. Synexa hosts image, video and audio models behind a single REST endpoint with a Python SDK, and each job has a known cost, which makes budgeting a matter of counting runs instead.
Stop measuring text and run something
Synexa puts image, video and audio models behind one REST endpoint and a Python SDK, billed per run. One request, one known cost, and output you can judge.
See Synexa →