Malay, code and text with many numbers or symbols often use more tokens per word than English prose. Treat the result as an estimate only.
Frequently asked questions
What is a token?
A piece of text a language model reads, often a short word or part of a word. On average, one token is about 4 characters or three quarters of a word in English.
How accurate is this?
It is a rough guide. The real count depends on the model’s tokenizer, the language and the content, and code, numbers and other languages use more tokens.
How can I get an exact count?
Use the tokenizer or token-counting tool from your AI provider, which uses the model’s own rules.