Skip to content
ToolShelf

Character counter

Count characters with and without spaces against a limit you set, and see the four different answers to how long a piece of text is when emoji are involved.

Your text

Your text stays in this tab. The work is done by JavaScript in your browser. None of it is uploaded, logged or saved, and the tool keeps working with the network off.

Against a limit

Optional. Platform limits change, so this is yours to set.

What counts as a character

Measure

Characters

0

Start typing and the count follows.

With spaces
0
Without spaces
0
Words
0
Lines
0

The four measures

UTF-16 code units
0
Code points
0
Grapheme clusters
0
UTF-8 bytes
0

All four agree, which is what happens with plain unaccented text. Add an emoji or an accent and they separate.

Four answers, all of them correct

Ask how many characters a piece of text is and there are four reasonable replies. For plain unaccented English they all agree, which is why the question rarely comes up. Add one emoji and they separate immediately.

  • UTF-16 code units. What String.length returns in JavaScript, and therefore what most form validation counts. An emoji is two.
  • Code points. Unicode characters. An emoji is one, a flag is two, and an é is one or two depending on how it was typed.
  • Grapheme clusters. What a reader calls a character and what the caret moves over. A family emoji is one.
  • UTF-8 bytes. What a database column holds. An emoji is four, and a Chinese character is three.

The family emoji makes the point on its own: one thing on screen, seven code points, eleven UTF-16 code units, twenty-five bytes. Every one of those numbers is right.

Which one your limit is using

If a form keeps rejecting text that looks short enough, it is almost certainly counting UTF-16 code units, front-end validation nearly always checks value.length, and a message with a few emoji in it is longer by that measure than it looks. If a database is complaining, try bytes: a VARCHAR(255) may mean 255 bytes rather than 255 characters depending on the engine and its encoding.

Why there are no platform presets

No dropdown of “X post” or “SMS” limits, on purpose. Those numbers change, and several of them are not plain counts at all:

  • X weights most non-Latin characters as two, so the limit depends on what you wrote in.
  • An SMS holds 160 characters in the GSM alphabet and 70 as soon as one character falls outside it, a single curly apostrophe can halve the capacity.
  • Google truncates meta descriptions by pixel width, not by character count, so the commonly quoted figure is a rule of thumb rather than a limit.

A dropdown of stale numbers would look authoritative and mislead. The limit field is yours to set, against whichever measure applies.

It counts in your browser

Nothing is uploaded, logged or stored. The text you are checking against a limit is often a message you have not sent yet, and it stays in the tab.

Questions

Is my text sent anywhere?
No. Every count is worked out in your browser and nothing is uploaded, logged or stored. You can disconnect the network after the page loads and it keeps counting.
Which characters count as spaces?
Every kind of whitespace: ordinary spaces, tabs, line breaks and non-breaking spaces. The 'without spaces' figure removes all of them, so a paragraph split over several lines does not get credit for its line breaks.
Why do different tools give different character counts?
Because 'character' has four reasonable meanings and tools pick different ones. A single emoji is one character to a reader, one code point, two UTF-16 code units and four UTF-8 bytes. All four figures are shown here so you can match whichever your limit actually uses.
Which measure should I use for a form that keeps rejecting my text?
UTF-16 code units, most likely. Front-end validation almost always checks value.length in JavaScript, which counts code units, so a message full of emoji is twice as long as it looks. If a database is rejecting it instead, try UTF-8 bytes.
Why is an emoji two characters?
Because JavaScript strings are UTF-16, and anything outside the first 65,536 Unicode characters is stored as a pair of code units. Emoji all live above that line. Some emoji are worse: a flag is two code points, and a family emoji is several people joined by invisible characters, making eleven code units in total.
Why are there no presets for X, SMS or meta descriptions?
Because they change, and several of them are not plain counts. X weights most non-Latin characters as two. An SMS holds 160 characters in the GSM alphabet but only 70 as soon as one character falls outside it. Google truncates meta descriptions by pixel width, not characters. A dropdown of stale numbers would look authoritative and be wrong; the field is yours to set.
What is a grapheme cluster?
What a person means by a character: one thing on screen, one press of the arrow key to move past. An 'é' typed as e plus a combining accent is two code points and one grapheme. This is the figure closest to what you see, and the one almost no counter reports.
Why does the grapheme count sometimes say it is unavailable?
It needs Intl.Segmenter, which every current browser has but some older ones do not. Rather than substitute a different measure and label it wrongly, the figure is left out and the other three still work.

More tools