Text Length Calculator

Paste any string to see its character count, word count, byte length in UTF-8 and UTF-16, and Unicode code point count. Useful for database column limits, API payload sizes, and SMS message budgeting.

Length

0
Characters
0
Characters (no spaces)
0
Words
0
UTF-8 bytes
0
UTF-16 bytes
0
Unicode code points

Character count vs byte length vs code points

For ASCII text (English letters, digits, basic punctuation) all three numbers are identical: one character = one code point = one byte in UTF-8 and two bytes in UTF-16. For anything outside ASCII the numbers diverge, sometimes dramatically.

An emoji like πŸš€ is one perceived character but four UTF-8 bytes, two UTF-16 code units, and one Unicode code point. A composed accented letter like Γ© (using a combining acute accent) is one perceived character but two code points and four UTF-8 bytes. A flag emoji πŸ‡ΊπŸ‡Έ is one perceived character but two code points and eight UTF-8 bytes.

The right metric depends on where the text is going:

  • Database VARCHAR length. Use character count (the JS String.length, which is UTF-16 code units, matches most databases' VARCHAR semantics).
  • API payload size. Use UTF-8 bytes - that's what gets transmitted over the wire.
  • SMS budget. Use UTF-16 bytes if any non-GSM-7 character is present (the whole message switches to UCS-2 encoding).
  • Twitter character count. Twitter uses a custom weighting close to Unicode code points but with some quirks; the count on this tool will be within a few of Twitter's display.

Word count and how it differs from character count

Word count is straightforward in English: a word is a maximal run of non-whitespace characters separated by whitespace. The calculator uses that definition (the \S+ regex). Hyphenated words like "well-being" count as one word; em-dash-joined phrases like "well - being" count as three tokens, which is the modern convention.

Some surfaces care about characters (most social platforms, SMS, meta descriptions), some care about words (essays, articles, paid writing). For full word-counter features including reading time, speaking time, and keyword density, use the main word counter.

Common length limits and how to budget them

The numbers you bump into most often:

  • Twitter / X: 280 characters (standard), 25,000 (Premium).
  • Instagram caption: 2,200 characters.
  • LinkedIn post: 3,000 characters.
  • Facebook post: 63,206 characters (effectively unlimited).
  • Meta title (SEO): ~60 characters before truncation.
  • Meta description (SEO): ~155 characters visible.
  • SMS (GSM-7): 160 characters.
  • SMS (UCS-2): 70 characters once any non-GSM character is present.
  • Email subject: 78 characters before truncation in most clients.
  • Database VARCHAR: anywhere from 32 to 65,535 depending on schema.
  • Twitter username: 15 characters.

If you need live limit-checking against a specific platform, use the character counter which highlights when you cross the limit. For Twitter specifically, use the Twitter character counter.

Why the byte counts matter for developers

If you're sizing a database column, an API request body, or a URL parameter, you need bytes, not characters. JavaScript's String.length returns UTF-16 code units, which over-counts non-BMP characters (emoji, rare scripts) by counting each surrogate pair as two. The UTF-8 byte count is what almost every server-side language uses for storage and bandwidth. See also: find-and-replace.

The Unicode code point count (the third metric on this page) is what you want for length validation that respects user perception - it counts emoji as 1, not 2. This is the same count String.prototype.codePointAt + iterator-based length give you in JavaScript.

Frequently asked questions

What's the difference between characters and bytes?

For ASCII text (A-Z, 0-9, basic punctuation), one character = one byte in UTF-8. For non-ASCII text (accents, emoji, Chinese characters), one character can be 2, 3, or 4 bytes in UTF-8. Use characters for visible length, bytes for storage and transmission.

Does this counter match Twitter's character count?

Close but not identical. Twitter uses a custom weighting that counts URLs as 23 characters regardless of length and applies Unicode normalization. This tool's character count uses JavaScript's String.length, which is UTF-16 code units - within a few of Twitter's display for most text.

Why are emoji counted differently in different metrics?

An emoji like πŸš€ is one perceived character but two UTF-16 code units (a surrogate pair) and four UTF-8 bytes. A flag emoji like πŸ‡ΊπŸ‡Έ is one perceived character but two code points and eight UTF-8 bytes. The "Unicode code points" metric is the closest to perceived length.

What's a code point?

A Unicode code point is a single abstract character, regardless of how many bytes it takes to encode. "A" is code point U+0041. "Γ©" is code point U+00E9. Most modern programming languages count code points when you ask for "string length" in their idiomatic API.

Does the count include whitespace?

Yes for character count. The "characters (no spaces)" metric excludes whitespace if you need length without it. Word count splits on whitespace, so leading and trailing whitespace are stripped before counting words.

What VARCHAR length should I use for a column?

A safe default is the character count of your longest expected input Γ— 1.5 for buffer. For internationalized text, allow 4 bytes per character (the UTF-8 max) and size the BYTE limit accordingly.

Is text uploaded for the count?

No. The count runs entirely in JavaScript in your browser. The input never leaves the page.

Does the tool handle very large texts?

For texts under ~1 million characters, the count is instant. Above that, JavaScript's string operations may slow down on older devices. For massive corpora, run a server-side counter.

Related tools

More wordcounter.ai tools

Other tools you might find useful.

Browse the full catalog β†’