Text Similarity Checker

Paste two blocks of text and the checker scores how similar they are. Three algorithms run side-by-side: word Jaccard, character Levenshtein, and cosine. Differences are highlighted word-by-word so you can see exactly where the texts agree and disagree.

Copied to clipboard
0Words in A 0Words in B 0Shared words 0Unique to A 0Unique to B

What each score means

Jaccard (0-100%): set overlap of unique words. Levenshtein: character edit distance; 0% means very different, 100% means identical. Cosine: term-frequency vector angle; like Jaccard but weighted by how often words repeat.

Use case

Compare two drafts to see what changed. Check if a translation preserves wording. Verify a paraphrase is paraphrased enough. Compare two product descriptions for duplicate content risk.

Private by default

Everything runs in your browser.

How to compare two texts for similarity

1. Paste both texts

Drop Text A in the left input box and Text B in the right one. The two can be any length: a sentence vs another sentence, two paragraphs, two articles, two product descriptions.

2. Read the three scores

Three algorithms run side-by-side and give you a percentage match from 0% (completely different) to 100% (identical). Each one weighs the comparison differently; together they triangulate.

3. Check the word diff

The diff panel shows three groups: words that appear in both texts (green), words unique to Text A (blue), words unique to Text B (red). For long texts only the first 50 unique words per side are shown to keep the diff readable.

4. Edit and recompute

The scores update live as you edit either text. Swap with the Swap A <-> B button if you want to flip them. Copy the score breakdown for a report or pull the unique-word lists for further analysis.

Three algorithms, three perspectives

Jaccard similarity (word set overlap)

Builds a set of unique words from each text and computes the overlap as a percentage of the combined set. Jaccard = (words in both) / (words in either). Insensitive to word order and word frequency. 100% means the two texts use exactly the same set of unique words. See also: Sentence Counter.

Levenshtein similarity (character edit distance)

Counts the minimum number of single-character edits (insert, delete, substitute) needed to turn Text A into Text B. Normalized as 1 - (edits / max-length). Sensitive to character-level differences. Useful for catching typos, near-duplicates, or rewordings.

Cosine similarity (term-frequency vector)

Builds a vector of word frequencies for each text, then computes the cosine of the angle between the two vectors. Like Jaccard but weighted by how often each word appears. Words used multiple times count more.

Which score should you trust?

The three scores often disagree, and disagreement is informative.

High Jaccard, low Levenshtein

The texts use the same words but in very different order or with different connecting language. Typical of paraphrases. Use this combo to verify a rewrite has new sentence structure.

Low Jaccard, high Levenshtein

The texts are similar character-by-character but use different vocabulary. Unusual; usually means the texts share a template or boilerplate frame with different content slotted in.

High cosine, lower Jaccard

The texts use roughly the same words and similar frequencies. Common with summaries vs originals: same vocabulary, fewer occurrences.

For most purposes, Jaccard is the easiest to communicate and the most defensible. Levenshtein catches close textual matches. Cosine catches term-frequency similarity even when the absolute texts differ.

Common use cases

Duplicate content audit

Compare two pages on your own site to see whether they accidentally cover the same content. Anything above ~70% Jaccard is a candidate for consolidation, redirect, or rewrite.

Paraphrase verification

If you rewrote a paragraph to avoid copying, run the original and rewrite through the checker. A 30-50% Jaccard means the wording is different; below 30% means you have substantially rewritten.

Translation roundtrip

Translate a paragraph to another language and back to English. The roundtripped version should be highly similar to the original; if Jaccard drops below 50%, the translation introduced too much drift.

Product description differentiation

If you sell multiple variants of similar products, compare their descriptions to make sure you have not just copy-pasted. Aim for under 50% Jaccard between siblings.

Document version diff

Compare draft N and draft N+1 of the same document to see how much changed. Useful when reviewing collaborative writing or external edits.

Limits of similarity scoring

Similarity scores are mechanical. They measure word overlap, not meaning.

Synonyms count as different

"Car" and "automobile" mean the same thing but score as different words. A paraphrase that swaps synonyms can have low Jaccard while being semantically near-identical. Semantic similarity (embedding-based) is a different tool.

Stop words drag scores up

The, of, and, is, in - common words appear in almost every English text and inflate Jaccard. The scores in this tool keep stop words. If you need a stopword-stripped variant, paste through remove duplicate words first or strip via find-and-replace.

Tokenization is lowercase-on-words

The checker tokenizes by stripping punctuation and lowercasing. Apostrophes inside words are kept (don't, won't); en-dashes and em-dashes are treated as word boundaries. URLs and numbers are tokenized as their alphanumeric parts.

Plagiarism vs similarity: what this tool is not

A similarity checker is not a plagiarism checker. The tool compares two texts you supply. It cannot tell you whether one of those texts copies from a third source elsewhere on the internet.

For actual plagiarism detection (comparing one text against the open web), specialized tools (Copyscape, Turnitin, originality.ai) crawl a much larger corpus and report hits. Use this tool for the two-text use case: drafts, paraphrases, variants, document versions.

If you do suspect plagiarism: a 90%+ Jaccard between your text and a competitor's known publication is a strong signal but not legal evidence. Save screenshots, archive URLs, and consult a plagiarism-detection service or legal counsel for actionable proof.

Related tools

Frequently asked questions

What is a high similarity score?

Depends on context. Two paraphrases of the same idea typically land at 30-60% Jaccard. Two versions of the same document with minor edits land at 80-95%. Identical texts score 100%. For duplicate-content audit, anything above 70% Jaccard is worth investigating.

Why do my three scores differ so much?

Each formula weighs word overlap differently. Jaccard measures unique-word set overlap. Levenshtein measures character-level edit distance. Cosine measures term-frequency vectors. Disagreement is informative: high Jaccard + low Levenshtein means same words in different order (paraphrase). Read each score in context.

Does it detect plagiarism?

No. The tool compares two texts that you supply. It cannot detect plagiarism against the open web; that requires a tool with a large crawled corpus (Copyscape, Turnitin, originality.ai). Use this tool for two-text comparisons: drafts, paraphrases, variants.

How are synonyms handled?

They are not. Car and automobile score as different words. The checker measures surface-level lexical similarity, not semantic similarity. For meaning-based similarity, you would need embeddings or a transformer-based tool.

Does case matter?

No. Both texts are lowercased before tokenization, so Apple and apple are the same word for matching purposes. Hyphenated and apostrophe-containing words are tokenized as single units.

Does the text leave my browser?

No. The checker runs entirely in your browser using JavaScript. Both texts and the comparison results are never sent to our servers and are not logged.

Why does Levenshtein get slow on long texts?

Levenshtein is O(n*m) where n and m are the lengths of the two texts. For multi-megabyte inputs, the computation can take several seconds. If you only need Jaccard and cosine (which are O(n)), they remain fast at any input size.

How do I lower the similarity between two texts?

Three levers: (1) use different vocabulary (swap synonyms) - drops Jaccard. (2) Change sentence structure - drops Levenshtein. (3) Change word frequencies - drops cosine. Real rewrites usually do all three. Run the checker after each pass to verify.

More wordcounter.ai tools

Other tools you might find useful.

Browse the full catalog →