Some characters look identical but are different Unicode code points. A Cyrillic а looks just like a Latin a, and a zero-width space is invisible. Attackers use these to fake domain names, usernames and filenames, and they also cause hard-to-find bugs in code. Paste text below to find look-alike characters, words that mix scripts, and hidden characters, or compare two strings to see if they only look the same.
Many Unicode characters are drawn almost identically to one another, and some have no visible shape at all. The checker looks at every character in your text, then judges it in the context of the word it sits in, because the same letter can be harmless in one place and dangerous in another.
A strong sign that the text is not what it appears to be. This level is used for:
Why it matters: this is how fake domains, impersonated usernames, disguised file extensions and "Trojan Source" code attacks work. Do not trust the text until you have checked it.
Unusual, and sometimes intentional, but not proof of a problem on its own. This level is used for:
Why it matters: a look-alike semicolon in source code can break a build, and styled letters can slip past search, filters and username rules.
The character resembles a Latin letter, but the surrounding text is ordinary writing in one script. This level is used for:
Why it matters: it usually does not. These characters are highlighted so you can see them, and the cleaner leaves them unchanged so genuine text is never damaged.
Cyrillic, Greek, Armenian, Cherokee, Lisu and other letters that resemble Latin letters, digits and punctuation, based on Unicode's confusables data (UTS #39).
Fullwidth, mathematical, circled, small-form and Roman numeral characters, found through Unicode compatibility normalization (NFKC).
Words containing letters from more than one script. Normal combinations such as Han with Hiragana and Katakana (Japanese) or Han with Hangul (Korean) are not flagged.
Zero-width characters, bidirectional controls, soft hyphens, filler characters and tag characters. Emoji joiners and flag sequences are recognised and ignored.
The built-in list covers the most common look-alikes and is not a complete copy of the Unicode confusables file. Look-alikes within plain ASCII, such as 0 and O or l, I and 1, and letter pairs such as rn and m, are not reported. How similar two characters look also depends on the font. A clean result is therefore a good sign, but not a guarantee that text is safe.