Unicode lets the same visual character be encoded in more than one way. This can create invisible mismatches that break search, file systems, and security checks.
One letter, two encodings
Take the letter é. It can be encoded as a single code point U+00E9 (composed), or as the letter e followed by a combining acute accent U+0301 (decomposed). They look the same. Many systems treat them as different strings.
What normalization does
Unicode defines four standard forms: NFC, NFD, NFKC, and NFKD. They describe how to convert equivalent sequences into a canonical form. NFC is the most common choice on the web.
Why it matters
- Two URLs that look identical may route to different servers because one uses a composed character and the other uses decomposed characters.
- A username check may fail if one side normalizes and the other does not.
- Search indexes may miss content encoded in a different normalization form.
What the scanner does today
This tool reports when the text contains characters that would change under NFC normalization. It does not rewrite the text automatically; it only points out where a mismatch could happen. You decide whether to normalize.
Tip: If you are building a system that compares text, normalize both sides to the same form before comparing. Do not assume two visually identical strings have the same bytes.