Scan, check, and remove hidden Unicode characters, bidirectional tricks, and mixed-script lookalikes locally in your browser. Use it as an AI watermark detector, invisible character checker, or text cleaner. No upload. No account.
Drop a plain-text file here to check and clean, or click to browse
.txt, .md, .csv, .json, .html, .css, .js, .ts, .py, .java, .go, .xml, .log
Drop in text from a chat app, a document, a code file, or a CSV export. Everything stays in your browser — the file is never sent to a server.
The scanner looks for invisible characters, BiDi overrides, mixed-script identifiers, and normalization gaps. Each hit gets a severity, a rule ID, and an exact position.
Read the report first. Remove only the characters you want, copy the cleaned text, or download the result. The original stays untouched.
When people search for an AI watermark detector, checker, or remover, they often mean very different things. This tool does one thing well: it finds the hidden text signals that can travel with copied content, then lets you remove them. Use it as an invisible character checker, Unicode text cleaner, or watermark scanner. It does not score how "AI-like" a paragraph reads, and it does not access any vendor's watermark key.
| Signal | What it means | Example |
|---|---|---|
| Invisible characters | Zero-width spaces, word joiners, soft hyphens, BOM, and other non-printing code points that can be inserted between visible characters. | U+200B, U+FEFF, U+202F |
| Bidirectional controls | Characters that change visual reading order without changing the logical byte order. Known as Trojan Source vectors in code. | U+202E Right-to-Left Override |
| Mixed scripts / lookalikes | Identifiers or URLs that mix writing systems or use characters that look identical to Latin letters, such as Cyrillic а vs Latin a. | pаypal with Cyrillic а |
| Normalization gaps | The same visual letter can be encoded as one code point or as a base letter plus a combining mark. This creates invisible mismatches in search, URLs, or comparisons. | é as U+00E9 vs e + U+0301 |
What it will not do: It cannot detect statistical watermarks like SynthID Text or pixel-level watermarks in images. Those methods do not leave characters in the text. No scanner without the vendor's key can detect them. We say this up front because a misleading claim is worse than a missing feature.
Check and clean text before it goes into a CMS, email template, or documentation site. Remove hidden spaces that can break URLs, formatting, and search indexing.
BiDi overrides can make source code look safe to a human reviewer while compiling into something different. The scanner flags these as high-risk findings so you can remove them before merging.
Byte order marks and non-breaking spaces in CSV, JSON, or log files often break parsers. Check for them and remove them before they reach your database.
If you suspect copied text carries invisible marks, this tool gives you a concrete report. It does not give a human-or-AI verdict — just evidence you can check and remove.
A hidden character is a real Unicode code point in the text, such as a zero-width space. Some AI systems may place these in generated output, but many do not. A statistical watermark, such as Google's SynthID Text, changes the probability of which word the model chooses next. It leaves no character to scan. This tool finds the first kind, not the second.
No. It only means this tool did not find any of the signals it looks for. The text may never have been marked, the marks may have been stripped by reformatting, or the marking method may not leave characters at all.
Only if the vendor left visible or invisible characters in the text. The major model providers mostly use statistical watermarking that requires their own keys to detect. Any third-party website that claims to detect those watermarks without access to the vendor's system is making a claim it cannot verify.
It is a way to hide malicious logic in source code using bidirectional Unicode characters. The code looks correct to a reviewer but compiles differently. This scanner flags BiDi controls near code so you can review them before merging.
Some characters from different writing systems look identical. A Cyrillic lowercase "a" looks like a Latin "a" but has a different code point. This can be used to spoof domains, identifiers, or URLs. The scanner warns when a single word mixes scripts that contain lookalikes.
Unicode allows the same visual character to be encoded in more than one way. Search engines, file systems, and databases may treat the two forms as different strings. The scanner reports when the text contains a form that could mismatch a normalized copy.
No. The scan runs entirely in your browser using JavaScript. Your text, files, and results never leave your device. We do not store or log anything.
Yes. The tool is free to use for personal and commercial work. Once the page loads, you can disconnect from the internet and it will still work because nothing is processed online.
Any plain-text file: .txt, .md, .csv, .json, .html, .css, .js, .ts, .py, .java, .go, .xml, .log, and similar. The file is read locally in your browser. You can also paste text directly.
After the scan, choose which rules to apply and click the clean action. The cleaned text appears below the report so you can compare it to the original before copying or downloading.
Yes. After the scan, enable every rule and run the clean action to remove all flagged characters in one pass. You can also keep specific rules disabled if you want to preserve certain formatting, such as non-breaking spaces in a prepared document.
It is both a checker and a remover for the kind of watermark that exists as characters in the text. You can scan first, review the findings, and then remove the flagged characters. It is not a remover for statistical watermarks, because those do not leave characters to delete.