AI Watermark Detector, Checker & Text Cleaner

Scan, check, and remove hidden Unicode characters, bidirectional tricks, and mixed-script lookalikes locally in your browser. Use it as an AI watermark detector, invisible character checker, or text cleaner. No upload. No account.

Drop a plain-text file here to check and clean, or click to browse

.txt, .md, .csv, .json, .html, .css, .js, .ts, .py, .java, .go, .xml, .log

▶ Scan rules (4 groups enabled)

How the Scan Works

Paste or Upload

Drop in text from a chat app, a document, a code file, or a CSV export. Everything stays in your browser — the file is never sent to a server.

Check Four Signals

The scanner looks for invisible characters, BiDi overrides, mixed-script identifiers, and normalization gaps. Each hit gets a severity, a rule ID, and an exact position.

Review Before Cleaning

Read the report first. Remove only the characters you want, copy the cleaned text, or download the result. The original stays untouched.

What the Scanner Actually Checks

When people search for an AI watermark detector, checker, or remover, they often mean very different things. This tool does one thing well: it finds the hidden text signals that can travel with copied content, then lets you remove them. Use it as an invisible character checker, Unicode text cleaner, or watermark scanner. It does not score how "AI-like" a paragraph reads, and it does not access any vendor's watermark key.

SignalWhat it meansExample
Invisible characters Zero-width spaces, word joiners, soft hyphens, BOM, and other non-printing code points that can be inserted between visible characters. U+200B, U+FEFF, U+202F
Bidirectional controls Characters that change visual reading order without changing the logical byte order. Known as Trojan Source vectors in code. U+202E Right-to-Left Override
Mixed scripts / lookalikes Identifiers or URLs that mix writing systems or use characters that look identical to Latin letters, such as Cyrillic а vs Latin a. pаypal with Cyrillic а
Normalization gaps The same visual letter can be encoded as one code point or as a base letter plus a combining mark. This creates invisible mismatches in search, URLs, or comparisons. é as U+00E9 vs e + U+0301

What it will not do: It cannot detect statistical watermarks like SynthID Text or pixel-level watermarks in images. Those methods do not leave characters in the text. No scanner without the vendor's key can detect them. We say this up front because a misleading claim is worse than a missing feature.

Why This Exists

For Publishing

Check and clean text before it goes into a CMS, email template, or documentation site. Remove hidden spaces that can break URLs, formatting, and search indexing.

For Code Review

BiDi overrides can make source code look safe to a human reviewer while compiling into something different. The scanner flags these as high-risk findings so you can remove them before merging.

For Data Import

Byte order marks and non-breaking spaces in CSV, JSON, or log files often break parsers. Check for them and remove them before they reach your database.

For Verification

If you suspect copied text carries invisible marks, this tool gives you a concrete report. It does not give a human-or-AI verdict — just evidence you can check and remove.

Common Questions

What is the difference between an AI watermark and a hidden character?

A hidden character is a real Unicode code point in the text, such as a zero-width space. Some AI systems may place these in generated output, but many do not. A statistical watermark, such as Google's SynthID Text, changes the probability of which word the model chooses next. It leaves no character to scan. This tool finds the first kind, not the second.

Does "no hidden characters" mean a human wrote the text?

No. It only means this tool did not find any of the signals it looks for. The text may never have been marked, the marks may have been stripped by reformatting, or the marking method may not leave characters at all.

Can this detect ChatGPT, Claude, or Gemini watermarks?

Only if the vendor left visible or invisible characters in the text. The major model providers mostly use statistical watermarking that requires their own keys to detect. Any third-party website that claims to detect those watermarks without access to the vendor's system is making a claim it cannot verify.

What is a Trojan Source attack?

It is a way to hide malicious logic in source code using bidirectional Unicode characters. The code looks correct to a reviewer but compiles differently. This scanner flags BiDi controls near code so you can review them before merging.

What are mixed-script lookalikes?

Some characters from different writing systems look identical. A Cyrillic lowercase "a" looks like a Latin "a" but has a different code point. This can be used to spoof domains, identifiers, or URLs. The scanner warns when a single word mixes scripts that contain lookalikes.

Why does normalization matter?

Unicode allows the same visual character to be encoded in more than one way. Search engines, file systems, and databases may treat the two forms as different strings. The scanner reports when the text contains a form that could mismatch a normalized copy.

Is my data sent anywhere?

No. The scan runs entirely in your browser using JavaScript. Your text, files, and results never leave your device. We do not store or log anything.

Can I use this commercially or offline?

Yes. The tool is free to use for personal and commercial work. Once the page loads, you can disconnect from the internet and it will still work because nothing is processed online.

What file types are supported?

Any plain-text file: .txt, .md, .csv, .json, .html, .css, .js, .ts, .py, .java, .go, .xml, .log, and similar. The file is read locally in your browser. You can also paste text directly.

How do I clean the text after scanning?

After the scan, choose which rules to apply and click the clean action. The cleaned text appears below the report so you can compare it to the original before copying or downloading.

Can I remove every hidden character at once?

Yes. After the scan, enable every rule and run the clean action to remove all flagged characters in one pass. You can also keep specific rules disabled if you want to preserve certain formatting, such as non-breaking spaces in a prepared document.

Is this an AI watermark remover or just a checker?

It is both a checker and a remover for the kind of watermark that exists as characters in the text. You can scan first, review the findings, and then remove the flagged characters. It is not a remover for statistical watermarks, because those do not leave characters to delete.

From the Docs