Back

Hidden character inspector

Text tools

Loading

Loading tool

The tool is loaded only when you open it.

All processing for this tool happens in your browser. Your input is not sent to a server.

About this tool

Text can look identical while containing different code points. Inspect pasted text or a local UTF-8 file to locate zero-width characters, non-breaking spaces, directional controls and malformed UTF-16 units. The inspector uses pinned Unicode 17.0.0 data and shows ASCII-only escaped evidence so directional characters cannot reorder the diagnostic text. Every finding starts unselected. Decide which exact occurrences to remove while keeping the original unchanged; a finding is a review prompt, not proof of malicious text.

Common uses

  • Find a non-breaking space or zero-width space that makes an identifier, search term or copied value fail an exact comparison.
  • Review directional controls and unexpected separators in copied source text without executing it or displaying those controls inside the report.
  • Distinguish intentional emoji joiners, script-shaping joiners and variation selectors from unwanted copy-and-paste artifacts before deleting anything.

How to use it

  1. 1.Paste text or open one strictly decoded UTF-8 text file. Inspect the original source before making a selection. Literal spellings such as \u200B are ordinary ASCII input; examples below use escaped JSON strings only to make the exact test data readable. Run the inspection and review code point, Unicode name or diagnostic label, category and exact location. The table uses 50 rows per page, but counts and the JSON report cover every occurrence. The escaped source and cleaned previews cover at most 4,000 input UTF-16 units each.
  2. 2.Select individual rows by their original UTF-16 offsets. Nothing is preselected, and selecting one occurrence does not select every matching character. Only those complete units or code points are removed; spaces, line endings and all other original text stay as supplied.
  3. 3.Review the escaped cleaned preview and the remaining findings. Download the complete ASCII-only JSON report for diagnostics. Explicitly review before copying or downloading raw cleaned text; those actions remain unavailable if any unpaired surrogate survives. Input changes invalidate the previous inspection and selection.

Exact, reproducible character inspections

Keep NBSP and remove a zero-width space

input (ASCII JSON): "A\u00A0B\u200BC"
selected UTF-16 offsets: [3]
{"codePointCount":5,"occurrences":[{"codePoint":"U+00A0","utf16Offset":1,"codePointIndex":1,"line":1,"column":2},{"codePoint":"U+200B","utf16Offset":3,"codePointIndex":3,"line":1,"column":4}],"cleanedAscii":"\"A\\u00A0BC\"","rawExportable":true}

Only offset 3 is selected. The NBSP at offset 1 remains, so the cleaned output is not the ordinary ASCII spelling A BC.

Remove only the first identical occurrence

input (ASCII JSON): "x\u200B\u200By"
selected UTF-16 offsets: [1]
{"codePointCount":4,"occurrences":[{"codePoint":"U+200B","utf16Offset":1,"codePointIndex":1,"line":1,"column":2},{"codePoint":"U+200B","utf16Offset":2,"codePointIndex":2,"line":1,"column":3}],"cleanedAscii":"\"x\\u200By\"","rawExportable":true}

Two consecutive U+200B characters produce two rows. Selecting offset 1 leaves the character originally at offset 2 intact.

Review and explicitly remove directional controls

input (ASCII JSON): "a\u202Eb\u202Cc"
selected UTF-16 offsets: [1,3]
{"codePointCount":5,"occurrences":[{"codePoint":"U+202E","utf16Offset":1,"codePointIndex":1,"line":1,"column":2},{"codePoint":"U+202C","utf16Offset":3,"codePointIndex":3,"line":1,"column":4}],"cleanedAscii":"\"abc\"","rawExportable":true}

The source contains RIGHT-TO-LEFT OVERRIDE and POP DIRECTIONAL FORMATTING. Both selected controls are removed; the letters remain in original logical order.

Keep emoji joiners and a variation selector

input (ASCII JSON): "\uD83D\uDC69\u200D\uD83D\uDCBB\u0645\u06CC\u200C\u0631\u0648\u0645\u2764\uFE0F"
selected UTF-16 offsets: []
{"codePointCount":11,"occurrences":[{"codePoint":"U+200D","utf16Offset":2,"codePointIndex":1,"line":1,"column":2},{"codePoint":"U+200C","utf16Offset":7,"codePointIndex":5,"line":1,"column":6},{"codePoint":"U+FE0F","utf16Offset":12,"codePointIndex":10,"line":1,"column":11}],"cleanedAscii":"\"\\uD83D\\uDC69\\u200D\\uD83D\\uDCBB\\u0645\\u06CC\\u200C\\u0631\\u0648\\u0645\\u2764\\uFE0F\"","rawExportable":true}

The woman-technologist emoji uses ZWJ, the Persian word uses ZWNJ, and the heart uses VS16. Supplementary emoji occupy two UTF-16 units each. All three format characters are reported but remain unselected, preserving these intentional sequences.

Remove a space and tab while retaining CRLF

input (ASCII JSON): "a b\tc\r\nd"
selected UTF-16 offsets: [1,3]
{"codePointCount":8,"occurrences":[{"codePoint":"U+0020","utf16Offset":1,"codePointIndex":1,"line":1,"column":2},{"codePoint":"U+0009","utf16Offset":3,"codePointIndex":3,"line":1,"column":4},{"codePoint":"U+000D","utf16Offset":5,"codePointIndex":5,"line":1,"column":6},{"codePoint":"U+000A","utf16Offset":6,"codePointIndex":6,"line":1,"column":7}],"cleanedAscii":"\"abc\\r\\nd\"","rawExportable":true}

SPACE and TAB are selected. CR and LF remain byte-for-byte when encoded as UTF-8; their original positions share line 1 at columns 6 and 7.

Distinguish CRLF, line separator and paragraph separator

input (ASCII JSON): "A\r\nB\u2028C\u2029D"
selected UTF-16 offsets: [4]
{"codePointCount":8,"occurrences":[{"codePoint":"U+000D","utf16Offset":1,"codePointIndex":1,"line":1,"column":2},{"codePoint":"U+000A","utf16Offset":2,"codePointIndex":2,"line":1,"column":3},{"codePoint":"U+2028","utf16Offset":4,"codePointIndex":4,"line":2,"column":2},{"codePoint":"U+2029","utf16Offset":6,"codePointIndex":6,"line":3,"column":2}],"cleanedAscii":"\"A\\r\\nBC\\u2029D\"","rawExportable":true}

The original line locations count CRLF once and each Unicode separator once. Removing U+2028 joins B and C; the retained U+2029 still separates the next paragraph.

Keep decomposed and composed spellings distinct

input (ASCII JSON): "e\u0301\u00E9"
selected UTF-16 offsets: []
{"codePointCount":3,"occurrences":[],"cleanedAscii":"\"e\\u0301\\u00E9\"","rawExportable":true}

The combining acute accent U+0301 is outside this detection set, and no character is selected. The decomposed e plus accent and the precomposed U+00E9 remain distinct; there is no normalization.

One remaining unpaired surrogate blocks raw export

input (ASCII JSON): "A\uD800B\uDC00C"
selected UTF-16 offsets: [1]
{"codePointCount":5,"occurrences":[{"codePoint":"U+D800","utf16Offset":1,"codePointIndex":1,"line":1,"column":2},{"codePoint":"U+DC00","utf16Offset":3,"codePointIndex":3,"line":1,"column":4}],"cleanedAscii":"\"AB\\uDC00C\"","rawExportable":false}

The high surrogate is removed but the low surrogate remains unpaired. The escaped diagnostic output is available; raw copy and UTF-8 download remain blocked.

Remove both malformed units explicitly

input (ASCII JSON): "A\uD800B\uDC00C"
selected UTF-16 offsets: [1,3]
{"codePointCount":5,"occurrences":[{"codePoint":"U+D800","utf16Offset":1,"codePointIndex":1,"line":1,"column":2},{"codePoint":"U+DC00","utf16Offset":3,"codePointIndex":3,"line":1,"column":4}],"cleanedAscii":"\"ABC\"","rawExportable":true}

Both unpaired surrogate occurrences are selected. The output is exactly ABC, and raw export becomes eligible after the separate review step.

Retain a supplementary variation selector after removing BOM

input (ASCII JSON): "\uFEFF\uD83D\uDE00\uDB40\uDD00X"
selected UTF-16 offsets: [0]
{"codePointCount":4,"occurrences":[{"codePoint":"U+FEFF","utf16Offset":0,"codePointIndex":0,"line":1,"column":1},{"codePoint":"U+E0100","utf16Offset":3,"codePointIndex":2,"line":1,"column":3}],"cleanedAscii":"\"\\uD83D\\uDE00\\uDB40\\uDD00X\"","rawExportable":true}

U+FEFF is selected. The emoji and U+E0100 remain; the selector starts at UTF-16 offset 3 but code-point index 2, demonstrating why those coordinates differ.

Common inspection mistakes

  • Deleting every finding, including intentional spaces, newlines, joiners and variation selectors, just because it appears in the table.
  • Treating a UTF-16 offset as a byte offset or a screen column; emoji and combining sequences make those measurements differ.
  • Pasting the six ASCII characters \u200B and expecting a real zero-width space to be decoded automatically.
  • Assuming that a 50-row page or a shortened 4,000-unit preview means the report or occurrence counts are incomplete.
  • Sharing an escaped report as though it had removed secrets; ASCII escaping preserves content in a readable reversible form.
  • Expecting removal to normalize equivalent spellings or convert line endings; only explicitly selected occurrences change.

Limits and notes

  • Detection covers Unicode 17.0.0 general categories Cc, Cf, Zs, Zl and Zp, the Default_Ignorable_Code_Point property, and unpaired UTF-16 surrogates. Ordinary U+0020 SPACE, tabs and line breaks are therefore findings too. Overlapping properties create one occurrence, not duplicate rows. This is not an exhaustive test of visually invisible glyphs, confusable letters, malicious code or safe identifiers; a font or application can render the same text differently. Zero-width joiner (ZWJ), zero-width non-joiner (ZWNJ), variation selectors and bidi controls have legitimate uses in emoji, writing systems and mixed-direction text. Removing them can change spelling, shaping, emoji presentation or reading order. General combining marks are not all findings. The inspector does not implement a security verdict, grapheme segmentation or a preview of the Unicode Bidirectional Algorithm.
  • UTF-16 offsets and code-point indices are zero-based. Line and column are one-based, with columns counting code points rather than grapheme clusters, terminal cells or pixel width. A supplementary character occupies two UTF-16 units but one code point. An unpaired surrogate is counted once as a diagnostic pseudo-unit, not a Unicode scalar value. CRLF is one line break; CR and LF themselves appear on the old line at consecutive columns. LF, lone CR, VT, FF, NEL, LS and PS each start a new line afterward; a tab occupies one code-point column and empty input has one line.
  • No Unicode normalization, trimming, whitespace replacement or implicit line-ending conversion is performed by the inspector. Selected removals operate on original offsets and retain every unselected UTF-16 unit. The original remains available. Browser paste or clipboard handling can transform text outside this tool; strict file input and a reviewed download are preferable when exact line endings matter.
  • Processing is synchronous and local: at most 100,000 UTF-16 input units and a 400,000-byte UTF-8 file, with no uploads, network lookup or persistence. Invalid UTF-8 files are rejected rather than replaced with U+FFFD. The BOM, if present, is inspected as U+FEFF rather than silently stripped. Text already altered by another decoder cannot be reconstructed. The ASCII JSON report is limited to 40,000,000 characters; exceeding that bound stops export rather than returning a partial report.
  • Previews and diagnostic JSON escape non-ASCII and control characters; raw cleaned output deliberately contains the remaining original characters and can still affect another editor’s display. The complete report can contain the complete escaped source and cleaned text, so escaping is not redaction or encryption. Review private content before sharing. Raw UTF-8 export is blocked while unpaired surrogates remain to avoid silent replacement during encoding. Unicode security guidance has broader scope than character-property detection. UTR #36 is a stabilized historical report, not a current comprehensive checklist; Unicode points readers to newer material including UTS #39. This tool does not implement the full security mechanisms in those documents and cannot certify a string, file or repository as safe.

Frequently asked questions

Why is an ordinary space or newline listed?

The declared scope includes all Cc control characters and Zs/Zl/Zp separators, including U+0020, TAB, CR and LF. They are listed for a complete property-based inspection, not automatically marked for deletion. Leave intentional whitespace unselected.

Can a finding or a clean report establish whether text is safe?

No. A zero-width joiner can form an emoji sequence, a non-joiner can affect a word’s shaping, and a variation selector can request a specific presentation. Review the surrounding text and its intended use before selecting any removal. No. The report only covers the pinned Unicode properties and malformed UTF-16 described here. It does not detect all confusables, unusual combining sequences, font-dependent empty glyphs or harmful program behavior. Use language-aware tooling and human review for security decisions.

Can I remove one repeated character without deleting the others?

Yes. Selection is per occurrence at its original UTF-16 offset. For x\u200B\u200By, selecting offset 1 removes the first zero-width space and retains the one originally at offset 2. Output is rebuilt from the original so earlier removals do not shift later selections.

Why are raw copy and download sometimes unavailable?

An unpaired surrogate has no valid UTF-8 encoding. If one remains in the cleaned text, raw export is blocked rather than silently substituting U+FFFD. The ASCII-only diagnostic report remains useful; remove the specific malformed units only if that change is appropriate.

Related tools