Back

Unicode normalization

Text tools

Loading

Loading tool

The tool is loaded only when you open it.

All processing for this tool happens in your browser. Your input is not sent to a server.

About this tool

Visually similar text can use different Unicode sequences. Compare the original with NFC, NFD, NFKC and NFKD in one local inspection, keeping the original intact. Review exact escaped text, code points, UTF-16 units and UTF-8 byte lengths before choosing an output. Canonical normalization composes or decomposes canonically equivalent sequences; compatibility normalization also folds distinctions such as some ligatures, fullwidth characters and circled numbers, which can discard meaning or formatting that matters to your application.

Common uses

  • Diagnose why an accented name copied from two sources looks the same but fails an exact string comparison.
  • Compare composed Hangul syllables with decomposed conjoining jamo and see how code-point and UTF-8 lengths change.
  • Review the broader effects of NFKC or NFKD on compatibility characters before choosing a normalization policy for an import or search workflow.

How to use it

  1. 1.Paste literal text or open one strictly decoded UTF-8 text file, then run the comparison. The original and all four forms are computed from the same complete input. A spelling such as \u0301 is ordinary ASCII unless you supply the actual combining character; the examples below use ASCII JSON solely to display exact test data.
  2. 2.Compare changed/unchanged status, escaped previews, code-point previews and UTF-16/UTF-8 counts. Each text preview covers at most 2,000 source UTF-16 units and 12,000 rendered ASCII characters; each U+ sequence shows at most 80 code points. Truncation is labeled and never shortens the stored result or complete report. Select the form appropriate to the receiving application and review its meaning.
  3. 3.After explicitly reviewing the selected result, copy or download the complete text. Download the ASCII-only JSON report to retain the original and all four forms. Editing the input invalidates the previous comparison and review; malformed UTF-16 blocks raw text export but remains losslessly available in escaped JSON.

Exact normalization examples

Compose an accented sequence

input (ASCII JSON): "e\u0301"
{"original":{"text":"e\u0301","codePoints":2,"utf16Units":2,"utf8Bytes":3},"NFC":{"text":"\u00E9","codePoints":1,"utf16Units":1,"utf8Bytes":2},"NFD":{"text":"e\u0301","codePoints":2,"utf16Units":2,"utf8Bytes":3},"NFKC":{"text":"\u00E9","codePoints":1,"utf16Units":1,"utf8Bytes":2},"NFKD":{"text":"e\u0301","codePoints":2,"utf16Units":2,"utf8Bytes":3}}

NFC and NFKC turn e plus COMBINING ACUTE ACCENT into U+00E9. NFD and NFKD retain the two-code-point canonical decomposition.

Decompose a precomposed accent

input (ASCII JSON): "\u00E9"
{"original":{"text":"\u00E9","codePoints":1,"utf16Units":1,"utf8Bytes":2},"NFC":{"text":"\u00E9","codePoints":1,"utf16Units":1,"utf8Bytes":2},"NFD":{"text":"e\u0301","codePoints":2,"utf16Units":2,"utf8Bytes":3},"NFKC":{"text":"\u00E9","codePoints":1,"utf16Units":1,"utf8Bytes":2},"NFKD":{"text":"e\u0301","codePoints":2,"utf16Units":2,"utf8Bytes":3}}

NFD and NFKD expand U+00E9 into e plus the combining acute accent. The visible spelling can remain similar while the byte length changes.

Decompose a Hangul syllable

input (ASCII JSON): "\uAC01"
{"original":{"text":"\uAC01","codePoints":1,"utf16Units":1,"utf8Bytes":3},"NFC":{"text":"\uAC01","codePoints":1,"utf16Units":1,"utf8Bytes":3},"NFD":{"text":"\u1100\u1161\u11A8","codePoints":3,"utf16Units":3,"utf8Bytes":9},"NFKC":{"text":"\uAC01","codePoints":1,"utf16Units":1,"utf8Bytes":3},"NFKD":{"text":"\u1100\u1161\u11A8","codePoints":3,"utf16Units":3,"utf8Bytes":9}}

The syllable U+AC01 decomposes into three conjoining jamo U+1100 U+1161 U+11A8 in NFD and NFKD, occupying nine UTF-8 bytes.

Review ligature, width and circled-number loss

input (ASCII JSON): "\uFB03\uFF21\u2460"
{"original":{"text":"\uFB03\uFF21\u2460","codePoints":3,"utf16Units":3,"utf8Bytes":9},"NFC":{"text":"\uFB03\uFF21\u2460","codePoints":3,"utf16Units":3,"utf8Bytes":9},"NFD":{"text":"\uFB03\uFF21\u2460","codePoints":3,"utf16Units":3,"utf8Bytes":9},"NFKC":{"text":"ffiA1","codePoints":5,"utf16Units":5,"utf8Bytes":5},"NFKD":{"text":"ffiA1","codePoints":5,"utf16Units":5,"utf8Bytes":5}}

Only the compatibility forms turn U+FB03 U+FF21 U+2460 into ffiA1. The original distinctions are not recoverable from that result alone.

Reorder combining marks without a length change

input (ASCII JSON): "q\u0307\u0323"
{"original":{"text":"q\u0307\u0323","codePoints":3,"utf16Units":3,"utf8Bytes":5},"NFC":{"text":"q\u0323\u0307","codePoints":3,"utf16Units":3,"utf8Bytes":5},"NFD":{"text":"q\u0323\u0307","codePoints":3,"utf16Units":3,"utf8Bytes":5},"NFKC":{"text":"q\u0323\u0307","codePoints":3,"utf16Units":3,"utf8Bytes":5},"NFKD":{"text":"q\u0323\u0307","codePoints":3,"utf16Units":3,"utf8Bytes":5}}

All forms put COMBINING DOT BELOW before COMBINING DOT ABOVE. The sequence changes while all three length measurements stay the same.

Retain CRLF and emoji sequences

input (ASCII JSON): "A\r\n\uD83D\uDC69\u200D\uD83D\uDCBB\u2764\uFE0F"
{"original":{"text":"A\r\n\uD83D\uDC69\u200D\uD83D\uDCBB\u2764\uFE0F","codePoints":8,"utf16Units":10,"utf8Bytes":20},"NFC":{"text":"A\r\n\uD83D\uDC69\u200D\uD83D\uDCBB\u2764\uFE0F","codePoints":8,"utf16Units":10,"utf8Bytes":20},"NFD":{"text":"A\r\n\uD83D\uDC69\u200D\uD83D\uDCBB\u2764\uFE0F","codePoints":8,"utf16Units":10,"utf8Bytes":20},"NFKC":{"text":"A\r\n\uD83D\uDC69\u200D\uD83D\uDCBB\u2764\uFE0F","codePoints":8,"utf16Units":10,"utf8Bytes":20},"NFKD":{"text":"A\r\n\uD83D\uDC69\u200D\uD83D\uDCBB\u2764\uFE0F","codePoints":8,"utf16Units":10,"utf8Bytes":20}}

All forms preserve the CRLF pair, woman-technologist ZWJ sequence and heart variation selector. Ten UTF-16 units represent eight code points and twenty UTF-8 bytes.

Preserve an unpaired surrogate in JSON

input (ASCII JSON): "A\uD800B"
{"original":{"text":"A\uD800B","codePoints":3,"utf16Units":3,"utf8Bytes":null},"NFC":{"text":"A\uD800B","codePoints":3,"utf16Units":3,"utf8Bytes":null},"NFD":{"text":"A\uD800B","codePoints":3,"utf16Units":3,"utf8Bytes":null},"NFKC":{"text":"A\uD800B","codePoints":3,"utf16Units":3,"utf8Bytes":null},"NFKD":{"text":"A\uD800B","codePoints":3,"utf16Units":3,"utf8Bytes":null}}

All forms retain the unpaired U+D800 unit. UTF-8 length is null and raw copy/download stay blocked; ASCII JSON preserves the exact string.

Keep a literal escape spelling literal

input (ASCII JSON): "e\\u0301"
{"original":{"text":"e\\u0301","codePoints":7,"utf16Units":7,"utf8Bytes":7},"NFC":{"text":"e\\u0301","codePoints":7,"utf16Units":7,"utf8Bytes":7},"NFD":{"text":"e\\u0301","codePoints":7,"utf16Units":7,"utf8Bytes":7},"NFKC":{"text":"e\\u0301","codePoints":7,"utf16Units":7,"utf8Bytes":7},"NFKD":{"text":"e\\u0301","codePoints":7,"utf16Units":7,"utf8Bytes":7}}

The input contains seven ASCII characters, including a backslash. The tool does not evaluate the spelling as a combining accent.

Inspect empty input

input (ASCII JSON): ""
{"original":{"text":"","codePoints":0,"utf16Units":0,"utf8Bytes":0},"NFC":{"text":"","codePoints":0,"utf16Units":0,"utf8Bytes":0},"NFD":{"text":"","codePoints":0,"utf16Units":0,"utf8Bytes":0},"NFKC":{"text":"","codePoints":0,"utf16Units":0,"utf8Bytes":0},"NFKD":{"text":"","codePoints":0,"utf16Units":0,"utf8Bytes":0}}

All four forms of the empty string remain empty, with zero code points, UTF-16 units and UTF-8 bytes.

Common normalization mistakes

  • Choosing NFKC solely because it changes more characters, without checking lost typography or distinctions.
  • Assuming a composed character, a combining sequence and an emoji cluster each have the same UTF-16 or UTF-8 length.
  • Treating unchanged counts as unchanged text; canonical reordering can preserve every length while changing the sequence.
  • Expecting literal ASCII \uXXXX spellings to be decoded automatically.
  • Using normalization as a security sanitizer, translation step, case-folding operation or proof of identity.
  • Sharing a complete escaped report as though escaping had removed secrets.

Limits and notes

  • NFC uses canonical decomposition followed by canonical composition; NFD uses canonical decomposition. NFKC and NFKD additionally apply compatibility decomposition, with NFKC then composing. Compatibility changes are not generally reversible: ligatures, width variants, circled digits and other distinctions can collapse. Normalization does not strip accents, translate, transliterate, convert letter case or establish that two different strings have the same intended meaning.
  • The implementation uses this browser runtime’s String.prototype.normalize and its Unicode support. It does not ship a Unicode 17 normalization database or promise a pinned Unicode version. A newer character can behave differently when a runtime does not yet know its mapping. Use an explicitly versioned downstream implementation when reproducibility for newly assigned characters is required.
  • Code-point counts differ from visible characters, grapheme clusters and display width; one emoji sequence can contain several code points. Supplementary characters occupy two UTF-16 units. An unpaired surrogate is retained and counted once as a diagnostic pseudo-unit, not a Unicode scalar value. Its UTF-8 length is unavailable, and raw copy/download are blocked instead of silently substituting U+FFFD. Change details show a single encompassing span after a shared code-point prefix and suffix, not a minimal edit script; unchanged sections inside that span can be included. Span indices are zero-based with exclusive ends and never split a valid surrogate pair.
  • Normalization is not sanitization, confusable or homoglyph detection, malware analysis, identifier validation or a security verdict. It does not remove arbitrary invisible characters, bidi controls, emoji joiners or variation selectors. CR and LF remain as supplied to the core; browser paste or clipboard handling can change line endings outside the tool. Strict file import and downloaded output are preferable for exact line endings.
  • Processing is local and bounded: at most 20,000 UTF-16 input units, an 80,000-byte UTF-8 file, 360,000 UTF-16 units per normalized form and 1,460,000 total units across the original and four forms. The ASCII JSON report is limited to 10,000,000 characters. There are no input uploads, network lookups or persistence. Invalid UTF-8 files are rejected rather than replacement-decoded, and a leading BOM is preserved as U+FEFF. Exceeding a processing or report bound causes an explicit error, never a silently truncated export. ASCII escaping is reversible representation, not redaction or encryption: the report contains the complete original and normalized text, including any private information.

Frequently asked questions

Which normalization form should I choose?

Follow the contract of the system receiving the text. NFC commonly gives a composed canonical representation, while NFD gives a decomposed one. NFKC/NFKD also collapse compatibility distinctions and need a separate semantic review. The tool compares all forms rather than declaring one universally correct. No. Composition is constrained by Unicode normalization rules, and some characters are excluded from composition. Combining marks may reorder even when code-point and byte counts are unchanged. Compare the exact output rather than using length as proof that nothing changed.

Will NFKC make suspicious text safe or merge all look-alike letters?

No. Compatibility normalization is not a full security mechanism or a confusable detector. For example, Latin and Cyrillic letters can look similar while remaining different. Review application-specific identifier rules and appropriate security tooling separately.

Why can the JSON report be exported when text download is blocked?

An unpaired UTF-16 surrogate cannot be encoded as valid UTF-8 without a replacement. The diagnostic JSON uses ASCII escapes to preserve that exact code unit in the original and all forms. Escaping does not repair it, and it does not make the text safe.

Are spaces, emoji and line endings automatically cleaned?

There is no extra cleanup. Compatibility normalization can change a character with a compatibility mapping, such as a non-breaking space, but canonical normalization does not perform a general whitespace replacement. CRLF, emoji joiners and variation selectors are not stripped. Inspect the actual code points before exporting.

Related tools