Text interoperability

Unicode normalization, hidden characters and text encoding

Decoding, normalization and character removal solve different problems. Keep the original bytes or text and make each transformation explicit.

Content updated: · Maintainer and corrections

Choose a normalization contract

NFC and NFD handle canonical equivalence with composition or decomposition. NFKC and NFKD also fold compatibility distinctions: ① can become 1. That may change meaning in identifiers, mathematical text or typography; do not apply compatibility normalization blindly.

Code points are not user-perceived characters. A combining sequence or emoji can contain several code points; a positional comparison is not an edit-distance alignment. The tools retain the original. Native String.normalize uses your browser's Unicode implementation, not a bundled current-version database.

Inspect before removing

Zero-width joiners, non-joiners, variation selectors and bidirectional controls can affect valid language or emoji rendering. Keep them unless the receiving contract explicitly forbids them. The inspector covers a documented character set, not every invisible glyph or spoofing technique.

UTF-16 offsets and code-point positions differ after supplementary characters. Use the displayed index convention when finding a character in an editor. Review each selected character type and the complete cleaned result; removing a control is not sanitization or proof that content is trustworthy.

Decoding needs a known encoding

Bytes are not text until decoded. Select the source encoding explicitly; a BOM can identify some Unicode byte orders but is not a universal encoding detector. TextDecoder's legacy mappings follow the browser Encoding Standard, which may differ from a vendor's original encoding.

Strict decoding reports invalid sequences, but a wrong legacy encoding can still decode successfully into incorrect text. Inspect known words and source documentation. UTF-8 export creates new bytes; it cannot preserve original encoding, metadata or malformed sequences. Browser support is checked before use.

e + U+0301 → NFC: é
① → NFKC: 1
👩 + U+200D + 💻 → keep the joiner to retain the intended emoji
Big5 A4 40 → 一; UTF-8 export: E4 B8 80

Data provenance

Sources