String.prototype.normalize()

The answer to "these two strings look the same but are not equal". An accented letter can be stored as one code point or as a base letter plus a combining mark, and === cannot tell you which you have.

String methodES2015Live demo
Common call
s.normalize('NFC')
Returns
a new string in that normal form
Replaces
nothing; there is no other way to do this
Watch out
NFKC is LOSSY — it rewrites characters, not just recomposes them
string.normalize([form])
→ string

Demo

Live evaluation
Try:
Inputs
sstringtext with an accent
formstringNFC, NFD, NFKC, NFKD
Output
[...'é'.normalize('NFC')].map(c => c.codePointAt(0).toString(16))
['e9']

The output lists the code points in hex, because that is the only way to see what happened — every form of 'é' renders identically on screen. NFC gives a single code point e9. NFD gives two: 65, a plain letter e, followed by 301, a combining acute accent. Both display as é and they are NOT ===. The last two cases show the K forms going further: the fi ligature becomes two ordinary letters, and the circled digit ① becomes a plain 1. That is useful for search and destructive for display — NFKC discards the distinction permanently.

Parameters

NameTypeRequiredDescription
formstringno ('NFC')One of 'NFC', 'NFD', 'NFKC', 'NFKD'. The C forms compose, the D forms decompose; the K forms additionally fold compatibility characters, which loses information.

Return value

string — The string converted to the requested Unicode normalization form. Default is NFC, the composed form most systems expect.

Common patterns

Normalise before comparing or storing
NFC is what most systems and databases expect.
const key = input.normalize('NFC');
Strip accents for search
Decompose, then drop the combining marks.
const bare = s.normalize('NFD').replace(/\p{Diacritic}/gu, '');
Compare without normalising by hand
localeCompare handles it and more.
a.localeCompare(b, undefined, {sensitivity: 'base'}) === 0;

Examples

1. Composed length
'\u00e9'.length
Returns
1
2. Decomposed length
'e\u0301'.length
Returns
2
3. They are not equal
'\u00e9' === 'e\u0301'
Returns
false
4. Normalised, they are
'e\u0301'.normalize() === '\u00e9'
Returns
true
5. Ligature folded
'\ufb01'.normalize('NFKC')
Returns
'fi'
6. Circled digit
'\u2460'.normalize('NFKC')
Returns
'1'

Pitfalls

1. Two identical-looking strings are not equal
The bug this method exists to fix. Text typed on a Mac often arrives decomposed while the same text from elsewhere is composed, so a username or filename matches visually and fails every comparison, lookup and uniqueness check.
Looks equal, is not
'\u00e9' === 'e\u0301'
false
Normalise both
'\u00e9'.normalize() === 'e\u0301'.normalize()
true
2. NFKC and NFKD are lossy
They fold compatibility characters — superscripts, ligatures, circled digits, full-width forms — into plain equivalents. Excellent for building a search key, destructive if you store the result, because the original characters cannot be recovered.
Information gone
'x\u00b2'.normalize('NFKC')
'x2' // the superscript is lost
Store NFC
'x\u00b2'.normalize('NFC')
'x\u00b2'
3. It changes the length, so cached indices break
NFD can double the length of accented text and NFC can shrink it. Any offset computed before normalising points somewhere else afterwards — normalise first, then index.
Length changed
'\u00e9'.normalize('NFD').length
2
Normalise first
const t = s.normalize('NFC');
t.indexOf(x);
consistent
4. Normalising is not case folding or accent stripping
NFC and NFD keep every accent — they only change how it is encoded. Removing accents needs NFD followed by an explicit strip of the combining marks, and case still needs toLowerCase on top.
Accent kept
'\u00e9'.normalize('NFD') === 'e'
false
Strip marks
'\u00e9'.normalize('NFD').replace(/\p{Diacritic}/gu, '') === 'e'
true

When to use

Use it
  • Before comparing, hashing or storing text from users or files
  • Deduplicating names, tags or filenames
  • Building a search key, with the K forms
  • As the first step of accent stripping
Reach for something else
  • You only need a case-insensitive compare → toLowerCase
  • You want proper linguistic comparison → localeCompare with sensitivity
  • Storing display text → never store the K forms
  • Pure ASCII data → normalisation is a no-op, skip it

Notes

Complexity
O(n), with a table lookup per character
Return
A new string; the original is untouched
CPython impl
V8: Builtins-string-normalize / ICU
Memory
Allocates the result; the length often differs from the input
Thread-safe
Single-threaded; the string is only read

FAQ

NFC for anything you store or transmit — it is the web default, what most databases expect, and the shortest. NFD when you need to inspect or remove combining marks. The K forms only for building a comparison key you throw away afterwards.

store(input.normalize('NFC'));
searchKey = input.normalize('NFKC').toLowerCase();

History

ES2015
normalize added, exposing the Unicode normalization forms to JavaScript for the first time.