String.prototype.isWellFormed()
JavaScript strings can hold sequences that are not valid Unicode. These two methods let you find out before handing such a string to something that will throw.
Demo
The demo builds a one-code-unit string, because a lone surrogate cannot survive being typed into a text box. Ordinary characters are well-formed. The two surrogate cases are not: 55357 and 56832 are the two halves of an emoji, and either one alone is an incomplete character that no valid Unicode encoding can represent. A string containing one is perfectly usable inside JavaScript — it has a length, you can slice it — and it will throw the moment you try to encode it as UTF-8.
Common patterns
const safe = s.isWellFormed() ? s : s.toWellFormed();
const safe = s.toWellFormed();
const head = [...s].slice(0, n).join('');
Examples
Pitfalls
'\u{1F600}ab'.slice(0, 1).isWellFormed()
[...'\u{1F600}ab'].slice(0, 1).join('').isWellFormed()
'\ud83d'.toWellFormed().charCodeAt(0).toString(16)
[...s].slice(0, n).join('')
encodeURIComponent('\ud83d')
encodeURIComponent('\ud83d'.toWellFormed())
s.isWellFormed()
!/[\ud800-\udbff](?![\udc00-\udfff])|(?:[^\ud800-\udbff]|^)[\udc00-\udfff]/.test(s)
When to use
- Validating text before encoding it as UTF-8 or a URI
- Sanitising strings that were truncated by length elsewhere
- Guarding data that will be stored or transmitted
- Diagnosing a URIError that appears to come from nowhere
- You control the slicing → iterate by code point and the problem cannot arise
- You want to count characters → spread, or Intl.Segmenter
- Checking for a specific character → includes or a regex
- Targeting runtimes older than 2023 → the surrogate-range regex
Notes
FAQ
Almost always by cutting between the halves of a surrogate pair — slice, substring, a fixed-width column, or a length-limited input. It can also come from decoding corrupted data, or from fromCharCode with a surrogate value.
'\u{1F600}'.slice(0, 1); // half an emoji