String.prototype.isWellFormed()

JavaScript strings can hold sequences that are not valid Unicode. These two methods let you find out before handing such a string to something that will throw.

String methodES2024Live demo
Common call
if (!s.isWellFormed()) s = s.toWellFormed();
Returns
boolean — and toWellFormed returns a repaired string
Replaces
a hand-rolled surrogate-range regex
Watch out
repair is LOSSY — the broken character becomes U+FFFD
string.isWellFormed()
→ boolean

Demo

Live evaluation
Try:
Inputs
codenumbera UTF-16 code unit
Output
String.fromCharCode(65).isWellFormed()
true

The demo builds a one-code-unit string, because a lone surrogate cannot survive being typed into a text box. Ordinary characters are well-formed. The two surrogate cases are not: 55357 and 56832 are the two halves of an emoji, and either one alone is an incomplete character that no valid Unicode encoding can represent. A string containing one is perfectly usable inside JavaScript — it has a length, you can slice it — and it will throw the moment you try to encode it as UTF-8.

Common patterns

Sanitise before encoding
The main use — avoid a throw from encodeURIComponent.
const safe = s.isWellFormed() ? s : s.toWellFormed();
Repair unconditionally
toWellFormed on a good string returns it unchanged.
const safe = s.toWellFormed();
Avoid creating them
Slice by code point, not code unit.
const head = [...s].slice(0, n).join('');

Examples

1. Ordinary text
'abc'.isWellFormed()
Returns
true
2. A whole emoji
'\u{1F600}'.isWellFormed()
Returns
true
3. Half of one
'\ud83d'.isWellFormed()
Returns
false
4. Repaired
'\ud83d'.toWellFormed()
Returns
'\ufffd'
5. The replacement char
'\ud83d'.toWellFormed().charCodeAt(0).toString(16)
Returns
'fffd'
6. Created by slicing
'\u{1F600}'.slice(0, 1).isWellFormed()
Returns
false

Pitfalls

1. Slicing by code unit creates them
This is where lone surrogates come from in practice. Truncating user text to a fixed length with slice, substring or a database column limit cuts an emoji in half, and the damaged string then breaks something downstream rather than at the point of the cut.
Broken half
'\u{1F600}ab'.slice(0, 1).isWellFormed()
false
Slice code points
[...'\u{1F600}ab'].slice(0, 1).join('').isWellFormed()
true
2. Repair is lossy and irreversible
toWellFormed replaces each lone surrogate with U+FFFD, the replacement character. The original code unit is gone, so repairing is a last resort at a boundary — better to avoid creating the damage at all.
Original lost
'\ud83d'.toWellFormed().charCodeAt(0).toString(16)
'fffd'
Do not split pairs
[...s].slice(0, n).join('')
nothing to repair
3. The failure surfaces far from its cause
A malformed string moves through your program without complaint until it reaches something that must encode it — encodeURIComponent, TextEncoder, JSON sent over the wire. The stack trace points at the encoder, not at the slice that caused it.
Throws later
encodeURIComponent('\ud83d')
URIError: URI malformed
Check at the boundary
encodeURIComponent('\ud83d'.toWellFormed())
'%EF%BF%BD'
4. ES2024 — check your runtime
Node 20+ and 2023-era browsers. Before that, detection meant a regex over the surrogate ranges, which is easy to get subtly wrong.
Missing
s.isWellFormed()
TypeError: s.isWellFormed is not a function
Regex fallback
!/[\ud800-\udbff](?![\udc00-\udfff])|(?:[^\ud800-\udbff]|^)[\udc00-\udfff]/.test(s)
equivalent, fiddly

When to use

Use it
  • Validating text before encoding it as UTF-8 or a URI
  • Sanitising strings that were truncated by length elsewhere
  • Guarding data that will be stored or transmitted
  • Diagnosing a URIError that appears to come from nowhere
Reach for something else
  • You control the slicing → iterate by code point and the problem cannot arise
  • You want to count characters → spread, or Intl.Segmenter
  • Checking for a specific character → includes or a regex
  • Targeting runtimes older than 2023 → the surrogate-range regex

Notes

Complexity
O(n) — a single scan for unpaired surrogates
Return
A boolean; toWellFormed returns a new string, or the original when already valid
CPython impl
V8: Builtins-string-iswellformed
Memory
isWellFormed allocates nothing
Thread-safe
Single-threaded; the string is only read

FAQ

Almost always by cutting between the halves of a surrogate pair — slice, substring, a fixed-width column, or a length-limited input. It can also come from decoding corrupted data, or from fromCharCode with a surrogate value.

'\u{1F600}'.slice(0, 1);   // half an emoji

History

ES2024
isWellFormed and toWellFormed added together to make surrogate validation a one-line check.