String.prototype.codePointAt()
The Unicode-aware replacement for charCodeAt. It reads a surrogate pair as one character, and returns undefined rather than NaN for a bad index.
Demo
For ordinary characters this matches charCodeAt exactly. The emoji is the reason the method exists: at index 0 it reads both code units and returns 128512, the real code point, where charCodeAt would return 55357 — a meaningless surrogate half. The fourth case shows the catch: the index is still measured in code units, so index 1 lands on the LOW half of the pair and returns that half alone. Out of range gives undefined rather than NaN, which is the other improvement over charCodeAt.
Parameters
| Name | Type | Required | Description |
|---|---|---|---|
| index | number | no (0) | Position in UTF-16 code units. If a high surrogate sits there, the pair is read as one code point; landing on the LOW half instead gives just that half. |
Return value
number | undefined — The full Unicode code point starting at that index, up to 0x10FFFF. undefined when the index is out of range — not NaN.
Common patterns
for (const ch of s) console.log(ch.codePointAt(0));
const isAstral = s.codePointAt(0) > 0xffff;
const realLength = [...s].length;
Examples
Pitfalls
'\u{1F600}'.codePointAt(1)
[...'\u{1F600}'].map(c => c.codePointAt(0))
'\u{1F600}'.length
[...'\u{1F600}'].length
[...'\u{1F1EE}\u{1F1F1}'].length
[...new Intl.Segmenter().segment(s)].length
String.fromCharCode(128512).codePointAt(0)
String.fromCodePoint(128512).codePointAt(0)
When to use
- Any code-point work on text that may contain emoji or non-Latin scripts
- Detecting whether a character is outside the BMP
- Encoding and escaping routines that must be Unicode-correct
- You want the character itself → at, or for...of
- You want user-perceived characters → Intl.Segmenter
- Pure ASCII hashing where speed matters → charCodeAt is marginally simpler
- Comparing or sorting text → localeCompare
Notes
FAQ
Because JavaScript strings ARE sequences of UTF-16 code units — that is the underlying representation, and changing the indexing would break the language. codePointAt reads a pair when it finds one but cannot change how positions are counted. Iterating with for...of sidesteps the issue entirely.
for (const ch of s) { /* ch is a whole character */ }