String.prototype.codePointAt()

The Unicode-aware replacement for charCodeAt. It reads a surrogate pair as one character, and returns undefined rather than NaN for a bad index.

String methodES2015Live demo
Common call
s.codePointAt(0)
Returns
a number up to 0x10FFFF, or undefined
Replaces
charCodeAt, for real text
Watch out
the INDEX is still in code units — astral characters take two
string.codePointAt(indexindex — Position in UTF-16 code units. If a high surrogate sits there, the pair is read as one code point; landing on the LOW half instead gives just that half.type: number · default: 0)
→ number | undefined

Demo

Live evaluation
Try:
Inputs
sstringthe source string
indexnumberindex
Output
'ABC'.codePointAt(0)
65

For ordinary characters this matches charCodeAt exactly. The emoji is the reason the method exists: at index 0 it reads both code units and returns 128512, the real code point, where charCodeAt would return 55357 — a meaningless surrogate half. The fourth case shows the catch: the index is still measured in code units, so index 1 lands on the LOW half of the pair and returns that half alone. Out of range gives undefined rather than NaN, which is the other improvement over charCodeAt.

Parameters

NameTypeRequiredDescription
indexnumberno (0)Position in UTF-16 code units. If a high surrogate sits there, the pair is read as one code point; landing on the LOW half instead gives just that half.

Return value

number | undefined — The full Unicode code point starting at that index, up to 0x10FFFF. undefined when the index is out of range — not NaN.

Common patterns

Iterate real characters
for...of iterates by code point, so no index arithmetic.
for (const ch of s) console.log(ch.codePointAt(0));
Detect an astral character
Anything above 0xFFFF takes two code units.
const isAstral = s.codePointAt(0) > 0xffff;
Count real characters
length counts code units; spread counts code points.
const realLength = [...s].length;

Examples

1. Capital A
'A'.codePointAt(0)
Returns
65
2. Whole emoji
'\u{1F600}'.codePointAt(0)
Returns
128512
3. charCodeAt halves it
'\u{1F600}'.charCodeAt(0)
Returns
55357
4. The low half
'\u{1F600}'.codePointAt(1)
Returns
56832
5. Out of range
'abc'.codePointAt(99)
Returns
undefined
6. charCodeAt gives NaN
'abc'.charCodeAt(99)
Returns
NaN

Pitfalls

1. The index is still in code units
This is the part people miss. codePointAt reads a whole character but is addressed by code unit, so walking a string with i++ lands on the low half of every astral character. Iterate with for...of or spread instead of indexing.
Hits the low half
'\u{1F600}'.codePointAt(1)
56832
Iterate by character
[...'\u{1F600}'].map(c => c.codePointAt(0))
[128512]
2. length is not the number of characters
It counts UTF-16 code units, so a single emoji has length 2 and a flag emoji can have length 4 or more. Any limit enforced against length — a tweet counter, a database column — will cut real text short.
Counts units
'\u{1F600}'.length
2
Counts characters
[...'\u{1F600}'].length
1
3. Even code points are not user-perceived characters
A family emoji or a flag is several code points joined by zero-width joiners, and an accented letter may be a base plus a combining mark. For "what the user sees as one character", Intl.Segmenter is the correct tool.
Three code points
[...'\u{1F1EE}\u{1F1F1}'].length
2 // one flag
Segment it
[...new Intl.Segmenter().segment(s)].length
1
4. Pairing it with String.fromCharCode
codePointAt returns values above 0xFFFF, and fromCharCode truncates to 16 bits — so the round trip silently destroys the character. Use String.fromCodePoint, its proper inverse.
Truncated
String.fromCharCode(128512).codePointAt(0)
62976 // the low 16 bits
Correct inverse
String.fromCodePoint(128512).codePointAt(0)
128512

When to use

Use it
  • Any code-point work on text that may contain emoji or non-Latin scripts
  • Detecting whether a character is outside the BMP
  • Encoding and escaping routines that must be Unicode-correct
Reach for something else
  • You want the character itself → at, or for...of
  • You want user-perceived characters → Intl.Segmenter
  • Pure ASCII hashing where speed matters → charCodeAt is marginally simpler
  • Comparing or sorting text → localeCompare

Notes

Complexity
O(1)
Return
A number or undefined; nothing is allocated
CPython impl
V8: Builtins-string-codepointat
Memory
No allocation
Thread-safe
Single-threaded; the string is only read

FAQ

Because JavaScript strings ARE sequences of UTF-16 code units — that is the underlying representation, and changing the indexing would break the language. codePointAt reads a pair when it finds one but cannot change how positions are counted. Iterating with for...of sidesteps the issue entirely.

for (const ch of s) { /* ch is a whole character */ }

History

ES2015
codePointAt, String.fromCodePoint and code-point iteration added together to make the language Unicode-correct.