UnicodeDecodeError

The bytes are not valid in the codec you picked — almost always the data was written in a different encoding than you assumed. Positions are byte offsets, not character indexes.

InheritsBaseException›Exception›ValueError›UnicodeError›UnicodeDecodeError
Type / value exceptionPython 3 (all)Live demo
UnicodeDecodeError(encodingencoding — Codec name as shown in the message, e.g. 'utf-8'. Stored in e.encoding.type: str · required, objectobject — The whole input being decoded (any bytes-like object is stored as bytes). Stored in e.object.type: bytes · required, startstart — Index of the first bad byte in object.type: int · required, endend — Index just after the last bad byte. end - start == 1 gives the "byte 0x.. in position N" wording; a longer range gives "bytes in position S-E".type: int · required, reasonreason — Why decoding failed, e.g. 'invalid start byte'.type: str · required)
Raised by
b.decode('utf-8'), str(b, 'ascii'), open(p, encoding='utf-8').read()
Message
'utf-8' codec can't decode byte 0xe9 in position 3: invalid continuation byte
Quick fix
decode with the real encoding, or errors='replace'
Watch out
position is a byte offset into e.object, not an index into the text

Demo

Live evaluation
The classic wrong-codec bug: text saved as Latin-1 (one byte per accented letter) and read back as UTF-8.
Try:
Inputs
textstrtext to round-trip
Code
'hello'.encode('latin-1').decode('utf-8')
Result
'hello'

In Trigger, é is the single byte 0xe9 in Latin-1. UTF-8 reads 0xe9 as the start of a 3-byte sequence: at the end of the data that is unexpected end of data, followed by a space it is an invalid continuation byte. '5 €' fails one step earlier — € does not exist in Latin-1, so encode() raises UnicodeEncodeError. In Handle, the cut-off emoji f0 9f 98 is one bad sequence: one U+FFFD, three escaped bytes.

Constructor

NameTypeRequiredDescription
encodingstryesCodec name as shown in the message, e.g. 'utf-8'. Stored in e.encoding.
objectbytesyesThe whole input being decoded (any bytes-like object is stored as bytes). Stored in e.object.
startintyesIndex of the first bad byte in object.
endintyesIndex just after the last bad byte. end - start == 1 gives the "byte 0x.. in position N" wording; a longer range gives "bytes in position S-E".
reasonstryesWhy decoding failed, e.g. 'invalid start byte'.

Attributes

AttributeTypeMeaning
encodingstrThe codec that failed, e.g. 'utf-8' — or 'charmap' for table-based codecs such as cp1252.
objectbytesThe complete input that was being decoded.
startintByte offset of the first invalid byte. object[start:end] is the offending slice.
endintByte offset just past the invalid sequence.
reasonstrShort explanation: 'invalid start byte', 'invalid continuation byte', 'unexpected end of data', 'ordinal not in range(128)', …
argstupleAll five constructor arguments, in order.

Common patterns

Always pass encoding= to open()
Without it, open() uses the platform's locale encoding (often cp1252 on Windows, UTF-8 elsewhere), so the same file can decode on one machine and fail on another.
with open(path, encoding='utf-8') as f:
    text = f.read()
Try UTF-8, fall back to a legacy codec
Valid UTF-8 rarely happens by accident, so try it first. cp1252 accepts almost any byte, which makes it a reasonable fallback for Western European legacy files.
def read_text(data: bytes) -> str:
    try:
        return data.decode('utf-8')
    except UnicodeDecodeError:
        return data.decode('cp1252', errors='replace')
Show where decoding failed
The attributes pinpoint the bad bytes — print a little context around them.
try:
    text = data.decode('utf-8')
except UnicodeDecodeError as e:
    context = e.object[max(0, e.start - 10):e.end + 10]
    raise ValueError(f'bad UTF-8 at byte {e.start}: {context!r}') from e
Round-trip unknown bytes with surrogateescape
errors=surrogateescape decodes every invalid byte to a lone surrogate and encodes it back to the same byte — lossless for file names and pass-through data.
text = data.decode('utf-8', errors='surrogateescape')
assert text.encode('utf-8', errors='surrogateescape') == data

Examples

1. Latin-1 byte read as UTF-8
b'caf\xe9'.decode('utf-8')
Returns
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 in position 3: unexpected end of data
2. Decode with the right codec
b'caf\xe9'.decode('latin-1')
Returns
'café'
3. Inspect the attributes
try: b'ab\xffcd'.decode('utf-8') except UnicodeDecodeError as e: info = (e.encoding, e.start, e.end, e.reason, e.object[e.start:e.end]) info
Returns
('utf-8', 2, 3, 'invalid start byte', b'\xff')
4. Positions count bytes
data = 'día'.encode('utf-8') + b'\xff' try: data.decode('utf-8') except UnicodeDecodeError as e: pos = (len('día'), e.start) pos
Returns
(3, 4)
5. Truncated multi-byte char
'😀'.encode('utf-8')[:3].decode('utf-8')
Returns
UnicodeDecodeError: 'utf-8' codec can't decode bytes in position 0-2: unexpected end of data
6. UTF-16 data read as UTF-8
'\ufeffhi'.encode('utf-16-le').decode('utf-8')
Returns
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte
7. Reading a file
with open('menu.txt', 'wb') as f: f.write(b'caf\xe9 au lait') with open('menu.txt', encoding='utf-8') as f: f.read()
Returns
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 in position 3: invalid continuation byte
8. cp1252 has holes: 'charmap'
b'\x9d'.decode('cp1252')
Returns
UnicodeDecodeError: 'charmap' codec can't decode byte 0x9d in position 0: character maps to <undefined> decoding with 'cp1252' codec failed

Pitfalls

1. errors='ignore' silently loses letters
Suppressing the error does not fix the encoding mismatch — it deletes the characters that proved it. Find the real codec instead.
errors='ignore'
b'caf\xe9'.decode('utf-8', errors='ignore')
'caf'
The real codec
b'caf\xe9'.decode('cp1252')
'café'
2. Decoding a stream chunk by chunk
A chunk boundary can split a multi-byte character; decoding each chunk separately then fails even though the data is valid. An incremental decoder keeps the partial bytes for the next call.
chunk.decode()
data = 'é'.encode('utf-8')
data[:1].decode('utf-8') + data[1:].decode('utf-8')
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xc3 in position 0: unexpected end of data
Incremental decoder
import codecs
data = 'é'.encode('utf-8')
dec = codecs.getincrementaldecoder('utf-8')()
dec.decode(data[:1]) + dec.decode(data[1:])
'é'
3. str(bytes) does not decode
Trying to dodge the error with str() gives the repr of the bytes object — b'…' ends up inside your text.
str(data)
str(b'caf\xc3\xa9')
"b'caf\\xc3\\xa9'"
data.decode()
b'caf\xc3\xa9'.decode('utf-8')
'café'

When to use

Use it
  • Catch it where outside bytes (files, sockets, subprocess output) become text
  • Use its start/end/object attributes to report exactly which bytes are bad
  • Raise it from a custom codec (with all five arguments)
Reach for something else
  • Hiding it with errors='ignore' when the real problem is the wrong codec
  • Catching it just to retry with random encodings — latin-1 never fails, so it 'succeeds' with garbage
  • Decoding at all when the data is binary (images, zip, pickle) — keep it as bytes

Notes

CPython impl
Objects/unicodeobject.c + Objects/stringlib/codecs.h — the UTF-8 decoder reports the longest valid prefix of a broken sequence as one error (start..end)
Catch via
except UnicodeError (all three unicode errors) or except ValueError
charmap
Table-based codecs (cp1252, cp437, …) report their encoding as 'charmap' and the reason as 'character maps to <undefined>'; when called via bytes.decode() the traceback adds a note naming the real codec (decoding with 'cp1252' codec failed)
BOM
Data starting with ff fe or fe ff is UTF-16; ef bb bf is a UTF-8 BOM — decode with 'utf-8-sig' to drop it

FAQ

Byte number N (counting from 0) of the input is not valid UTF-8 at that point. The reason tells you more: invalid start byte means a byte that can never begin a character (0x80-0xc1, 0xf5-0xff — 0xff and 0xfe usually mean UTF-16 data); invalid continuation byte means a multi-byte sequence was cut short by an ordinary byte, the usual sign of Latin-1/cp1252 text; unexpected end of data means the input stops in the middle of a character.