UnicodeDecodeError
The bytes are not valid in the codec you picked — almost always the data was written in a different encoding than you assumed. Positions are byte offsets, not character indexes.
Demo
'hello'.encode('latin-1').decode('utf-8')
In Trigger, é is the single byte 0xe9 in Latin-1. UTF-8 reads 0xe9 as the start of a 3-byte sequence: at the end of the data that is unexpected end of data, followed by a space it is an invalid continuation byte. '5 €' fails one step earlier — € does not exist in Latin-1, so encode() raises UnicodeEncodeError. In Handle, the cut-off emoji f0 9f 98 is one bad sequence: one U+FFFD, three escaped bytes.
Constructor
| Name | Type | Required | Description |
|---|---|---|---|
| encoding | str | yes | Codec name as shown in the message, e.g. 'utf-8'. Stored in e.encoding. |
| object | bytes | yes | The whole input being decoded (any bytes-like object is stored as bytes). Stored in e.object. |
| start | int | yes | Index of the first bad byte in object. |
| end | int | yes | Index just after the last bad byte. end - start == 1 gives the "byte 0x.. in position N" wording; a longer range gives "bytes in position S-E". |
| reason | str | yes | Why decoding failed, e.g. 'invalid start byte'. |
Attributes
| Attribute | Type | Meaning |
|---|---|---|
| encoding | str | The codec that failed, e.g. 'utf-8' — or 'charmap' for table-based codecs such as cp1252. |
| object | bytes | The complete input that was being decoded. |
| start | int | Byte offset of the first invalid byte. object[start:end] is the offending slice. |
| end | int | Byte offset just past the invalid sequence. |
| reason | str | Short explanation: 'invalid start byte', 'invalid continuation byte', 'unexpected end of data', 'ordinal not in range(128)', … |
| args | tuple | All five constructor arguments, in order. |
Common patterns
with open(path, encoding='utf-8') as f: text = f.read()
def read_text(data: bytes) -> str: try: return data.decode('utf-8') except UnicodeDecodeError: return data.decode('cp1252', errors='replace')
try: text = data.decode('utf-8') except UnicodeDecodeError as e: context = e.object[max(0, e.start - 10):e.end + 10] raise ValueError(f'bad UTF-8 at byte {e.start}: {context!r}') from e
text = data.decode('utf-8', errors='surrogateescape') assert text.encode('utf-8', errors='surrogateescape') == data
Examples
Pitfalls
b'caf\xe9'.decode('utf-8', errors='ignore')
b'caf\xe9'.decode('cp1252')
data = 'é'.encode('utf-8') data[:1].decode('utf-8') + data[1:].decode('utf-8')
import codecs data = 'é'.encode('utf-8') dec = codecs.getincrementaldecoder('utf-8')() dec.decode(data[:1]) + dec.decode(data[1:])
str(b'caf\xc3\xa9')
b'caf\xc3\xa9'.decode('utf-8')
When to use
- Catch it where outside bytes (files, sockets, subprocess output) become text
- Use its start/end/object attributes to report exactly which bytes are bad
- Raise it from a custom codec (with all five arguments)
- Hiding it with errors='ignore' when the real problem is the wrong codec
- Catching it just to retry with random encodings — latin-1 never fails, so it 'succeeds' with garbage
- Decoding at all when the data is binary (images, zip, pickle) — keep it as bytes
Notes
FAQ
Byte number N (counting from 0) of the input is not valid UTF-8 at that point. The reason tells you more: invalid start byte means a byte that can never begin a character (0x80-0xc1, 0xf5-0xff — 0xff and 0xfe usually mean UTF-16 data); invalid continuation byte means a multi-byte sequence was cut short by an ordinary byte, the usual sign of Latin-1/cp1252 text; unexpected end of data means the input stops in the middle of a character.