bytes.decode()
The boundary between bytes and text. Bytes carry no record of their own encoding, so decode is you asserting one — and the wrong assertion fails loudly or, worse, quietly.
Demo
The demo encodes your text as UTF-8, then decodes it with the codec you name — so the last three cases are the same bytes read three ways. Decoding as utf-8 gives the text back. Decoding as ascii RAISES, because the accented character needs a byte above 127. Decoding as latin-1 is the dangerous one: it succeeds and returns mojibake, because latin-1 maps every possible byte to some character and therefore can never fail.
Parameters
| Name | Type | Required | Description |
|---|---|---|---|
| encoding | str | no ('utf-8') | Codec name. 'utf-8' is the sane default; 'ascii' is strict; 'latin-1' accepts any byte and never raises. |
| errors | str | no ('strict') | How to handle undecodable bytes: 'strict' raises, 'replace' inserts a replacement character, 'ignore' drops them. |
Return value
str — The text the bytes represent under the given encoding. Raises UnicodeDecodeError when the bytes are not valid for that encoding and errors is strict.
Common patterns
text = response.content.decode('utf-8')
text = data.decode('utf-8', errors='replace')
assert text.encode('utf-8').decode('utf-8') == text
Examples
Pitfalls
'héllo'.encode('utf-8').decode('latin-1')
'héllo'.encode('utf-8').decode('utf-8')
str(b'abc')
b'abc'.decode('utf-8')
data.decode() # hoping it is utf-8
charset = resp.headers.get_content_charset('utf-8') data.decode(charset)
b'a\xffb'.decode('utf-8', errors='ignore')
b'a\xffb'.decode('utf-8', errors='replace')
When to use
- Converting a payload from a file, socket or subprocess into text
- Any boundary where bytes become something a person will read
- Round-tripping text through a byte-oriented channel
- The data is genuinely binary — decoding an image is meaningless
- You only need a debug view → repr shows the bytes safely
- You do not know the encoding → find it out rather than guessing latin-1
Notes
FAQ
Mojibake is text decoded with the wrong codec — recognisable as sequences like é where an accented character should be. latin-1 causes it because it maps all 256 byte values to characters, so it can never report an error; it just produces the wrong characters confidently.
'é'.encode('utf-8').decode('latin-1') # 'é'