bytes.decode()

The boundary between bytes and text. Bytes carry no record of their own encoding, so decode is you asserting one — and the wrong assertion fails loudly or, worse, quietly.

Bytes methodPython 3.0+Live demo
Common call
data.decode('utf-8')
Returns
str — text, no longer bytes
Replaces
str(data, encoding), which does the same thing
Watch out
str(data) without an encoding gives "b'abc'", not the text
bytes.decode(encodingencoding — Codec name. 'utf-8' is the sane default; 'ascii' is strict; 'latin-1' accepts any byte and never raises.type: str · default: 'utf-8'='utf-8', errorserrors — How to handle undecodable bytes: 'strict' raises, 'replace' inserts a replacement character, 'ignore' drops them.type: str · default: 'strict'='strict')
→ str

Demo

Live evaluation
Try:
Inputs
sstrtext (encoded as utf-8 first)
encodingstrcodec to decode with
Output
bytes('abc', 'utf-8').decode('utf-8')
'abc'

The demo encodes your text as UTF-8, then decodes it with the codec you name — so the last three cases are the same bytes read three ways. Decoding as utf-8 gives the text back. Decoding as ascii RAISES, because the accented character needs a byte above 127. Decoding as latin-1 is the dangerous one: it succeeds and returns mojibake, because latin-1 maps every possible byte to some character and therefore can never fail.

Parameters

NameTypeRequiredDescription
encodingstrno ('utf-8')Codec name. 'utf-8' is the sane default; 'ascii' is strict; 'latin-1' accepts any byte and never raises.
errorsstrno ('strict')How to handle undecodable bytes: 'strict' raises, 'replace' inserts a replacement character, 'ignore' drops them.

Return value

str — The text the bytes represent under the given encoding. Raises UnicodeDecodeError when the bytes are not valid for that encoding and errors is strict.

Common patterns

Decode a network or file payload
The standard boundary crossing, with the encoding stated explicitly.
text = response.content.decode('utf-8')
Survive imperfect input
replace keeps going and marks the damage rather than crashing.
text = data.decode('utf-8', errors='replace')
Round-trip through bytes
encode and decode are exact inverses when the codec matches.
assert text.encode('utf-8').decode('utf-8') == text

Examples

1. ASCII
b'abc'.decode('utf-8')
Returns
'abc'
2. Default is utf-8
b'abc'.decode()
Returns
'abc'
3. Wrong codec raises
'é'.encode().decode('ascii')
Returns
UnicodeDecodeError: 'ascii' codec can't decode byte 0xc3 in position 0
4. latin-1 never fails
'é'.encode().decode('latin-1')
Returns
'é' # mojibake
5. errors=replace
b'\xff'.decode('utf-8', errors='replace')
Returns
'\ufffd'
6. Empty
b''.decode()
Returns
''

Pitfalls

1. latin-1 never raises, so it hides the bug
Every one of the 256 byte values maps to a character in latin-1, so decoding always "works". People reach for it to silence a UnicodeDecodeError and end up storing mojibake, which surfaces much later and much further away.
Silently wrong
'héllo'.encode('utf-8').decode('latin-1')
'héllo'
Match the codec
'héllo'.encode('utf-8').decode('utf-8')
'héllo'
2. str(data) does not decode
Calling str on bytes without an encoding gives you the REPR — the literal text b'abc', complete with the prefix and quotes. It is a common accident because it produces a plausible-looking string instead of an error.
The repr
str(b'abc')
"b'abc'"
Decode properly
b'abc'.decode('utf-8')
'abc'
3. Bytes carry no encoding information
There is nothing in a bytes object recording how it was produced. decode is an assertion you are making, not a detection — and if the source and your assertion disagree, nothing checks it for you.
Guessing
data.decode()   # hoping it is utf-8
right until it is not
Get it from the source
charset = resp.headers.get_content_charset('utf-8')
data.decode(charset)
stated, not assumed
4. errors=ignore deletes data
It silently drops undecodable bytes, so the result is shorter than the input with nothing to indicate what went missing. replace at least leaves a visible marker.
Data lost
b'a\xffb'.decode('utf-8', errors='ignore')
'ab' # the bad byte vanished
Mark the damage
b'a\xffb'.decode('utf-8', errors='replace')
'a\ufffdb'

When to use

Use it
  • Converting a payload from a file, socket or subprocess into text
  • Any boundary where bytes become something a person will read
  • Round-tripping text through a byte-oriented channel
Reach for something else
  • The data is genuinely binary — decoding an image is meaningless
  • You only need a debug view → repr shows the bytes safely
  • You do not know the encoding → find it out rather than guessing latin-1

Notes

Complexity
O(n) — every byte is examined
Return
A new str; the bytes object is unchanged
CPython impl
Objects/bytesobject.c :: bytes_decode, dispatching to the codec registry
Memory
Allocates a new string, which may be larger or smaller than the input
Thread-safe
Yes — bytes and str are immutable

FAQ

Mojibake is text decoded with the wrong codec — recognisable as sequences like é where an accented character should be. latin-1 causes it because it maps all 256 byte values to characters, so it can never report an error; it just produces the wrong characters confidently.

'é'.encode('utf-8').decode('latin-1')
# 'é'

History

3.0
bytes and str split cleanly, making decode the required boundary crossing.
3.1
errors='surrogateescape' added, allowing lossless round trips of undecodable bytes.