bytes.lower()

The usual first step in case-insensitive comparison of ASCII data. Like upper, it maps exactly 26 byte values and ignores everything else.

Bytes methodPython 3.0+Live demo
Common call
data.lower()
Returns
a new bytes object; the original is unchanged
Replaces
a manual loop over bytes
Watch out
accented and non-Latin letters are NOT lowercased — decode first
bytes.lower()
→ bytes

Demo

Live evaluation
Try:
Inputs
sstrdata (encoded as utf-8)
Output
bytes('HELLO', 'utf-8').lower()
b'hello'

Only the 26 ASCII uppercase bytes change. Digits, punctuation and spaces pass through, and so does anything outside ASCII. In the accented case the H, L and O become lowercase but the É does not, because its two UTF-8 bytes are not in the A-Z range — the bytes method has no idea it is a letter at all.

Common patterns

Case-insensitive header lookup
HTTP header names are ASCII and case-insensitive by specification.
headers = {k.lower(): v for k, v in raw_headers}
Normalise a protocol keyword
Accept GET, get and Get alike.
if verb.lower() == b'get':
    ...
Lowercase real text
Decode first so non-ASCII letters are handled.
data.decode('utf-8').lower().encode('utf-8')

Examples

1. Uppercase
b'HELLO'.lower()
Returns
b'hello'
2. Mixed
b'HeLLo World'.lower()
Returns
b'hello world'
3. Digits stay
b'ABC123'.lower()
Returns
b'abc123'
4. Accent ignored
'HÉLLO'.encode().lower()
Returns
b'h\xc3\x89llo'
5. Empty
b''.lower()
Returns
b''
6. Compare
b'GET'.lower() == b'get'
Returns
True

Pitfalls

1. Non-ASCII letters are not touched
Only the 26 ASCII letters map. An accented capital stays a capital, producing a half-lowercased word.
Half converted
'HÉLLO'.encode().lower()
b'h\xc3\x89llo'
Work on text
'HÉLLO'.lower().encode()
b'h\xc3\xa9llo'
2. It returns a new object
Bytes are immutable. Calling lower without assigning the result changes nothing.
Result dropped
data.lower()
data
unchanged
Assign it
data = data.lower()
lowercased
3. Not a substitute for casefold
str.casefold handles cases like ß for aggressive matching; bytes has nothing equivalent. For real text comparison, decode and casefold.
ASCII only
'STRASSE'.encode().lower() == 'straße'.encode()
False
Casefold text
'STRASSE'.casefold() == 'straße'.casefold()
True

When to use

Use it
  • Case-insensitive matching of ASCII headers and keywords
  • Normalising identifiers that are ASCII by definition
  • Building lookup keys from ASCII data
Reach for something else
  • Real text → decode, lowercase or casefold as str, encode
  • Anything that might contain non-Latin letters
  • Binary formats, where an accidental A-Z byte would be corrupted

Notes

Complexity
O(n) — one pass over the bytes
Return
A new bytes object of the same length
CPython impl
Objects/bytesobject.c :: stringlib_lower
Memory
Allocates a buffer the same size as the input
Thread-safe
Yes — bytes are immutable

FAQ

Because bytes.lower only maps the 26 ASCII letters. The accented character is two bytes outside that range and is left alone. Decode to str, lowercase there, and encode again.

'HÉLLO'.lower().encode('utf-8')

History

3.0
bytes.lower arrived with the bytes type in the text/binary split.