bytes.upper()

Touches exactly 26 byte values, 0x61 to 0x7a. Anything encoded above ASCII passes through untouched, which is why an accented letter stays lowercase.

Bytes methodPython 3.0+Live demo
Common call
data.upper()
Returns
a new bytes object; the original is unchanged
Replaces
a manual loop over bytes with an ord/chr dance
Watch out
accented and non-Latin letters are NOT uppercased — decode first
bytes.upper()
→ bytes

Demo

Live evaluation
Try:
Inputs
sstrdata (encoded as utf-8)
Output
bytes('hello', 'utf-8').upper()
b'HELLO'

Only the 26 ASCII lowercase bytes change. Digits, punctuation and spaces pass through, and so does anything outside ASCII — look at the accented case, where the h, l and o become capitals but the é does not, because its two UTF-8 bytes are not in the a-z range. The str version would uppercase it; the bytes version has no idea it is a letter.

Common patterns

Normalise an ASCII protocol token
HTTP methods, hex digests and similar are ASCII by definition.
if method.upper() == b'GET':
    ...
Case-insensitive comparison
Uppercase both sides before comparing.
if header_name.upper() == expected.upper():
Uppercase real text
Decode first so non-ASCII letters are handled.
data.decode('utf-8').upper().encode('utf-8')

Examples

1. Lowercase
b'hello'.upper()
Returns
b'HELLO'
2. Mixed
b'HeLLo World'.upper()
Returns
b'HELLO WORLD'
3. Digits stay
b'abc123'.upper()
Returns
b'ABC123'
4. Accent ignored
'héllo'.encode().upper()
Returns
b'H\xc3\xa9LLO'
5. Empty
b''.upper()
Returns
b''
6. New object
d = b'a' d.upper() is d
Returns
False

Pitfalls

1. Non-ASCII letters are not touched
The bytes method only knows the 26 ASCII letters. An accented character is two UTF-8 bytes outside that range and stays exactly as it was, so the result is a half-uppercased word that looks like a bug.
Half converted
'héllo'.encode().upper()
b'H\xc3\xa9LLO'
Work on text
'héllo'.upper().encode()
b'H\xc3\x89LLO'
2. It returns a new object
Bytes are immutable, so upper cannot change the original. Calling it without assigning the result does nothing at all.
Result dropped
data.upper()
data
unchanged
Assign it
data = data.upper()
uppercased
3. Not a Unicode case mapping
There are no special cases — no ß to SS, no dotted I. Text with any of that needs decoding and the str method, which understands the full mapping.
Naive
'straße'.encode().upper()
ß bytes untouched
Unicode-aware
'straße'.upper()
'STRASSE'

When to use

Use it
  • Normalising ASCII protocol tokens and identifiers
  • Case-insensitive comparison of ASCII data
  • Hex digests and other guaranteed-ASCII content
Reach for something else
  • Real text → decode, uppercase as str, encode
  • Anything that might contain non-Latin letters
  • You need the original kept — assign the result to a new name

Notes

Complexity
O(n) — one pass over the bytes
Return
A new bytes object of the same length
CPython impl
Objects/bytesobject.c :: stringlib_upper
Memory
Allocates a buffer the same size as the input
Thread-safe
Yes — bytes are immutable

FAQ

Because bytes.upper only maps the 26 ASCII letters. An accented character is stored as two bytes outside that range, and the method leaves them alone. Decode to str, uppercase there, and encode again.

'héllo'.upper().encode('utf-8')

History

3.0
bytes.upper arrived with the bytes type in the text/binary split.