bytes.capitalize()
Sentence case for ASCII. The part people miss is the second half — it does not just uppercase the first byte, it lowercases all the rest.
Common call
data.capitalize()
Returns
a new bytes object; the original is unchanged
Replaces
data[:1].upper() + data[1:].lower()
Watch out
it LOWERCASES the rest — HELLO WORLD becomes Hello world
bytes.capitalize()
→ bytes
Demo
Live evaluation
Try:
Inputs
sstrdata (encoded as utf-8)
Output
bytes('hello world', 'utf-8').capitalize()
b'Hello world'
Watch the uppercase case: HELLO WORLD becomes Hello world, not HELLO WORLD with a capital H. capitalize uppercases the first byte and LOWERCASES every other letter — it is sentence case, not "make the first letter a capital". When the first byte is not a letter it is left alone and the rest are still lowercased.
Common patterns
Sentence-case an ASCII label
One call instead of slicing and two case conversions.
label = raw_label.capitalize()
Normalise inconsistent input
Shouted or mixed input all lands on the same form.
assert b'hELLO'.capitalize() == b'HELLO'.capitalize()
Capitalise real text
Decode first so non-ASCII letters take part.
data.decode('utf-8').capitalize().encode('utf-8')
Examples
1. Lowercase
b'hello world'.capitalize()
Returns
b'Hello world'2. Uppercase
b'HELLO WORLD'.capitalize()
Returns
b'Hello world'3. Mixed
b'hELLO'.capitalize()
Returns
b'Hello'4. Digit first
b'1abc'.capitalize()
Returns
b'1abc'5. Only first word
b'hello world'.capitalize()
Returns
b'Hello world' # not Hello World6. Empty
b''.capitalize()
Returns
b''Pitfalls
1. It lowercases everything after the first byte
The half people forget. Acronyms, proper nouns and anything else uppercase in the rest of the data get flattened.
Acronym flattened
b'hello NASA'.capitalize()
b'Hello nasa'
Just the first
b'hello NASA'[:1].upper() + b'hello NASA'[1:]
b'Hello NASA'
2. Only the first WORD is capitalised
capitalize is sentence case; title is word case. Expecting every word capitalised is the other common misread.
One capital
b'hello world'.capitalize()
b'Hello world'
Use title
b'hello world'.title()
b'Hello World'
3. Non-ASCII letters are not touched
A leading accented letter stays as it is, and accented letters later in the data are not lowercased either.
Accent ignored
'élan'.encode().capitalize()
b'\xc3\xa9lan'
Work on text
'élan'.capitalize().encode()
b'\xc3\x89lan'
When to use
Use it
- Sentence-casing ASCII labels and messages
- Normalising inconsistently cased ASCII input
Reach for something else
- Data with acronyms or proper nouns that must keep their case
- Capitalising every word → title
- Real text with non-ASCII letters → decode first
Notes
Complexity
O(n) — one pass over the bytes
Return
A new bytes object of the same length
CPython impl
Objects/bytesobject.c :: stringlib_capitalize
Memory
Allocates a buffer the same size as the input
Thread-safe
Yes — bytes are immutable
FAQ
Because capitalize is sentence case: it uppercases the first byte and lowercases every other letter. To uppercase only the first byte and leave the rest alone, slice and use upper on the first byte.
data[:1].upper() + data[1:]
History
3.0
bytes.capitalize arrived with the bytes type in the text/binary split.