bytes.splitlines()

The right way to split binary-mode lines, because it handles CRLF. But it recognises only three terminators, where str.splitlines recognises a dozen — a real, documented difference.

Bytes methodPython 3.0+Live demo
Common call
lines = data.splitlines()
Returns
list of bytes, terminators stripped
Replaces
split(b'\n') followed by rstrip(b'\r') on every line
Watch out
bytes splits ONLY on \n, \r, \r\n — not \x0b, \x0c, \x1c, \x85
bytes.splitlines(keependskeepends — When True, each line keeps its own terminator, so joining the parts reproduces the original exactly.type: bool · default: False=False)
→ list

Demo

Live evaluation
Try:
Inputs
sstrdata with line breaks (encoded as utf-8)
Output
bytes('a\nb\r\nc', 'utf-8').splitlines()
[b'a', b'b', b'c']

Both LF and CRLF are recognised and removed, so a\nb\r\nc gives three clean lines. A trailing terminator does NOT create an empty last element — a\nb\n gives two lines, not three — while a blank line in the middle IS kept as an empty element. The exotic case is the one that separates bytes from str: vertical tab and form feed are line breaks to str.splitlines but not to bytes.splitlines, so the data comes back as a single line.

Parameters

NameTypeRequiredDescription
keependsboolno (False)When True, each line keeps its own terminator, so joining the parts reproduces the original exactly.

Return value

list — A list of lines with the terminators removed, unless keepends is True. A trailing terminator does not produce an empty final line.

Common patterns

Split a binary-mode file into lines
Handles Windows and Unix endings in one call.
with open(path, 'rb') as f:
    lines = f.read().splitlines()
Round-trip with keepends
Lines keep their terminators, so joining restores the original.
parts = data.splitlines(keepends=True)
assert b''.join(parts) == data
Count lines
A trailing newline does not inflate the count.
n = len(data.splitlines())

Examples

1. LF and CRLF
b'a\nb\r\nc'.splitlines()
Returns
[b'a', b'b', b'c']
2. keepends
b'a\nb\r\nc'.splitlines(True)
Returns
[b'a\n', b'b\r\n', b'c']
3. Trailing newline
b'a\nb\n'.splitlines()
Returns
[b'a', b'b']
4. split differs
b'a\nb\n'.split(b'\n')
Returns
[b'a', b'b', b''] # extra empty
5. Exotic not split
b'a\x0bb'.splitlines()
Returns
[b'a\x0bb']
6. str DOES split
'a\x0bb'.splitlines()
Returns
['a', 'b']

Pitfalls

1. bytes and str disagree on what a line break is
str.splitlines treats vertical tab, form feed, file/group/record separators, NEL and the Unicode line and paragraph separators as line breaks. bytes.splitlines recognises only \n, \r and \r\n. The same data decodes into a different number of lines.
Different counts
len(b'a\x0cb'.splitlines()), len('a\x0cb'.splitlines())
(1, 2)
Pick one domain
data.decode().splitlines()   # if you want the str rules
consistent
2. A trailing newline does not give an empty last line
Opposite to split(b"\n"), which produces a trailing empty element. Code expecting the extra element, or expecting its absence, will be off by one depending on which method it assumed.
split has extra
b'a\n'.split(b'\n')
[b'a', b'']
splitlines does not
b'a\n'.splitlines()
[b'a']
3. A bare \r is a line break
Old Mac endings and stray carriage returns split lines on their own. Data that mixes CR into content, such as progress output, breaks into more lines than expected.
Splits on CR
b'50%\r100%'.splitlines()
[b'50%', b'100%']
Split on LF only
b'50%\r100%'.split(b'\n')
[b'50%\r100%']

When to use

Use it
  • Splitting binary-mode file or socket data into lines
  • Handling mixed LF and CRLF without normalising first
  • Counting lines without the trailing-newline off-by-one
Reach for something else
  • You need the str set of line breaks → decode first
  • A bare \r must NOT split → split(b"\n") and rstrip
  • You need the empty trailing element that split gives

Notes

Complexity
O(n) — one scan, plus an allocation per line
Return
A new list of new bytes objects
CPython impl
Objects/bytesobject.c :: stringlib_splitlines
Memory
Allocates the list and every line
Thread-safe
Yes — bytes are immutable

FAQ

Because bytes have no notion of Unicode, and most of the extra terminators are Unicode or control characters whose meaning depends on the encoding. The bytes version sticks to the three ASCII line endings every format agrees on. It is documented, but easy to miss.

b'a\x0cb'.splitlines()   # [b'a\x0cb']
'a\x0cb'.splitlines()    # ['a', 'b']

History

3.0
bytes.splitlines arrived with the bytes type in the text/binary split.
3.2
keepends became usable as a keyword argument.