bytes.splitlines()
The right way to split binary-mode lines, because it handles CRLF. But it recognises only three terminators, where str.splitlines recognises a dozen — a real, documented difference.
Demo
Both LF and CRLF are recognised and removed, so a\nb\r\nc gives three clean lines. A trailing terminator does NOT create an empty last element — a\nb\n gives two lines, not three — while a blank line in the middle IS kept as an empty element. The exotic case is the one that separates bytes from str: vertical tab and form feed are line breaks to str.splitlines but not to bytes.splitlines, so the data comes back as a single line.
Parameters
| Name | Type | Required | Description |
|---|---|---|---|
| keepends | bool | no (False) | When True, each line keeps its own terminator, so joining the parts reproduces the original exactly. |
Return value
list — A list of lines with the terminators removed, unless keepends is True. A trailing terminator does not produce an empty final line.
Common patterns
with open(path, 'rb') as f: lines = f.read().splitlines()
parts = data.splitlines(keepends=True) assert b''.join(parts) == data
n = len(data.splitlines())
Examples
Pitfalls
len(b'a\x0cb'.splitlines()), len('a\x0cb'.splitlines())
data.decode().splitlines() # if you want the str rules
b'a\n'.split(b'\n')
b'a\n'.splitlines()
b'50%\r100%'.splitlines()
b'50%\r100%'.split(b'\n')
When to use
- Splitting binary-mode file or socket data into lines
- Handling mixed LF and CRLF without normalising first
- Counting lines without the trailing-newline off-by-one
- You need the str set of line breaks → decode first
- A bare \r must NOT split → split(b"\n") and rstrip
- You need the empty trailing element that split gives
Notes
FAQ
Because bytes have no notion of Unicode, and most of the extra terminators are Unicode or control characters whose meaning depends on the encoding. The bytes version sticks to the three ASCII line endings every format agrees on. It is documented, but easy to miss.
b'a\x0cb'.splitlines() # [b'a\x0cb'] 'a\x0cb'.splitlines() # ['a', 'b']