bytes.expandtabs()

Not "tab becomes N spaces". Each tab advances to the next tab STOP, so the number of spaces depends on where the tab sits — which is what makes columns line up.

Bytes methodPython 3.0+Live demo
Common call
data.expandtabs(4)
Returns
a new bytes object; the original is unchanged
Replaces
data.replace(b'\t', b' '), which does NOT align columns
Watch out
a tab after 3 characters with tabsize 4 becomes ONE space, not four
bytes.expandtabs(tabsizetabsize — Distance between tab stops in bytes. A tabsize of 0 removes tabs entirely.type: int · default: 8=8)
→ bytes

Demo

Live evaluation
Try:
Inputs
sstrdata with tabs (encoded as utf-8)
tabsizeinttab stop spacing
Output
bytes('a\tb', 'utf-8').expandtabs(4)
b'a b'

Compare the first two cases. After one character, a tab with tabsize 4 becomes three spaces, landing on column 4. After three characters the same tab becomes ONE space, because column 4 is only one step away. That is the tab-stop rule: the tab fills up to the next multiple of tabsize, however far that is. With tabsize 0 the tabs are simply deleted.

Parameters

NameTypeRequiredDescription
tabsizeintno (8)Distance between tab stops in bytes. A tabsize of 0 removes tabs entirely.

Return value

bytes — A new bytes object with each tab replaced by enough spaces to reach the next multiple of tabsize. The column resets at each newline.

Common patterns

Normalise tabbed input before display
Keeps columns aligned regardless of the viewer's tab width.
clean = line.expandtabs(4)
Prepare data for a fixed-width parser
Parsers that count bytes cannot cope with tabs.
record = raw.expandtabs(8)
Strip tabs entirely
A tabsize of 0 removes them without a replace call.
no_tabs = data.expandtabs(0)

Examples

1. Tab at column 1
b'a\tb'.expandtabs(4)
Returns
b'a b'
2. Tab at column 3
b'abc\td'.expandtabs(4)
Returns
b'abc d'
3. Default is 8
b'a\tb'.expandtabs()
Returns
b'a b'
4. Resets at newline
b'ab\tc\nd\te'.expandtabs(4)
Returns
b'ab c\nd e'
5. Tabsize 0 removes
b'a\tb'.expandtabs(0)
Returns
b'ab'
6. No tabs
b'abc'.expandtabs(4)
Returns
b'abc'

Pitfalls

1. A tab is not a fixed number of spaces
The single most common misunderstanding. Each tab advances to the next tab stop, so the same tabsize produces a different number of spaces depending on the column. Using replace instead breaks alignment.
Fixed count
b'abc\td'.replace(b'\t', b'    ')
b'abc d' # misaligned
Tab stops
b'abc\td'.expandtabs(4)
b'abc d' # lands on column 4
2. Columns are counted in bytes
A multi-byte character advances the column by its byte count, not by one, so tabbed text containing non-ASCII misaligns compared with how it displays.
Byte columns
'é\tb'.encode().expandtabs(4)
b'\xc3\xa9 b' # 2 bytes, then 2 spaces
Expand text
'é\tb'.expandtabs(4).encode()
character columns
3. It returns a new object
Bytes are immutable. The result must be assigned or it is lost.
Result dropped
data.expandtabs(4)
data
unchanged
Assign it
data = data.expandtabs(4)
expanded

When to use

Use it
  • Aligning tabbed input for display or a fixed-width parser
  • Normalising whitespace before comparing lines
  • Removing tabs with tabsize 0
Reach for something else
  • You want a fixed number of spaces per tab → replace
  • Text with non-ASCII characters → decode first
  • The data is binary — tab bytes may be meaningful

Notes

Complexity
O(n + inserted spaces) — one pass tracking the column
Return
A new bytes object; longer than the input by the spaces added
CPython impl
Objects/bytesobject.c :: stringlib_expandtabs
Memory
Allocates a buffer sized for the expanded result
Thread-safe
Yes — bytes are immutable

FAQ

Because the tab only needs to advance to the NEXT tab stop. If the data is already at column 3 and stops are every 4, one space reaches column 4. That is what keeps the following text aligned across lines with different prefixes.

b'a\tb'.expandtabs(4)     # b'a   b'
b'abc\tb'.expandtabs(4)   # b'abc b'

History

3.0
bytes.expandtabs arrived with the bytes type in the text/binary split.