bytes.startswith()
The standard way to sniff a file format from its magic number. The tuple form checks several signatures in one call.
Common call
data.startswith(b'\x89PNG')
Returns
bool — True or False, never an error for a missing prefix
Replaces
data[:len(prefix)] == prefix, which allocates a slice
Watch out
a LIST of prefixes is a TypeError; it must be a tuple
bytes.startswith(prefix[, start[, end]])
→ bool
Demo
Live evaluation
Try:
Inputs
sstrdata (encoded as utf-8)
prefixstrprefix to test
Output
bytes('abcdef', 'utf-8').startswith(bytes('abc', 'utf-8'))
True
A straightforward prefix test returning True or False. Two edge cases are worth knowing: the empty prefix is always True, because every buffer begins with nothing, and a prefix longer than the data is simply False rather than an error. Comparison is byte for byte, so this is exactly what you want for checking file signatures.
Parameters
| Name | Type | Required | Description |
|---|---|---|---|
| prefix | bytes | tuple | yes | Bytes to test for, or a TUPLE of alternatives. A list raises TypeError. |
| start | int | no (0) | Byte offset to test from, as though the buffer began there. |
| end | int | no (len) | Byte offset to stop at, exclusive. |
Return value
bool — True if the buffer begins with prefix. A tuple of prefixes returns True if ANY of them match.
Common patterns
Detect a file format
Magic numbers at the head of a file identify its type.
if data.startswith(b'\x89PNG\r\n\x1a\n'): kind = 'png'
Check several signatures at once
The tuple form avoids a chain of or clauses.
if data.startswith((b'GIF87a', b'GIF89a')): kind = 'gif'
Skip a known header
Test then slice past it.
if data.startswith(MAGIC): body = data[len(MAGIC):]
Examples
1. Matches
b'abcdef'.startswith(b'abc')
Returns
True2. Does not match
b'abcdef'.startswith(b'xyz')
Returns
False3. Empty is always True
b'abc'.startswith(b'')
Returns
True4. Too long
b'ab'.startswith(b'abc')
Returns
False5. Tuple of options
b'abc'.startswith((b'x', b'ab'))
Returns
True6. From an offset
b'abcdef'.startswith(b'cd', 2)
Returns
TruePitfalls
1. A list of prefixes is a TypeError
Only a tuple works. The distinction looks arbitrary and catches people building the alternatives dynamically, since a list comprehension is the natural way to produce them.
List rejected
b'abc'.startswith([b'a', b'b'])
TypeError: startswith first arg must be bytes or a tuple of bytes, not list
Convert to tuple
b'abc'.startswith(tuple(prefixes))
True
2. The empty prefix is always True
Every buffer begins with nothing, so an accidentally empty prefix makes every test pass. This turns a format check into a no-op that silently accepts everything.
Always passes
b'abc'.startswith(b'')
True
Require a prefix
if prefix and data.startswith(prefix):
meaningful
3. The prefix must be bytes
A str raises rather than comparing as False, which is at least loud. It is the usual failure when text-oriented code meets a file opened in binary mode.
str rejected
b'abc'.startswith('abc')
TypeError: startswith first arg must be bytes or a tuple of bytes, not str
Bytes literal
b'abc'.startswith(b'abc')
True
4. start shifts the test, not just the search
With a start offset the method behaves as though the buffer began there, so the prefix is matched at that position rather than anywhere after it.
Not a search
b'abcdef'.startswith(b'cd', 1)
False # position 1 is b
Correct offset
b'abcdef'.startswith(b'cd', 2)
True
When to use
Use it
- Identifying a file format from its magic number
- Checking a protocol header or record marker
- Testing several possible signatures with the tuple form
Reach for something else
- The match could be anywhere → find or the in operator
- The data is text → decode first and use str.startswith
- You need the matched length → compare against the prefix directly
Notes
Complexity
O(len(prefix)) — stops at the first differing byte
Return
A bool; never raises for a non-matching prefix
CPython impl
Objects/bytesobject.c :: bytes_startswith
Memory
No allocation — unlike slicing, which copies
Thread-safe
Yes — bytes are immutable
FAQ
Historical, and now fixed by the API. A tuple signals a fixed set of alternatives, and it is hashable so CPython can treat it as a constant. If your prefixes come from a list, wrap it with tuple().
data.startswith(tuple(prefixes))
History
3.0
bytes.startswith arrived with the bytes type in the text/binary split.