Sniffer.has_header

A vote per column: if a column below the first row is all numbers, or all strings of one length, the first row gets +1 when it breaks that pattern and -1 when it fits. More plus than minus means "header".

Sniffer methodPython 2.3+Live demo
Common call
csv.Sniffer().has_header(sample)
Returns
True / False
Replaces
asking whether the file has a header
Watch out
text-only columns of varying length do not vote at all
Sniffer.has_header(samplesample — The first lines of the file. It is sniffed for the dialect first (and raises csv.Error if that fails).type: str · required)
→ bool

Demo

Live evaluation
Sample text (\n = line break) → the guess.
Try:
Inputs
textstrsample, \n = line break
Code
import csv
text = 'name,age\\nAda,36\\nBob,41'.replace('\\n', '\n')
csv.Sniffer().has_header(text)
Result
True

In "numeric column" age is a number in every data row and "age" is not → +1, and Ada/Bob are both 3 characters long while "name" has 4 → +1. In "same lengths" city (London 6, Paris 5) varies and drops out, but the 3-letter names still vote +1. "a,b / c,d" scores -2: every value is one character, header included. With one row only, no column gets a type, and the code's attempt to call None counts as a vote for the header — so True. A sample without a detectable delimiter fails in sniff first.

Parameters

NameTypeRequiredDescription
samplestryesThe first lines of the file. It is sniffed for the dialect first (and raises csv.Error if that fails).

Return value

bool — True when the first row looks different from the rows below it.

Common patterns

Choose reader or DictReader
Sniff the dialect and the header from one sample.
import csv
with open('upload.csv', newline='', encoding='utf-8') as f:
    sample = f.read(4096)
    f.seek(0)
    sniffer = csv.Sniffer()
    dialect = sniffer.sniff(sample)
    if sniffer.has_header(sample):
        rows = list(csv.DictReader(f, dialect=dialect))
    else:
        rows = list(csv.reader(f, dialect))
Skip the header if there is one
next() on the reader drops the first row.
rows = csv.reader(f, dialect)
if csv.Sniffer().has_header(sample):
    next(rows)

Examples

1. Numbers under a text header
import csv csv.Sniffer().has_header('name,age\nAda,36\nBob,41\n')
Returns
True
2. All numeric rows
import csv csv.Sniffer().has_header('1,2\n3,4\n5,6\n')
Returns
False
3. Text of equal length
import csv csv.Sniffer().has_header('code,n\nAB,1\nCD,2\n')
Returns
True
4. Complex numbers count as numbers
import csv csv.Sniffer().has_header('a,b\n1j,2\n3+4j,5\n')
Returns
True
5. Fails when sniff fails
import csv csv.Sniffer().has_header('x\ny\nz')
Returns
_csv.Error: Could not determine delimiter

Pitfalls

1. Text columns with varying lengths
Columns whose values differ in length and are not numbers cast no vote. If every column is like that, the answer is False even with an obvious header.
text only
import csv
csv.Sniffer().has_header('first,last\nAda,Lovelace\nAlan,Turing\n')
False
decide yourself
import csv
header = next(csv.reader(['first,last']))
header == ['first', 'last']
True
2. Numeric headers
A header made of numbers (years, IDs) fits the numeric columns below it and is voted down.
year columns
import csv
csv.Sniffer().has_header('2023,2024\n10,12\n11,15\n')
False
known layout
import csv
rows = list(csv.reader(['2023,2024', '10,12', '11,15']))
header, data = rows[0], rows[1:]
header
['2023', '2024']

When to use

Use it
  • Mixed numeric/text files of unknown origin
Reach for something else
  • Files you produce yourself — you know whether there is a header
  • All-text or all-numeric data, where the heuristic has nothing to compare

Notes

CPython impl
Lib/csv.py — reads the sample with the sniffed dialect, checks up to 21 rows after the first; per column the type is complex (if complex(value) parses) or len(value); columns with mixed types are dropped; rows with a different number of fields are skipped
Vote
Each remaining column: +1 if the header value breaks the column type, -1 if it fits; has_header returns sum > 0

FAQ

csv.Sniffer().has_header(sample) with the first few KB of the file. It is a heuristic: reliable when some columns are numeric or fixed-length, unreliable for all-text files.