Sniffer.has_header
A vote per column: if a column below the first row is all numbers, or all strings of one length, the first row gets +1 when it breaks that pattern and -1 when it fits. More plus than minus means "header".
Demo
import csv text = 'name,age\\nAda,36\\nBob,41'.replace('\\n', '\n') csv.Sniffer().has_header(text)
In "numeric column" age is a number in every data row and "age" is not → +1, and Ada/Bob are both 3 characters long while "name" has 4 → +1. In "same lengths" city (London 6, Paris 5) varies and drops out, but the 3-letter names still vote +1. "a,b / c,d" scores -2: every value is one character, header included. With one row only, no column gets a type, and the code's attempt to call None counts as a vote for the header — so True. A sample without a detectable delimiter fails in sniff first.
Parameters
| Name | Type | Required | Description |
|---|---|---|---|
| sample | str | yes | The first lines of the file. It is sniffed for the dialect first (and raises csv.Error if that fails). |
Return value
bool — True when the first row looks different from the rows below it.
Common patterns
import csv with open('upload.csv', newline='', encoding='utf-8') as f: sample = f.read(4096) f.seek(0) sniffer = csv.Sniffer() dialect = sniffer.sniff(sample) if sniffer.has_header(sample): rows = list(csv.DictReader(f, dialect=dialect)) else: rows = list(csv.reader(f, dialect))
rows = csv.reader(f, dialect) if csv.Sniffer().has_header(sample): next(rows)
Examples
Pitfalls
import csv csv.Sniffer().has_header('first,last\nAda,Lovelace\nAlan,Turing\n')
import csv header = next(csv.reader(['first,last'])) header == ['first', 'last']
import csv csv.Sniffer().has_header('2023,2024\n10,12\n11,15\n')
import csv rows = list(csv.reader(['2023,2024', '10,12', '11,15'])) header, data = rows[0], rows[1:] header
When to use
- Mixed numeric/text files of unknown origin
- Files you produce yourself — you know whether there is a header
- All-text or all-numeric data, where the heuristic has nothing to compare
Notes
FAQ
csv.Sniffer().has_header(sample) with the first few KB of the file. It is a heuristic: reliable when some columns are numeric or fixed-length, unreliable for all-text files.