csv.Sniffer

Two heuristics from Lib/csv.py: first look for quoted fields and the character next to them; failing that, find the character that appears equally often on most lines. A guess, not a parser — restrict it with delimiters= when you can.

csv classPython 2.3+Live demo
Common call
dialect = csv.Sniffer().sniff(f.read(4096)); f.seek(0)
Returns
a Dialect class → csv.reader(f, dialect)
Replaces
asking the user which delimiter the file uses
Watch out
can guess a letter as the delimiter — pass delimiters=",;\t|"
csv.Sniffer().sniff(sample, delimitersdelimiters — Only these characters may be chosen as the delimiter.type: str | None · default: None=None)
→ type[csv.Dialect]

Demo

Live evaluation
Sample text (\n = line break) → the guessed delimiter, quotechar, doublequote and skipinitialspace.
Try:
Inputs
textstrsample, \n = line break
Code
import csv
text = 'name;age\\nAda;36\\nBob;41'.replace('\\n', '\n')
d = csv.Sniffer().sniff(text)
(d.delimiter, d.quotechar, d.doublequote, d.skipinitialspace)
Result
(';', '"', False, False)

"single\nvalue" has no delimiter at all, yet sniff answers "l": the letter l appears exactly once on both lines, which is all the frequency heuristic checks. With delimiters= it correctly gives up: csv.Error: Could not determine delimiter. In "two candidates" both : and , occur once per line; ties are broken by the preferred order , then tab, ; space :. In "comma + space" every comma is followed by a space, so skipinitialspace is True. doublequote is only True when a quoted field with a doubled quote is found.

Parameters

NameTypeRequiredDescription
samplestryesThe first few kilobytes of the file, as text.
delimitersstr | Noneno (None)Only these characters may be chosen as the delimiter.

Return value

type[csv.Dialect] — A new Dialect subclass (named "dialect") with delimiter, quotechar, doublequote and skipinitialspace guessed; \r\n and QUOTE_MINIMAL otherwise.

Common patterns

Open a CSV file of unknown format
Sniff a sample, rewind, read with the result. Restrict the candidates.
import csv
with open('upload.csv', newline='', encoding='utf-8') as f:
    dialect = csv.Sniffer().sniff(f.read(4096), delimiters=',;\t|')
    f.seek(0)
    rows = list(csv.reader(f, dialect))
Fall back when sniffing fails
sniff raises csv.Error when no candidate is consistent.
import csv
try:
    dialect = csv.Sniffer().sniff(sample, delimiters=',;\t')
except csv.Error:
    dialect = csv.excel
Header too
has_header uses the same sample.
import csv
sniffer = csv.Sniffer()
dialect = sniffer.sniff(sample)
header = sniffer.has_header(sample)

Examples

1. Detect semicolons
import csv csv.Sniffer().sniff('name;age\nAda;36\n').delimiter
Returns
';'
2. Use the result
import csv d = csv.Sniffer().sniff('a;b\n1;2\n') list(csv.reader(['x;"y;z"'], d))
Returns
[['x', 'y;z']]
3. It returns a class
import csv d = csv.Sniffer().sniff('a;b\n1;2\n') (d.__name__, issubclass(d, csv.Dialect), d.lineterminator)
Returns
('dialect', True, '\r\n')
4. Tabs
import csv csv.Sniffer().sniff('a\tb\n1\t2\n').delimiter
Returns
'\t'
5. The preferred order
import csv csv.Sniffer().preferred
Returns
[',', '\t', ';', ' ', ':']
6. Nothing consistent
import csv csv.Sniffer().sniff('x\ny\nz')
Returns
_csv.Error: Could not determine delimiter

Pitfalls

1. Trusting the guess on odd samples
Any character that occurs equally often on every line can win — even a letter.
unrestricted
import csv
csv.Sniffer().sniff('hello\nworld').delimiter
'o'
delimiters=
import csv
try:
    d = csv.Sniffer().sniff('hello\nworld', delimiters=',;\t').delimiter
except csv.Error as e:
    d = str(e)
d
'Could not determine delimiter'
2. Reading after sniffing without seek(0)
f.read(n) moved the file position; the reader starts after the sample.
no seek
import csv
with open('d.csv', 'w', newline='') as f:
    f.write('a;b\r\n1;2\r\n')
with open('d.csv', newline='') as f:
    d = csv.Sniffer().sniff(f.read(1024))
    rows = list(csv.reader(f, d))
rows
[]
f.seek(0)
import csv
with open('d.csv', 'w', newline='') as f:
    f.write('a;b\r\n1;2\r\n')
with open('d.csv', newline='') as f:
    d = csv.Sniffer().sniff(f.read(1024))
    f.seek(0)
    rows = list(csv.reader(f, d))
rows
[['a', 'b'], ['1', '2']]

When to use

Use it
  • User uploads in unknown formats (comma vs semicolon vs tab)
  • Quick scripts over files from many sources
Reach for something else
  • Files whose format you know — just pass the delimiter
  • Single-column files: there is no delimiter to find

Notes

CPython impl
Lib/csv.py — sniff() tries _guess_quote_and_delimiter (four regexes looking for quote-delimited text), then _guess_delimiter (per-line character frequency over the 127 ASCII characters, in chunks of 10 lines, accepting a mode that holds on at least 90% of lines)
Result
A new class "dialect" (a csv.Dialect subclass, _name = "sniffed") — not an instance and not registered
Sample size
More lines make the frequency heuristic more reliable; a partial last line can mislead it

FAQ

csv.Sniffer().sniff(sample).delimiter, where sample is the first few KB of the file. Pass delimiters=',;\t|' to limit the candidates, and seek(0) before reading the file.