csv.Sniffer
Two heuristics from Lib/csv.py: first look for quoted fields and the character next to them; failing that, find the character that appears equally often on most lines. A guess, not a parser — restrict it with delimiters= when you can.
Demo
import csv text = 'name;age\\nAda;36\\nBob;41'.replace('\\n', '\n') d = csv.Sniffer().sniff(text) (d.delimiter, d.quotechar, d.doublequote, d.skipinitialspace)
"single\nvalue" has no delimiter at all, yet sniff answers "l": the letter l appears exactly once on both lines, which is all the frequency heuristic checks. With delimiters= it correctly gives up: csv.Error: Could not determine delimiter. In "two candidates" both : and , occur once per line; ties are broken by the preferred order , then tab, ; space :. In "comma + space" every comma is followed by a space, so skipinitialspace is True. doublequote is only True when a quoted field with a doubled quote is found.
Parameters
| Name | Type | Required | Description |
|---|---|---|---|
| sample | str | yes | The first few kilobytes of the file, as text. |
| delimiters | str | None | no (None) | Only these characters may be chosen as the delimiter. |
Return value
type[csv.Dialect] — A new Dialect subclass (named "dialect") with delimiter, quotechar, doublequote and skipinitialspace guessed; \r\n and QUOTE_MINIMAL otherwise.
Common patterns
import csv with open('upload.csv', newline='', encoding='utf-8') as f: dialect = csv.Sniffer().sniff(f.read(4096), delimiters=',;\t|') f.seek(0) rows = list(csv.reader(f, dialect))
import csv try: dialect = csv.Sniffer().sniff(sample, delimiters=',;\t') except csv.Error: dialect = csv.excel
import csv sniffer = csv.Sniffer() dialect = sniffer.sniff(sample) header = sniffer.has_header(sample)
Examples
Pitfalls
import csv csv.Sniffer().sniff('hello\nworld').delimiter
import csv try: d = csv.Sniffer().sniff('hello\nworld', delimiters=',;\t').delimiter except csv.Error as e: d = str(e) d
import csv with open('d.csv', 'w', newline='') as f: f.write('a;b\r\n1;2\r\n') with open('d.csv', newline='') as f: d = csv.Sniffer().sniff(f.read(1024)) rows = list(csv.reader(f, d)) rows
import csv with open('d.csv', 'w', newline='') as f: f.write('a;b\r\n1;2\r\n') with open('d.csv', newline='') as f: d = csv.Sniffer().sniff(f.read(1024)) f.seek(0) rows = list(csv.reader(f, d)) rows
When to use
- User uploads in unknown formats (comma vs semicolon vs tab)
- Quick scripts over files from many sources
- Files whose format you know — just pass the delimiter
- Single-column files: there is no delimiter to find
Notes
FAQ
csv.Sniffer().sniff(sample).delimiter, where sample is the first few KB of the file. Pass delimiters=',;\t|' to limit the candidates, and seek(0) before reading the file.