urllib.parse.urlparse

urlparse cuts a URL at its delimiters without decoding or validating much: the parts are plain strings, % escapes stay as they are, and a URL without '//' has no netloc at all. The extra attributes do the fiddly work — .hostname is lowercased, .port is an int (or a ValueError).

urllib.parse functionPython 3.0+Live demo
Common call
urlparse(url).hostname
Returns
'example.com' — lowercased, without port or user info
Replaces
url.split('/')[2] and other string surgery
Watch out
'example.com/path' (no scheme, no //) is all path; .port raises ValueError for bad ports
urllib.parse.urlparse(urlurl — The URL. Leading C0 control characters and spaces are stripped (3.12+), and tab, CR and LF are removed everywhere (3.10+).type: str | bytes · required, schemescheme — The scheme to report when the URL has none. It never overrides a scheme that is present.type: str | bytes · default: ''='', allow_fragmentsallow_fragments — False leaves #... inside the path or query instead of splitting it into fragment.type: bool · default: True=True)
→ ParseResult

Demo

Live evaluation
The six components. Try a URL without a scheme, or with ;params.
Try:
Inputs
urlstra URL
Code
from urllib.parse import urlparse
urlparse('https://user:pw@Example.com:8080/a/b;v=1?q=x#top')
Result
ParseResult(scheme='https', netloc='user:pw@Example.com:8080', path='/a/b', params='v=1', query='q=x', fragment='top')

Without '//' there is no netloc: 'example.com/path' is entirely path, and 'mailto:ada@example.com' is a scheme plus a path. .hostname lowercases the host and drops the IPv6 brackets; .port converts to int and raises ValueError when the text is not digits ('Port could not be cast to integer value as ...') or the number is above 65535 ('Port out of range 0-65535'). An unbalanced [ in the netloc fails already in urlparse: 'Invalid IPv6 URL'.

Parameters

NameTypeRequiredDescription
urlstr | bytesyesThe URL. Leading C0 control characters and spaces are stripped (3.12+), and tab, CR and LF are removed everywhere (3.10+).
schemestr | bytesno ('')The scheme to report when the URL has none. It never overrides a scheme that is present.
allow_fragmentsboolno (True)False leaves #... inside the path or query instead of splitting it into fragment.

Return value

ParseResult — A named 6-tuple (scheme, netloc, path, params, query, fragment) with .hostname, .port, .username, .password and .geturl(). ParseResultBytes for bytes input.

Common patterns

Domain of a URL
.hostname, not .netloc — netloc still has the port and user info.
from urllib.parse import urlparse
host = urlparse(url).hostname  # None if the URL has no //host
Port with a default
.port is None when absent.
from urllib.parse import urlparse
u = urlparse(url)
port = u.port or (443 if u.scheme == 'https' else 80)
Accept URLs typed without a scheme
Prefix '//' (or 'https://') so the host lands in netloc.
from urllib.parse import urlparse
if '//' not in text:
    text = 'https://' + text
u = urlparse(text)
Change one part
ParseResult is a namedtuple: _replace returns a new one, geturl() rebuilds the string.
from urllib.parse import urlparse
clean = urlparse(url)._replace(query='', fragment='').geturl()

Examples

1. All six components
from urllib.parse import urlparse urlparse('https://user:pw@Example.com:8080/a/b;v=1?q=x#top')
Returns
ParseResult(scheme='https', netloc='user:pw@Example.com:8080', path='/a/b', params='v=1', query='q=x', fragment='top')
2. netloc helpers
from urllib.parse import urlparse u = urlparse('https://user:pw@Example.com:8080/a/b?q=x#top') (u.hostname, u.port, u.username, u.password)
Returns
('example.com', 8080, 'user', 'pw')
3. No // means no netloc
from urllib.parse import urlparse urlparse('example.com/path')
Returns
ParseResult(scheme='', netloc='', path='example.com/path', params='', query='', fragment='')
4. 'host:port' looks like a scheme
from urllib.parse import urlparse urlparse('localhost:8000')
Returns
ParseResult(scheme='localhost', netloc='', path='8000', params='', query='', fragment='')
5. Port out of range
from urllib.parse import urlparse urlparse('https://example.com:99999/').port
Returns
ValueError: Port out of range 0-65535
6. urlunparse
from urllib.parse import urlunparse urlunparse(('https', 'example.com', '/a', '', 'q=1', ''))
Returns
'https://example.com/a?q=1'
7. ParseResult.geturl
from urllib.parse import ParseResult ParseResult('https', 'example.com', '/a', '', '', 'top').geturl()
Returns
'https://example.com/a#top'
8. Tabs and newlines are removed
from urllib.parse import urlparse urlparse('https://exa\tmple.com/').netloc
Returns
'example.com'

Pitfalls

1. Parsing a URL without a scheme
Only '//' starts a netloc. Text typed by users ('example.com/page') has no host as far as urlparse is concerned.
bare domain
from urllib.parse import urlparse
u = urlparse('example.com/page')
(u.hostname, u.path)
(None, 'example.com/page')
add '//'
from urllib.parse import urlparse
u = urlparse('//example.com/page')
(u.hostname, u.path)
('example.com', '/page')
2. Using netloc as the host name
netloc is everything between // and the path: user info, host and port, with the original capitals.
netloc
from urllib.parse import urlparse
urlparse('https://ada@Example.com:8080/').netloc
'ada@Example.com:8080'
hostname
from urllib.parse import urlparse
urlparse('https://ada@Example.com:8080/').hostname
'example.com'
3. urlunparse wants six parts
urlunparse takes the 6-tuple of urlparse (with params). A 5-tuple from urlsplit belongs in urlunsplit.
5 parts
from urllib.parse import urlunparse
urlunparse(('https', 'example.com', '/a', '', 'q=1'))
ValueError: not enough values to unpack (expected 7, got 6)
6 parts
from urllib.parse import urlunparse
urlunparse(('https', 'example.com', '/a', '', 'q=1', ''))
'https://example.com/a?q=1'

When to use

Use it
  • Reading the scheme, host, port or path of a URL
  • Changing one part of a URL: _replace + geturl
  • Old ;params syntax matters to you (rare)
Reach for something else
  • Most new code → urlsplit (same thing without the rarely used params split)
  • Query parameters → parse_qs(urlsplit(url).query)
  • Validating that a string is a URL → urlparse accepts almost anything

Notes

CPython impl
urlparse calls urlsplit and then, for schemes in uses_params (http, https, ftp, …), splits ;params off the last path segment
Validation
Very little: unbalanced brackets, invalid bracketed hosts and NFKC-unsafe netlocs raise ValueError; everything else parses
Exceptions
ValueError from urlparse for bad IPv6 brackets; ValueError from .port; TypeError when str and bytes are mixed

FAQ

urlparse(url).hostname — it strips user info and port and lowercases the host. It is None when the URL has no '//' part, e.g. 'example.com/page'.