urllib.parse.urlparse
urlparse cuts a URL at its delimiters without decoding or validating much: the parts are plain strings, % escapes stay as they are, and a URL without '//' has no netloc at all. The extra attributes do the fiddly work — .hostname is lowercased, .port is an int (or a ValueError).
Demo
from urllib.parse import urlparse urlparse('https://user:pw@Example.com:8080/a/b;v=1?q=x#top')
Without '//' there is no netloc: 'example.com/path' is entirely path, and 'mailto:ada@example.com' is a scheme plus a path. .hostname lowercases the host and drops the IPv6 brackets; .port converts to int and raises ValueError when the text is not digits ('Port could not be cast to integer value as ...') or the number is above 65535 ('Port out of range 0-65535'). An unbalanced [ in the netloc fails already in urlparse: 'Invalid IPv6 URL'.
Parameters
| Name | Type | Required | Description |
|---|---|---|---|
| url | str | bytes | yes | The URL. Leading C0 control characters and spaces are stripped (3.12+), and tab, CR and LF are removed everywhere (3.10+). |
| scheme | str | bytes | no ('') | The scheme to report when the URL has none. It never overrides a scheme that is present. |
| allow_fragments | bool | no (True) | False leaves #... inside the path or query instead of splitting it into fragment. |
Return value
ParseResult — A named 6-tuple (scheme, netloc, path, params, query, fragment) with .hostname, .port, .username, .password and .geturl(). ParseResultBytes for bytes input.
Common patterns
from urllib.parse import urlparse host = urlparse(url).hostname # None if the URL has no //host
from urllib.parse import urlparse u = urlparse(url) port = u.port or (443 if u.scheme == 'https' else 80)
from urllib.parse import urlparse if '//' not in text: text = 'https://' + text u = urlparse(text)
from urllib.parse import urlparse clean = urlparse(url)._replace(query='', fragment='').geturl()
Examples
Pitfalls
from urllib.parse import urlparse u = urlparse('example.com/page') (u.hostname, u.path)
from urllib.parse import urlparse u = urlparse('//example.com/page') (u.hostname, u.path)
from urllib.parse import urlparse urlparse('https://ada@Example.com:8080/').netloc
from urllib.parse import urlparse urlparse('https://ada@Example.com:8080/').hostname
from urllib.parse import urlunparse urlunparse(('https', 'example.com', '/a', '', 'q=1'))
from urllib.parse import urlunparse urlunparse(('https', 'example.com', '/a', '', 'q=1', ''))
When to use
- Reading the scheme, host, port or path of a URL
- Changing one part of a URL: _replace + geturl
- Old ;params syntax matters to you (rare)
- Most new code → urlsplit (same thing without the rarely used params split)
- Query parameters → parse_qs(urlsplit(url).query)
- Validating that a string is a URL → urlparse accepts almost anything
Notes
FAQ
urlparse(url).hostname — it strips user info and port and lowercases the host. It is None when the URL has no '//' part, e.g. 'example.com/page'.