urllib.parse.urlsplit
The URL splitter to reach for first: five plain-string parts, nothing decoded, ;params left in the path. urlunsplit is its inverse — up to cosmetics: an empty '?' or '#' is dropped on the way back.
Demo
from urllib.parse import urlsplit urlsplit('HTTPS://Example.com:8080/a/b;v=1?q=x#top')
The scheme is lowercased ('HTTPS' → 'https') but the host is not — use .hostname for that. Leading spaces and control characters are stripped before parsing. Only the first # starts the fragment. On the way back, urlunsplit adds the '/' a path needs after a netloc, and an empty query or fragment simply disappears, so 'https://example.com/a?#' comes back as 'https://example.com/a'.
Parameters
| Name | Type | Required | Description |
|---|---|---|---|
| url | str | bytes | yes | The URL. Leading C0 control characters and spaces are stripped (3.12+); tab, CR and LF are removed everywhere (3.10+). The scheme is lowercased. |
| scheme | str | bytes | no ('') | Default scheme for URLs that have none. |
| allow_fragments | bool | no (True) | False keeps #... as part of the path or query. |
Return value
SplitResult — A named 5-tuple (scheme, netloc, path, query, fragment) with .hostname, .port, .username, .password and .geturl(). SplitResultBytes for bytes input.
Common patterns
from urllib.parse import urlsplit, urlunsplit s = urlsplit(url) base = urlunsplit((s.scheme, s.netloc, s.path, '', ''))
from urllib.parse import urlsplit, urlencode new_url = urlsplit(url)._replace(query=urlencode(params)).geturl()
from urllib.parse import urlsplit if urlsplit(url).scheme not in ('http', 'https'): raise ValueError('only http(s) URLs are allowed')
Examples
Pitfalls
from urllib.parse import urlsplit urlsplit('https://example.com/a')._replace(host='example.org')
from urllib.parse import urlsplit urlsplit('https://example.com/a')._replace(netloc='example.org').geturl()
from urllib.parse import urlunsplit urlunsplit(('https', 'example.com', '/a', 'q=1'))
from urllib.parse import urlunsplit urlunsplit(('https', 'example.com', '/a', 'q=1', ''))
from urllib.parse import urlsplit urlsplit('https://Example.COM/').netloc == 'example.com'
from urllib.parse import urlsplit urlsplit('https://Example.COM/').hostname == 'example.com'
When to use
- Any time you need the parts of a URL
- Rebuilding a URL after changing a part (_replace + geturl, or urlunsplit)
- Legacy ;params handling → urlparse
- Resolving a relative link against a page URL → urljoin
- Decoding the query → parse_qs on .query
Notes
FAQ
urlsplit. The Python docs say urlsplit() should generally be used instead of urlparse(); urlparse only adds the split of ;params off the last path segment, a feature from obsolete RFCs.