urllib.parse.urlsplit

The URL splitter to reach for first: five plain-string parts, nothing decoded, ;params left in the path. urlunsplit is its inverse — up to cosmetics: an empty '?' or '#' is dropped on the way back.

urllib.parse functionPython 3.0+Live demo
Common call
urlsplit(url).path
Returns
'/a/b;v=1' — params stay in the path
Replaces
urlparse when you do not need ;params split off
Watch out
urlunsplit takes exactly 5 parts — not the 6 of urlparse
urllib.parse.urlsplit(urlurl — The URL. Leading C0 control characters and spaces are stripped (3.12+); tab, CR and LF are removed everywhere (3.10+). The scheme is lowercased.type: str | bytes · required, schemescheme — Default scheme for URLs that have none.type: str | bytes · default: ''='', allow_fragmentsallow_fragments — False keeps #... as part of the path or query.type: bool · default: True=True)
→ SplitResult

Demo

Live evaluation
The five components. Compare with urlparse: ;params stay in path.
Try:
Inputs
urlstra URL
Code
from urllib.parse import urlsplit
urlsplit('HTTPS://Example.com:8080/a/b;v=1?q=x#top')
Result
SplitResult(scheme='https', netloc='Example.com:8080', path='/a/b;v=1', query='q=x', fragment='top')

The scheme is lowercased ('HTTPS' → 'https') but the host is not — use .hostname for that. Leading spaces and control characters are stripped before parsing. Only the first # starts the fragment. On the way back, urlunsplit adds the '/' a path needs after a netloc, and an empty query or fragment simply disappears, so 'https://example.com/a?#' comes back as 'https://example.com/a'.

Parameters

NameTypeRequiredDescription
urlstr | bytesyesThe URL. Leading C0 control characters and spaces are stripped (3.12+); tab, CR and LF are removed everywhere (3.10+). The scheme is lowercased.
schemestr | bytesno ('')Default scheme for URLs that have none.
allow_fragmentsboolno (True)False keeps #... as part of the path or query.

Return value

SplitResult — A named 5-tuple (scheme, netloc, path, query, fragment) with .hostname, .port, .username, .password and .geturl(). SplitResultBytes for bytes input.

Common patterns

Strip query and fragment
A canonical form for caching or comparison.
from urllib.parse import urlsplit, urlunsplit
s = urlsplit(url)
base = urlunsplit((s.scheme, s.netloc, s.path, '', ''))
Replace the query
_replace returns a new SplitResult; geturl() calls urlunsplit.
from urllib.parse import urlsplit, urlencode
new_url = urlsplit(url)._replace(query=urlencode(params)).geturl()
Check the scheme before fetching
Reject file:, javascript: and friends.
from urllib.parse import urlsplit
if urlsplit(url).scheme not in ('http', 'https'):
    raise ValueError('only http(s) URLs are allowed')

Examples

1. Five components
from urllib.parse import urlsplit urlsplit('https://example.com/a/b;v=1?q=x#top')
Returns
SplitResult(scheme='https', netloc='example.com', path='/a/b;v=1', query='q=x', fragment='top')
2. The scheme is lowercased
from urllib.parse import urlsplit urlsplit('HTTPS://Example.COM/Path')[:2]
Returns
('https', 'Example.COM')
3. Tuple unpacking
from urllib.parse import urlsplit scheme, netloc, path, query, fragment = urlsplit('https://example.com/a?q=1') netloc
Returns
'example.com'
4. Empty ? and # are dropped
from urllib.parse import urlsplit, urlunsplit urlunsplit(urlsplit('https://example.com/a?#'))
Returns
'https://example.com/a'
5. urlunsplit adds the slash
from urllib.parse import urlunsplit urlunsplit(('https', 'example.com', 'a/b', 'x=1', ''))
Returns
'https://example.com/a/b?x=1'
6. SplitResult.geturl
from urllib.parse import SplitResult SplitResult('https', 'example.com', '/a', 'q=1', '').geturl()
Returns
'https://example.com/a?q=1'
7. _asdict()
from urllib.parse import urlsplit urlsplit('https://example.com/a')._asdict()
Returns
{'scheme': 'https', 'netloc': 'example.com', 'path': '/a', 'query': '', 'fragment': ''}

Pitfalls

1. Misspelling a field in _replace
The fields are scheme, netloc, path, query, fragment — there is no 'host'. Since 3.13 an unknown name raises TypeError.
host=
from urllib.parse import urlsplit
urlsplit('https://example.com/a')._replace(host='example.org')
TypeError: Got unexpected field names: ['host']
netloc=
from urllib.parse import urlsplit
urlsplit('https://example.com/a')._replace(netloc='example.org').geturl()
'https://example.org/a'
2. Passing four parts to urlunsplit
urlunsplit needs all five parts, fragment included — use an empty string for parts you do not have.
4 parts
from urllib.parse import urlunsplit
urlunsplit(('https', 'example.com', '/a', 'q=1'))
ValueError: not enough values to unpack (expected 6, got 5)
5 parts
from urllib.parse import urlunsplit
urlunsplit(('https', 'example.com', '/a', 'q=1', ''))
'https://example.com/a?q=1'
3. Expecting the host to be lowercased
urlsplit lowercases only the scheme. Compare hosts with .hostname, which is lowercased.
netloc
from urllib.parse import urlsplit
urlsplit('https://Example.COM/').netloc == 'example.com'
False
hostname
from urllib.parse import urlsplit
urlsplit('https://Example.COM/').hostname == 'example.com'
True

When to use

Use it
  • Any time you need the parts of a URL
  • Rebuilding a URL after changing a part (_replace + geturl, or urlunsplit)
Reach for something else
  • Legacy ;params handling → urlparse
  • Resolving a relative link against a page URL → urljoin
  • Decoding the query → parse_qs on .query

Notes

CPython impl
Strips leading C0/space and removes tab/CR/LF, takes the scheme if the text before ':' is a valid scheme starting with a letter, takes the netloc after '//' up to the first / ? or #, then splits off # and ?; results are cached (functools.lru_cache)
Errors
ValueError: 'Invalid IPv6 URL' for unbalanced brackets, '... does not appear to be an IPv4 or IPv6 address', 'An IPv4 address cannot be in brackets', and NFKC-unsafe netlocs
_replace
An unknown field name raises TypeError in 3.13 (ValueError before)

FAQ

urlsplit. The Python docs say urlsplit() should generally be used instead of urlparse(); urlparse only adds the split of ;params off the last path segment, a feature from obsolete RFCs.