urllib.parse

Two halves in one module. Parsing: urlsplit / urlparse cut a URL into named parts, urljoin resolves links, urlunsplit rebuilds. Quoting: quote / quote_plus / urlencode percent-encode, unquote / parse_qs decode. Nothing here touches the network — that is urllib.request.

Data formatsPython 3.0+Live demo
Import
import urllib.parse
from urllib.parse import urlsplit, urljoin, urlencode, parse_qs, quote
Parsing
urlsplit, urlparse, urlunsplit, urlunparse, urljoin, urldefrag — results are named tuples with .hostname, .port, .geturl()
Quoting
quote, quote_plus, quote_from_bytes, unquote, unquote_plus, unquote_to_bytes, urlencode, parse_qs, parse_qsl
Spec
RFC 3986 with deliberate leniency; not a WHATWG URL parser and not a validator
Python 2
The functions of Python 2 urlparse plus urllib.quote / urlencode moved here in Python 3

Demo

Live evaluation
urlsplit plus the netloc helpers: scheme, host, port, path, query, fragment.
Try:
Inputs
urlstra URL
Code
from urllib.parse import urlsplit
u = urlsplit('https://Shop.Example.com:8443/cart/items?id=7&qty=2#summary')
(u.scheme, u.hostname, u.port, u.path, u.query, u.fragment)
Result
('https', 'shop.example.com', 8443, '/cart/items', 'id=7&qty=2', 'summary')

urlsplit keeps the host as written in netloc; .hostname lowercases it and drops IPv6 brackets, and .port is an int or None — or a ValueError such as 'Port out of range 0-65535'. Without '//' there is no host at all. On the quoting side, quote keeps '/' unless safe='', and quote_plus writes spaces as + — the form style that parse_qs decodes.

Members

Functions8
LIVE
urllib.parse.parse_qs
parse_qs(qs, keep_blank_values=False, strict_parsing=False, encoding='utf-8', errors='replace', max_num_fields=None, separator='&')
Parse a query string into a dict of lists (parse_qs) or a list of (name, value) pairs (parse_qsl), decoding %XX escapes and + as you go.
LIVE
urllib.parse.quote
quote(string, safe='/', encoding=None, errors=None)
Percent-encode text for a URL: quote writes a space as %20 and keeps / by default; quote_plus writes a space as + for form data; quote_from_bytes takes bytes.
LIVE
urllib.parse.unquote
unquote(string, encoding='utf-8', errors='replace')
Decode %XX escapes back into characters. unquote_plus also turns + into a space (form data); unquote_to_bytes returns the raw bytes.
LIVE
urllib.parse.urldefrag
urldefrag(url)
Split the #fragment off a URL: returns a DefragResult(url, fragment) whose geturl() puts them back together.
LIVE
urllib.parse.urlencode
urlencode(query, doseq=False, safe='', encoding=None, errors=None, quote_via=quote_plus)
Turn a dict (or a list of pairs) into a query string: every key and value goes through quote_plus and is joined with = and &.
LIVE
urllib.parse.urljoin
urljoin(base, url, allow_fragments=True)
Resolve a link against the URL of the page it appears on, like a browser does: relative paths, ../, /absolute paths, //host and full URLs.
LIVE
urllib.parse.urlparse
urlparse(url, scheme='', allow_fragments=True)
Split a URL into six parts — scheme, netloc, path, params, query, fragment — as a ParseResult with hostname, port, username and password attributes; urlunparse puts them back together.
LIVE
urllib.parse.urlsplit
urlsplit(url, scheme='', allow_fragments=True)
Split a URL into five parts — scheme, netloc, path, query, fragment — as a SplitResult; urlunsplit joins five parts back into a URL.

Common patterns

Read query parameters
Split first, then parse the query.
from urllib.parse import urlsplit, parse_qs
params = parse_qs(urlsplit(url).query)
Build a URL with a query
urlencode quotes every key and value.
from urllib.parse import urlencode
url = 'https://example.com/search?' + urlencode({'q': term, 'page': 2})
Resolve a link
Relative href → absolute URL.
from urllib.parse import urljoin
absolute = urljoin(page_url, href)
Encode one path segment
safe='' so slashes in the value are encoded.
from urllib.parse import quote
url = 'https://example.com/users/' + quote(username, safe='')

Examples

1. Split a URL
from urllib.parse import urlsplit urlsplit('https://example.com:8080/a?q=1#top')
Returns
SplitResult(scheme='https', netloc='example.com:8080', path='/a', query='q=1', fragment='top')
2. Host and port
from urllib.parse import urlsplit u = urlsplit('https://Example.com:8080/') (u.hostname, u.port)
Returns
('example.com', 8080)
3. Resolve a relative link
from urllib.parse import urljoin urljoin('https://example.com/docs/guide.html', 'faq.html')
Returns
'https://example.com/docs/faq.html'
4. Build a query string
from urllib.parse import urlencode urlencode({'q': 'rock & roll', 'page': 2})
Returns
'q=rock+%26+roll&page=2'
5. Read a query string
from urllib.parse import parse_qs parse_qs('tag=a&tag=b&page=2')
Returns
{'tag': ['a', 'b'], 'page': ['2']}
6. Percent-encode
from urllib.parse import quote quote('café menu.pdf')
Returns
'caf%C3%A9%20menu.pdf'
7. Percent-decode
from urllib.parse import unquote unquote('caf%C3%A9%20menu.pdf')
Returns
'café menu.pdf'

Pitfalls

1. quote keeps slashes by default
quote's safe parameter defaults to '/', which is wrong for a single path segment or a query value.
quote(x)
from urllib.parse import quote
quote('a/b c')
'a/b%20c'
quote(x, safe='')
from urllib.parse import quote
quote('a/b c', safe='')
'a%2Fb%20c'
2. urljoin with a base that lacks a trailing slash
The last path segment of the base is replaced, not extended.
'/v1'
from urllib.parse import urljoin
urljoin('https://example.com/api/v1', 'users')
'https://example.com/api/users'
'/v1/'
from urllib.parse import urljoin
urljoin('https://example.com/api/v1/', 'users')
'https://example.com/api/v1/users'
3. Building query strings by hand
An & or = inside a value splits it into extra parameters on the server.
f-string
from urllib.parse import parse_qs
q = 'R&D'
parse_qs(f'dept={q}&page=1')
{'dept': ['R'], 'page': ['1']}
urlencode
from urllib.parse import parse_qs, urlencode
parse_qs(urlencode({'dept': 'R&D', 'page': 1}))
{'dept': ['R&D'], 'page': ['1']}

When to use

Use it
  • Reading or changing parts of a URL
  • Turning relative links into absolute ones
  • Encoding values for URLs and decoding query strings
Reach for something else
  • Fetching URLs → urllib.request, or the requests / httpx packages
  • Validating untrusted URLs for security decisions → parse, then check scheme and hostname explicitly
  • HTML escaping → html.escape

Notes

CPython impl
Lib/urllib/parse.py, pure Python. bytes input is decoded as ASCII, parsed as str and encoded back
Leniency
Almost any string parses; only unbalanced or invalid IPv6 brackets and NFKC-unsafe netlocs raise ValueError. Since 3.10 tab, CR and LF are removed; since 3.12 leading C0 control characters and spaces are stripped
Results
ParseResult, SplitResult and DefragResult are namedtuples: index, unpack, _replace(), _asdict(), geturl(); encode() gives the *Bytes classes

FAQ

from urllib.parse import urlsplit; u = urlsplit(url) — then u.scheme, u.hostname, u.port, u.path, u.query and u.fragment. parse_qs(u.query) decodes the query parameters.