urllib.parse.urljoin

urljoin is link resolution (RFC 3986), not string concatenation. The last path segment of base is a file name unless base ends with '/': urljoin('https://x/api/v1', 'users') drops v1. An absolute path or a full URL replaces the base path altogether.

urllib.parse functionPython 3.0+Live demo
Common call
urljoin('https://example.com/docs/', 'intro')
Returns
'https://example.com/docs/intro'
Replaces
base + '/' + path — and its doubled or missing slashes
Watch out
Without a trailing '/', the base's last segment is replaced
urllib.parse.urljoin(basebase — The URL of the document containing the link. If empty, url is returned unchanged.type: str | bytes · required, urlurl — The link: relative path, ../path, /path, ?query, #fragment, //host/path or a full URL. If empty, base is returned unchanged.type: str | bytes · required, allow_fragmentsallow_fragments — Passed to urlparse for both URLs: False keeps #... inside the path or query.type: bool · default: True=True)
→ str | bytes

Demo

Live evaluation
Resolve a link against a base, the way a browser resolves an <a href>.
Try:
Inputs
basestrthe page URL
urlstrthe link
Code
from urllib.parse import urljoin
urljoin('https://example.com/docs/guide.html', 'intro.html')
Result
'https://example.com/docs/intro.html'

Relative links are resolved against the base's directory: everything after the last '/' of the base path is dropped first. That is why 'https://example.com/api/v1' + 'users' gives /api/users, and why a link starting with '/' goes back to the host root. Extra ../ segments cannot climb above the root — they are ignored (RFC 3986, Python 3.5+).

Parameters

NameTypeRequiredDescription
basestr | bytesyesThe URL of the document containing the link. If empty, url is returned unchanged.
urlstr | bytesyesThe link: relative path, ../path, /path, ?query, #fragment, //host/path or a full URL. If empty, base is returned unchanged.
allow_fragmentsboolno (True)Passed to urlparse for both URLs: False keeps #... inside the path or query.

Return value

str | bytes — The absolute URL that url refers to when found on the page at base.

Common patterns

Absolute links from scraped HTML
Resolve every href against the URL the page was fetched from.
from urllib.parse import urljoin
links = [urljoin(page_url, a['href']) for a in anchors]
An API base URL
Keep the base ending in '/' and the endpoint without a leading '/'.
from urllib.parse import urljoin
API = 'https://api.example.com/v2/'
urljoin(API, 'users/42')
Follow a redirect
A Location header may be relative to the request URL.
from urllib.parse import urljoin
next_url = urljoin(request_url, response.headers['Location'])

Examples

1. Base without a trailing slash
from urllib.parse import urljoin urljoin('https://example.com/docs/guide', 'intro')
Returns
'https://example.com/docs/intro'
2. Base with a trailing slash
from urllib.parse import urljoin urljoin('https://example.com/docs/guide/', 'intro')
Returns
'https://example.com/docs/guide/intro'
3. A leading slash resets the path
from urllib.parse import urljoin urljoin('https://example.com/docs/guide/', '/intro')
Returns
'https://example.com/intro'
4. ../ goes up one level
from urllib.parse import urljoin urljoin('https://example.com/docs/guide/', '../api')
Returns
'https://example.com/docs/api'
5. //host keeps only the scheme
from urllib.parse import urljoin urljoin('https://example.com/docs/', '//cdn.example.net/x.js')
Returns
'https://cdn.example.net/x.js'
6. A query replaces the query
from urllib.parse import urljoin urljoin('https://example.com/a/b?x=1', '?y=2')
Returns
'https://example.com/a/b?y=2'
7. A fragment keeps the query
from urllib.parse import urljoin urljoin('https://example.com/a/b?x=1', '#frag')
Returns
'https://example.com/a/b?x=1#frag'
8. str and bytes do not mix
from urllib.parse import urljoin urljoin(b'https://example.com/a/', 'b')
Returns
TypeError: Cannot mix str and non-str arguments

Pitfalls

1. An API base without a trailing slash
The last segment of the base is treated as a file name and replaced — v1 disappears from the URL.
'.../v1'
from urllib.parse import urljoin
urljoin('https://example.com/api/v1', 'users')
'https://example.com/api/users'
'.../v1/'
from urllib.parse import urljoin
urljoin('https://example.com/api/v1/', 'users')
'https://example.com/api/v1/users'
2. A leading slash on the endpoint
A path that starts with / is absolute: it replaces the whole base path, trailing slash or not.
'/users'
from urllib.parse import urljoin
urljoin('https://example.com/api/v1/', '/users')
'https://example.com/users'
'users'
from urllib.parse import urljoin
urljoin('https://example.com/api/v1/', 'users')
'https://example.com/api/v1/users'
3. Trusting urljoin to keep you on your host
A user-supplied link that is a full URL or starts with // replaces the host. Check the result before fetching or redirecting.
user input
from urllib.parse import urljoin
urljoin('https://example.com/app/', '//evil.example/x')
'https://evil.example/x'
check netloc
from urllib.parse import urljoin, urlsplit
target = urljoin('https://example.com/app/', '//evil.example/x')
urlsplit(target).netloc == 'example.com'
False

When to use

Use it
  • Turning links found in a page into absolute URLs
  • Endpoint paths under an API base URL (base ending in /)
  • Relative redirects (Location headers)
Reach for something else
  • Joining file-system paths → os.path.join or pathlib
  • Appending a query string → urlencode + _replace(query=...)
  • Keeping user input on your own host without checking the result

Notes

CPython impl
urljoin parses both URLs with urlparse, returns url unchanged if its scheme differs or is not in uses_relative, takes a netloc from url if present, and otherwise merges the paths and removes . and .. segments
RFC 3986
Behaviour updated to match RFC 3986 in Python 3.5; surplus ../ segments are dropped
Schemes
Only schemes in urllib.parse.uses_relative (http, https, ftp, file, ws, wss, …) are resolved; for others such as mailto: the link is returned as is

FAQ

Because without a trailing '/', the last segment is a file name: urljoin('https://x.com/api/v1', 'users') resolves 'users' next to 'v1' and gives 'https://x.com/api/users'. End the base with '/' to keep it.