urllib.parse.urljoin
urljoin is link resolution (RFC 3986), not string concatenation. The last path segment of base is a file name unless base ends with '/': urljoin('https://x/api/v1', 'users') drops v1. An absolute path or a full URL replaces the base path altogether.
Demo
from urllib.parse import urljoin urljoin('https://example.com/docs/guide.html', 'intro.html')
Relative links are resolved against the base's directory: everything after the last '/' of the base path is dropped first. That is why 'https://example.com/api/v1' + 'users' gives /api/users, and why a link starting with '/' goes back to the host root. Extra ../ segments cannot climb above the root — they are ignored (RFC 3986, Python 3.5+).
Parameters
| Name | Type | Required | Description |
|---|---|---|---|
| base | str | bytes | yes | The URL of the document containing the link. If empty, url is returned unchanged. |
| url | str | bytes | yes | The link: relative path, ../path, /path, ?query, #fragment, //host/path or a full URL. If empty, base is returned unchanged. |
| allow_fragments | bool | no (True) | Passed to urlparse for both URLs: False keeps #... inside the path or query. |
Return value
str | bytes — The absolute URL that url refers to when found on the page at base.
Common patterns
from urllib.parse import urljoin links = [urljoin(page_url, a['href']) for a in anchors]
from urllib.parse import urljoin API = 'https://api.example.com/v2/' urljoin(API, 'users/42')
from urllib.parse import urljoin next_url = urljoin(request_url, response.headers['Location'])
Examples
Pitfalls
from urllib.parse import urljoin urljoin('https://example.com/api/v1', 'users')
from urllib.parse import urljoin urljoin('https://example.com/api/v1/', 'users')
from urllib.parse import urljoin urljoin('https://example.com/api/v1/', '/users')
from urllib.parse import urljoin urljoin('https://example.com/api/v1/', 'users')
from urllib.parse import urljoin urljoin('https://example.com/app/', '//evil.example/x')
from urllib.parse import urljoin, urlsplit target = urljoin('https://example.com/app/', '//evil.example/x') urlsplit(target).netloc == 'example.com'
When to use
- Turning links found in a page into absolute URLs
- Endpoint paths under an API base URL (base ending in /)
- Relative redirects (Location headers)
- Joining file-system paths → os.path.join or pathlib
- Appending a query string → urlencode + _replace(query=...)
- Keeping user input on your own host without checking the result
Notes
FAQ
Because without a trailing '/', the last segment is a file name: urljoin('https://x.com/api/v1', 'users') resolves 'users' next to 'v1' and gives 'https://x.com/api/users'. End the base with '/' to keep it.