urllib.parse.urldefrag
One job: separate the part a browser never sends to the server. urldefrag('https://x/page#top') gives url='https://x/page' and fragment='top' — handy for de-duplicating links while crawling.
Common call
urldefrag(link).url
Returns
'https://example.com/page' — no #fragment
Replaces
link.split('#')[0]
Watch out
The url part is rebuilt with urlunparse, so an empty '?' disappears too
urllib.parse.urldefrag(urlurl — The URL. Without a #, it is returned unchanged with an empty fragment.type: str | bytes · required)
→ DefragResult
Demo
Live evaluation
The URL and its fragment, separated.
Try:
Inputs
urlstra URL
Code
from urllib.parse import urldefrag urldefrag('https://example.com/page?x=1#section-2')
Result
DefragResult(url='https://example.com/page?x=1', fragment='section-2')
Only the first # starts the fragment, so 'a#b' is one fragment. When there is a #, the URL part is rebuilt from urlparse's pieces, which drops empty delimiters: 'https://example.com/a?#top' comes back as url 'https://example.com/a', and geturl() cannot restore the lone '?'.
Parameters
| Name | Type | Required | Description |
|---|---|---|---|
| url | str | bytes | yes | The URL. Without a #, it is returned unchanged with an empty fragment. |
Return value
DefragResult — A named 2-tuple (url, fragment); fragment is '' when there is none. DefragResultBytes for bytes input.
Common patterns
De-duplicate crawled links
page#a and page#b are the same document.
from urllib.parse import urldefrag seen = {urldefrag(link).url for link in links}
Unpack both parts
DefragResult is a tuple.
from urllib.parse import urldefrag url, fragment = urldefrag(link)
Examples
1. Split off the fragment
from urllib.parse import urldefrag
urldefrag('https://example.com/page?x=1#section-2')
Returns
DefragResult(url='https://example.com/page?x=1', fragment='section-2')2. No fragment
from urllib.parse import urldefrag
urldefrag('https://example.com/page')
Returns
DefragResult(url='https://example.com/page', fragment='')3. Tuple unpacking
from urllib.parse import urldefrag
url, frag = urldefrag('https://example.com/page#top')
(url, frag)
Returns
('https://example.com/page', 'top')4. geturl() rebuilds it
from urllib.parse import urldefrag
urldefrag('https://example.com/page#top').geturl()
Returns
'https://example.com/page#top'5. DefragResult directly
from urllib.parse import DefragResult
DefragResult('https://example.com/a', 'top').geturl()
Returns
'https://example.com/a#top'6. bytes input
from urllib.parse import urldefrag
urldefrag(b'https://example.com/#top')
Returns
DefragResultBytes(url=b'https://example.com/', fragment=b'top')Pitfalls
1. split('#') on a URL without a fragment
Indexing [1] fails when there is no #; urldefrag always returns two parts.
split('#')[1]
'https://example.com/page'.split('#')[1]
IndexError: list index out of range
urldefrag
from urllib.parse import urldefrag urldefrag('https://example.com/page').fragment
''
2. Expecting the URL part byte for byte
When a fragment is present, the URL part is re-assembled, which drops an empty ? (an equivalent URL, but not the same string).
compare strings
from urllib.parse import urldefrag urldefrag('https://example.com/a?#top').url == 'https://example.com/a?'
False
compare parsed
from urllib.parse import urldefrag, urlsplit urlsplit(urldefrag('https://example.com/a?#top').url) == urlsplit('https://example.com/a?')
True
When to use
Use it
- Removing #anchors before comparing, caching or fetching URLs
- Reading the fragment of a link
Reach for something else
- Other URL parts → urlsplit
- Changing the fragment → urlsplit(url)._replace(fragment=...).geturl()
Notes
CPython impl
If '#' is in the URL: urlparse it and urlunparse everything but the fragment; else return the URL unchanged with fragment ''
DefragResult
A namedtuple (url, fragment) since 3.2; geturl() is url + "#" + fragment, or just url when the fragment is empty
FAQ
urldefrag(url).url. It returns the URL without the fragment; .fragment holds the text after the first '#'.