UnicodeError
The class to catch when any text/bytes conversion fails — you almost never see it raised bare, but except UnicodeError handles decode, encode and translate errors alike.
Demo
'hello'.encode('utf-8').decode('ascii')
In Trigger, the position is the first non-ASCII BYTE: 'año' fails at 1 and 'naïve' at 2, both on 0xc3 — the first of the two bytes UTF-8 uses for ñ and ï. In Handle, the rocket in 'go 🚀' is four bytes, but ASCII decoding rejects bytes one at a time, so start/end cover only the first one. In Raise, even a plain '-' prints as '\x2d': the translate message always escapes the character.
Constructor
| Name | Type | Required | Description |
|---|---|---|---|
| *args | object | no | A bare UnicodeError takes any arguments, like Exception. The subclasses require theirs: (encoding, object, start, end, reason) for decode/encode, (object, start, end, reason) for UnicodeTranslateError. |
Attributes
| Attribute | Type | Meaning |
|---|---|---|
| args | tuple | Constructor arguments. A bare UnicodeError has only this — the attributes below exist on the three subclasses. |
| encoding | str | None | Codec name on decode/encode errors; always None on UnicodeTranslateError. |
| object | bytes | str | The input being processed: bytes for UnicodeDecodeError, str for encode and translate errors. |
| start | int | First bad index into object — a byte offset for decode, a character index otherwise. |
| end | int | Index just past the bad range. |
| reason | str | Codec-specific explanation, e.g. 'ordinal not in range(128)'. |
Common patterns
def convert(data, src='utf-8', dst='latin-1'): try: return data.decode(src).encode(dst) except UnicodeError as e: raise ValueError(f'cannot convert: {e}') from e
import codecs def dash(e): if isinstance(e, UnicodeEncodeError): return ('-' * (e.end - e.start), e.end) raise e codecs.register_error('dash', dash) safe = 'naïve'.encode('ascii', 'dash')
try: value = int(raw.decode('utf-8')) except UnicodeError: value = None # undecodable bytes except ValueError: value = 0 # decodable, but not a number
Examples
Pitfalls
try: b'\xff'.decode('utf-8') except ValueError: kind = 'bad number' except UnicodeError: kind = 'bad text' kind
try: b'\xff'.decode('utf-8') except UnicodeError: kind = 'bad text' except ValueError: kind = 'bad number' kind
data = 'día 1'.encode('utf-8') + b'\xff' try: data.decode('utf-8') except UnicodeError as e: good = data.decode('utf-8', 'replace')[:e.start] good
data = 'día 1'.encode('utf-8') + b'\xff' try: data.decode('utf-8') except UnicodeError as e: good = data[:e.start].decode('utf-8') good
When to use
- Catching any encode/decode failure with one except clause
- Type-checking the instance passed to a custom codecs error handler
- UnicodeTranslateError: raising it from a custom translation codec
- Raising a bare UnicodeError — raise the specific subclass with its five (or four) arguments
- Catching it when only one direction can fail — name UnicodeDecodeError or UnicodeEncodeError directly
- str.translate() problems — it never raises UnicodeTranslateError; unmapped characters are kept as-is
Notes
FAQ
UnicodeError is the common base class. UnicodeDecodeError is raised when bytes cannot be turned into text (bytes.decode, reading a file), UnicodeEncodeError when text cannot be turned into bytes (str.encode, writing or printing). In real code you almost always see one of the subclasses; catch UnicodeError when you want to handle both.