UnicodeError

The class to catch when any text/bytes conversion fails — you almost never see it raised bare, but except UnicodeError handles decode, encode and translate errors alike.

InheritsBaseException›Exception›ValueError›UnicodeError
Type / value exceptionPython 3 (all)Live demo
UnicodeError(*args)
Raised by
its subclasses: UnicodeDecodeError, UnicodeEncodeError, UnicodeTranslateError
Message
codec name, position and reason: 'ascii' codec can't decode byte 0xc3 in position 2
Quick fix
except UnicodeError catches all three
Watch out
decode positions are bytes, encode/translate positions are characters

Demo

Live evaluation
Encode text to UTF-8 bytes, then decode those bytes as ASCII. The failing position is a byte offset — count the bytes, not the letters.
Try:
Inputs
textstrtext to round-trip
Code
'hello'.encode('utf-8').decode('ascii')
Result
'hello'

In Trigger, the position is the first non-ASCII BYTE: 'año' fails at 1 and 'naïve' at 2, both on 0xc3 — the first of the two bytes UTF-8 uses for ñ and ï. In Handle, the rocket in 'go 🚀' is four bytes, but ASCII decoding rejects bytes one at a time, so start/end cover only the first one. In Raise, even a plain '-' prints as '\x2d': the translate message always escapes the character.

Constructor

NameTypeRequiredDescription
*argsobjectnoA bare UnicodeError takes any arguments, like Exception. The subclasses require theirs: (encoding, object, start, end, reason) for decode/encode, (object, start, end, reason) for UnicodeTranslateError.

Attributes

AttributeTypeMeaning
argstupleConstructor arguments. A bare UnicodeError has only this — the attributes below exist on the three subclasses.
encodingstr | NoneCodec name on decode/encode errors; always None on UnicodeTranslateError.
objectbytes | strThe input being processed: bytes for UnicodeDecodeError, str for encode and translate errors.
startintFirst bad index into object — a byte offset for decode, a character index otherwise.
endintIndex just past the bad range.
reasonstrCodec-specific explanation, e.g. 'ordinal not in range(128)'.

Common patterns

One handler for both directions
Code that decodes input and encodes output can treat every conversion failure the same way.
def convert(data, src='utf-8', dst='latin-1'):
    try:
        return data.decode(src).encode(dst)
    except UnicodeError as e:
        raise ValueError(f'cannot convert: {e}') from e
Custom error handler
codecs.register_error() receives the UnicodeError subclass instance and returns (replacement, resume_position). Check the type to support the directions you want.
import codecs

def dash(e):
    if isinstance(e, UnicodeEncodeError):
        return ('-' * (e.end - e.start), e.end)
    raise e

codecs.register_error('dash', dash)
safe = 'naïve'.encode('ascii', 'dash')
Specific before general
UnicodeError is a ValueError, so order the except clauses from narrow to broad.
try:
    value = int(raw.decode('utf-8'))
except UnicodeError:
    value = None  # undecodable bytes
except ValueError:
    value = 0     # decodable, but not a number

Examples

1. Catches encode errors
try: 'é'.encode('ascii') except UnicodeError as e: kind = type(e).__name__ kind
Returns
'UnicodeEncodeError'
2. Catches decode errors
try: b'\xff'.decode('utf-8') except UnicodeError as e: kind = type(e).__name__ kind
Returns
'UnicodeDecodeError'
3. It is a ValueError
issubclass(UnicodeError, ValueError)
Returns
True
4. Bare UnicodeError has no start
UnicodeError('bad text').start
Returns
AttributeError: 'UnicodeError' object has no attribute 'start'
5. UnicodeTranslateError message
str(UnicodeTranslateError('abc', 1, 2, 'bad'))
Returns
"can't translate character '\\x62' in position 1: bad"
6. Translate error has no encoding
e = UnicodeTranslateError('abc', 0, 3, 'bad') (e.encoding, e.object, str(e))
Returns
(None, 'abc', "can't translate characters in position 0-2: bad")
7. Error handlers accept it
import codecs codecs.replace_errors(UnicodeTranslateError('abc', 1, 2, 'bad'))
Returns
('�', 2)

Pitfalls

1. except ValueError first swallows UnicodeError
Clauses are tried in order and UnicodeError is a ValueError subclass, so a ValueError clause placed first catches encoding problems too.
General first
try:
    b'\xff'.decode('utf-8')
except ValueError:
    kind = 'bad number'
except UnicodeError:
    kind = 'bad text'
kind
'bad number'
Specific first
try:
    b'\xff'.decode('utf-8')
except UnicodeError:
    kind = 'bad text'
except ValueError:
    kind = 'bad number'
kind
'bad text'
2. Slicing text with a decode position
A decode error's start is an offset into the bytes (e.object). Using it on the decoded text misses whenever earlier characters took more than one byte.
text[:e.start]
data = 'día 1'.encode('utf-8') + b'\xff'
try:
    data.decode('utf-8')
except UnicodeError as e:
    good = data.decode('utf-8', 'replace')[:e.start]
good
'día 1�'
data[:e.start]
data = 'día 1'.encode('utf-8') + b'\xff'
try:
    data.decode('utf-8')
except UnicodeError as e:
    good = data[:e.start].decode('utf-8')
good
'día 1'

When to use

Use it
  • Catching any encode/decode failure with one except clause
  • Type-checking the instance passed to a custom codecs error handler
  • UnicodeTranslateError: raising it from a custom translation codec
Reach for something else
  • Raising a bare UnicodeError — raise the specific subclass with its five (or four) arguments
  • Catching it when only one direction can fail — name UnicodeDecodeError or UnicodeEncodeError directly
  • str.translate() problems — it never raises UnicodeTranslateError; unmapped characters are kept as-is

Notes

CPython impl
Objects/exceptions.c — UnicodeError is a plain ValueError subclass; the three subclasses add encoding/object/start/end/reason and their own __str__
Subclasses
UnicodeDecodeError (bytes → str), UnicodeEncodeError (str → bytes), UnicodeTranslateError (str → str)
Catch via
except UnicodeError, or except ValueError for the broader family
Translate error
Handlers like replace and backslashreplace accept it; xmlcharrefreplace raises TypeError: don't know how to handle UnicodeTranslateError in error callback

FAQ

UnicodeError is the common base class. UnicodeDecodeError is raised when bytes cannot be turned into text (bytes.decode, reading a file), UnicodeEncodeError when text cannot be turned into bytes (str.encode, writing or printing). In real code you almost always see one of the subclasses; catch UnicodeError when you want to handle both.