Garbled characters in a CSV opened in Excel (é instead of é)

The file is UTF-8, where é is stored as two bytes, but Excel opened it as Windows-1252, where those two bytes are the characters à and ©. Nothing is wrong with the data or with CSV syntax, so this sample parses without errors; the text has simply been decoded with the wrong character set (mojibake). The reverse mistake, reading a Windows-1252 file as UTF-8, is what produces Python’s UnicodeDecodeError.

Seen as:

  • José, Zürich, café, piñata
  • ’ “ – (instead of ’ “ –)
  • UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 in position 3: invalid continuation byte
  • UnicodeDecodeError: 'charmap' codec can't decode byte 0x9d in position 2187: character maps to <undefined>

Input

Settings

History

Load from URL

Common causes

1. A UTF-8 CSV without a byte-order mark opened in Excel

Excel on Windows assumes the system code page unless the file starts with a UTF-8 BOM. Export with a BOM (Python: encoding="utf-8-sig"; Excel: “CSV UTF-8”), or import through Data > From Text/CSV and pick 65001: Unicode (UTF-8).

Before
df.to_csv("customers.csv", index=False)
After
df.to_csv("customers.csv", index=False, encoding="utf-8-sig")

2. Text that was already double-encoded

If a pipeline read UTF-8 as Latin-1 and saved the result as UTF-8, the mojibake is now stored in the file. Reverse the mistake once: encode as Latin-1 (or cp1252) and decode as UTF-8.

Before
name,city
José,Zürich
After
name,city
José,Zürich

3. Curly quotes and dashes from Word

Typographic characters are three bytes in UTF-8, so a misread ’ becomes ’ and “ becomes “. The fix is the same as for accented letters: read the file as UTF-8.

Before
id,note
1,Customer’s order
After
id,note
1,Customer’s order

4. A Windows-1252 file read as UTF-8

Excel’s plain “CSV (Comma delimited)” format writes Windows-1252 on Windows. Reading it as UTF-8 fails on the first accented character. Tell the reader the real encoding.

Before
df = pd.read_csv("export.csv")
After
df = pd.read_csv("export.csv", encoding="cp1252")

Frequently asked questions

Why does the file look fine in a text editor but not in Excel?

Modern editors detect UTF-8 automatically. Excel on Windows does not unless the file starts with a byte-order mark, so the same bytes are shown differently.

How do I fix mojibake that is already in my data?

In Python, text.encode(“cp1252”).decode(“utf-8”) reverses one round of the mistake. The ftfy library detects and repairs mixed or repeated mojibake automatically.

Does adding a BOM break other tools?

Some strict parsers treat the BOM as part of the first column name. Most CSV readers strip it, and PasteKit does too. See the CSV and Excel encoding guide for a full comparison.

Related