Common causes
1. HTML entities in XML content
Text copied from web pages or written by people used to HTML contains named entities. Replace each with its numeric reference, or with the character itself if the file is UTF-8.
<title>Price list © 2026</title><title>Price list © 2026</title>2. Entities declared in a DTD the parser does not load
XHTML and DocBook documents declare their entities in an external DTD, which most parsers do not download for security and speed. Declare the few you need in an internal subset.
<article>
<p>Café menu</p>
</article><!DOCTYPE article [
<!ENTITY eacute "é">
]>
<article>
<p>Café menu</p>
</article>3. HTML-escaped text inserted into XML
A CMS or rich-text field often stores HTML-escaped text. When it is inserted into an RSS feed or sitemap, its entities come along. Unescape it to plain text first and let the XML serializer escape the five special characters.
<description>Fish & chips — £9</description><description>Fish & chips — £9</description>Frequently asked questions
What is the numeric reference for a non-breaking space?
Use (decimal) or (hexadecimal). Every Unicode character can be written this way in XML, without any declaration.
Why does the page display fine in a browser?
Browsers parse .html files with the forgiving HTML parser, which knows all HTML entities. Served as XML, for example as application/xhtml+xml or as an RSS feed, the same text is an error.
Is it safe to enable DTD loading to fix this?
Usually not for untrusted input: external entities enable XXE attacks and entity expansion enables “billion laughs” denial of service. Replace the entities in the content instead.