The five characters that actually break HTML#
&, <, >, " and ' are the characters HTML parses as syntax rather than content — < starts a tag, & starts an entity, and the quote characters close an attribute value early. Any of these appearing unescaped inside text that gets inserted into a page — a comment, a username, a search query echoed back — is how basic HTML injection happens. Basic mode escapes exactly these five and nothing else, which is what most real security-relevant escaping actually needs.
Why "all" mode exists beyond the basic five#
Encoding every non-ASCII character as an entity is not a security requirement in modern UTF-8 documents — it is a compatibility one. It guarantees the output is readable in any document, regardless of its declared character encoding, and it is still common in generated feeds, older CMS export formats, and email templates that cannot always guarantee UTF-8 will survive the whole pipeline intact. Where a character has a recognized name (é → é), all mode uses the name; where it does not, it falls back to a numeric entity (日) rather than leaving the raw character in place.
Named entities versus numeric entities when decoding#
A named entity like & or é is easier for a human to read in source; a numeric entity like & or & encodes the exact Unicode code point directly, in decimal or hexadecimal. Both decode to the identical character — this tool recognizes both forms, plus the small set of named entities that account for the overwhelming majority of real content: the five basic ones, common punctuation like em dashes and curly quotes, currency symbols, and the accented Latin-1 letters.