Skip to content

HTML Entity Encoder & Decoder

Encode text into HTML entities — the basic five that make markup safe, or every non-ASCII character too — or decode named and numeric entities (&, é, —) back into readable text.

  • Basic mode: just the 5 HTML-unsafe characters
  • All mode: every non-ASCII character too
  • Decodes named and numeric entities
  • Handles decimal and hexadecimal numeric entities
  • Runs fully client-side

Converter

Converted locally, nothing uploaded

Output

How to encode or decode HTML entities

  1. 01

    Pick encode or decode

    Encode escapes readable text into entities; decode reverses it.

  2. 02

    Choose a mode when encoding

    Basic escapes only the five characters that break HTML markup. All also converts accented letters and symbols to named or numeric entities.

  3. 03

    Copy the result

    The output updates as you type and can be copied in one click.

The five characters that actually break HTML#

&, <, >, " and ' are the characters HTML parses as syntax rather than content — < starts a tag, & starts an entity, and the quote characters close an attribute value early. Any of these appearing unescaped inside text that gets inserted into a page — a comment, a username, a search query echoed back — is how basic HTML injection happens. Basic mode escapes exactly these five and nothing else, which is what most real security-relevant escaping actually needs.

Why "all" mode exists beyond the basic five#

Encoding every non-ASCII character as an entity is not a security requirement in modern UTF-8 documents — it is a compatibility one. It guarantees the output is readable in any document, regardless of its declared character encoding, and it is still common in generated feeds, older CMS export formats, and email templates that cannot always guarantee UTF-8 will survive the whole pipeline intact. Where a character has a recognized name (é → &eacute;), all mode uses the name; where it does not, it falls back to a numeric entity (&#26085;) rather than leaving the raw character in place.

Named entities versus numeric entities when decoding#

A named entity like &amp; or &eacute; is easier for a human to read in source; a numeric entity like &#38; or &#x26; encodes the exact Unicode code point directly, in decimal or hexadecimal. Both decode to the identical character — this tool recognizes both forms, plus the small set of named entities that account for the overwhelming majority of real content: the five basic ones, common punctuation like em dashes and curly quotes, currency symbols, and the accented Latin-1 letters.

Frequently asked questions

What's the difference between basic and all encoding mode?

Basic mode escapes only the five characters that HTML parses as syntax: & < > " '. All mode does that plus every other non-ASCII character, using a named entity where one exists and a numeric entity otherwise — useful for maximum compatibility across encodings, not required for security.

Does this decode numeric entities like &#233; or &#x2014;?

Yes — both decimal (&#233;) and hexadecimal (&#x2014;) numeric entities are decoded, alongside named entities like &amp; and &eacute;.

Is encoding the basic five characters enough to prevent HTML injection?

For inserting plain text into HTML content, yes — those five are what HTML actually parses as markup. Safely inserting into an HTML attribute, a script context, or a URL has additional rules beyond simple entity encoding, which this tool does not attempt to cover.

Is my text sent anywhere?

No. Encoding and decoding happen entirely in your browser — nothing is uploaded.

Developers

UUID Generator

Generate cryptographically random UUID v4 or time-ordered UUID v7, in bulk.

Developers

Base64 Encoder & Decoder

Encode and decode Base64 with correct UTF-8 handling, including the URL-safe alphabet.

Developers

ULID Generator

Generate ULIDs — sortable by creation time like UUID v7, but Crockford Base32 instead of hex.