About the HTML Entity Encoder / Decoder
This HTML entity encoder escapes text so it can be shown safely inside a web page, and decodes entities back into readable characters. Encoding turns the five characters that have meaning in HTML — < > & " and ' — into < > & " and ', so code samples display as text instead of being parsed as tags. You can also convert every non-ASCII character (é, €, —, emoji) into named or numeric entities for systems that only accept plain ASCII.
Developers use it to show code snippets in blog posts and documentation, to escape user content in templates, to fix mangled text in CMS exports and emails, and to read entity-filled HTML scraped from other sites. Decoding understands named entities such as © and —, decimal references like € and hex references like 😀.
The decoder knows the 230-plus most common named entities (Latin-1, Greek, typographic, math and symbol names) plus every decimal and hex reference; an unrecognised name is left unchanged. It expects the trailing semicolon on named entities, so legacy forms such as & without “;” stay as typed, and numeric references €–Ÿ are decoded as their literal Unicode code points rather than remapped to Windows-1252 characters the way browsers do. Escaping text is only part of preventing cross-site scripting: use your framework’s automatic escaping, and encode differently for URLs, JavaScript and CSS contexts.
How to use the html entity encoder / decoder
- 1Choose Encode to escape text or Decode to turn entities back into characters.
- 2Paste your text, HTML or code snippet.
- 3For encoding, choose whether to convert only special characters or all non-ASCII characters too.
- 4Copy the result from the output box.
Formula and method
Encoding replaces each character that HTML would otherwise interpret with a character reference. The ampersand is always encoded (because it starts every entity), angle brackets are encoded so they cannot open tags, and both quote characters are encoded so the text is also safe inside attribute values.
Optionally, every character above ASCII is written as a named entity when one exists (© → ©) or as a numeric reference to its Unicode code point: decimal © or hexadecimal ©. Characters outside the Basic Multilingual Plane, such as emoji, become a single reference to their full code point. Decoding reverses the process; invalid code points become the replacement character U+FFFD.
- &name;
- Named character reference, e.g. © = ©
- &#n;
- Decimal code point reference, e.g. € = €
- &#xh;
- Hexadecimal code point reference, e.g. € = €
Worked examples
Escape an HTML snippet for display
The two < and two > characters, two double quotes and the ampersand are replaced — 7 entities in all — so the browser shows the tag as text. ©, — and £ stay as UTF-8 characters.
ASCII-only output with named entities
On top of the 7 special characters, ©, — and £ become ©, — and £, giving 10 entities and output that survives any ASCII-only system.
Emoji and accents as hex references
é is code point U+00E9 and 😀 is U+1F600, so they become é and 😀. The emoji is one reference even though JavaScript stores it as two UTF-16 code units.
Decode mixed entities
Named (& © ♥), decimal (– = en dash) and hex (™ = ™) references are all decoded — 6 in total.
Frequently asked questions
Which characters must be escaped in HTML?+
Always escape & and < in text content. Inside attribute values also escape the quote character used to delimit the attribute. Escaping > and both quote types everywhere is the simplest safe habit.
What is the difference between &nbsp; and a normal space?+
is a non-breaking space (U+00A0): browsers will not wrap a line at it and will not collapse several of them into one. Use it sparingly, e.g. between a number and its unit.
Should I use named or numeric entities?+
Both work in every browser. Named entities like © are easier to read; numeric references work for every Unicode character, including emoji, which have no names. On UTF-8 pages you rarely need either except for < > & and quotes.
Does HTML encoding prevent XSS?+
Encoding user input before putting it into HTML text or quoted attributes blocks most injection, but it is not enough inside script blocks, URLs, CSS or unquoted attributes. Use your framework’s context-aware escaping.
Why do I see &amp; in my page?+
The text was encoded twice: & became &, then the & of that entity was encoded again. Decode once to fix it, and make sure only one layer of your code escapes the output.