HTML: Character Sets and Encoding

, , …,
Want to master JavaScript in a week? Example based, to the point. Xah JavaScript Tutorial.

In HTML, you can declare the Character Set for the file. Here's example of setting it to be UTF-8 (Unicode):

<meta http-equiv="Content-Type" content="text/html;charset=utf-8">

If you are using HTML5, you can just say:

<meta charset="utf-8" />

If you don't understand what is Character Set and Encoding, see: UNICODE Basics: What's Character Encoding, UTF-8, and All That?.

Once you declared your character set, you can have characters from that character set in your HTML file.

UTF-8 (Unicode) contains all the world's language's characters. Here is a sample of characters from Unicode:

£ ¥ © é α ± °

For more examples, see: List of Unicode Symbols ♂ ♀ ♥ ¥ α © §.

Using Character Entity

Another way to show special characters in your file is by so-called “character entity”.

Decimal Form

For example, the bullet symbol is Unicode character number 8226. In HTML, you can write it as &#8226;.

Hexadecimal Form

The number 8226 in hexadecimal is 2022. Sometimes you only know the hexadecimal number of a character. You can write it using hexadecimal like this &#x2022;.

Named Form

For some commonly used characters, HTML provides “named entity” for them. For example, the bullet character can be written as &bull;.

For a complete list of named entities, see: HTML/XML Entities List.

HTML/HTTP's Charset is About Encoding, Not Character Set

HTTP's definition of charset (and the charset meta tag in HTML) is actually about character encoding.

Here's a excerpt from〔RFC 2616 @ http://www.w3.org/Protocols/rfc2616/rfc2616-sec3.html#sec3.4〕:

3.4 Character Sets

HTTP uses the same definition of the term “character set” as that described for MIME:

The term “character set” is used in this document to refer to a method used with one or more tables to convert a sequence of octets into a sequence of characters. Note that unconditional conversion in the other direction is not required, in that not all characters may be available in a given character set and a character set may provide more than one sequence of octets to represent a particular character. This definition is intended to allow various kinds of character encoding, from simple single-table mappings such as US-ASCII to complex table switching methods such as those that use ISO-2022's techniques. However, the definition associated with a MIME character set name MUST fully specify the mapping to be performed from octets to characters. In particular, use of external profiling information to determine the exact mapping is not permitted.

Note: This use of the term “character set” is more commonly referred to as a “character encoding.” However, since HTTP and MIME share the same registry, it is important that the terminology also be shared.

If you don't understand what is Character Set and Encoding, see: UNICODE Basics: What's Character Encoding, UTF-8, and All That?.

Reference & Notes

blog comments powered by Disqus