Unicode

You are not logged in.

Please Log In for full access to the web site.
Note that this link will take you to an external site (https://shimmer.mit.edu) to authenticate, and then you will be redirected back to this page.

The questions on this page are completely optional. They are for enrichment purposes only and do not have grades associated with them. You are not expected to know the contents of this exercise for the exam.

Back to Exercises

In this exercise set, we are going to explore how UTF-8 works.

The char data type in C can only store at most one byte (8 bits) of information, which limits the number of possible characters to just 256. In order to represent glyphs far beyond ASCII, including emojis, we need to use something bigger.

One idea is to use int to represent each character instead. This would allow a maximum of 2^{32} symbols. However, this method is very inefficient as all characters necessarily take 4 bytes each, especially as some symbols are almost never used.

Ideally, more common characters should take up less space. UTF-8 accomplishes this by assigning each symbol a "code point," which is an abstract numerical value. Each code point can be translated to a concrete binary encoding using the rules described below. The encoding scheme is designed so lower numerical values map to encoding with smaller number of bytes. UTF-8 uses between one to four bytes for each symbol.

UTF-8 Encoding

A code point is usually given in form of U+xxxx where xxxx is an unsigned hexadecimal number, written with at least four digits. Valid code points range from U+0000 to U+10FFFF. To translate a code point to a binary encoding, select the appropriate row in the following table and write the bits of the code point in each x position starting from the least significant digit for the rightmost bit to the most significant digit for the leftmost bit. Zero padding may be needed. We use unsigned binary representation.

From To Byte 1 Byte 2 Byte 3 Byte 4
U+0000 U+007F 0xxx xxxx
U+0080 0+07FF 110x xxxx 10xx xxxx
U+0800 U+FFFF 1110 xxxx 10xx xxxx 10xx xxxx
U+10000 U+10FFFF 1111 0xxx 10xx xxxx 10xx xxxx 10xx xxxx

For example, to encode the euro sign (€) which has code point U+20AC: First, note that the code point lies between U+0800 and U+FFFF, which means the encoding will have three bytes. Hexadecimal 20AC in binary is 0010_0000_1010_1100. Substituting the bits into 1110_xxxx__10xx_xxxx__10xx_xxxx gives 1110_0010__1000_0010__1010_1100 (0xE282AC in hexadecimal). You can read more examples here.

Note that some symbols are actually encoded using more than four bytes by combining multiple "characters" together. For example, 🇺🇸 (the United States flag) is actually two characters: U+1F1FA and U+1F1F8.

Exercise 1

Encode the character ʤ (Dezh Digraph) which has code point U+02A4 using UTF-8 encoding scheme.

(For questions with binary number input: you can write _ to separate the bits and make your answer more readable. This does not affect the correctness.)

What is 0x02A4 in unsigned binary? Write using 16 bits.

How many bytes does the UTF-8 encoding of U+02A4 require?

What is the resulting binary encoding? Write with 16 bits.

What is the resulting encoding in hexadecimal?

Exercise 2

The following sequence of bytes is a UTF-8 encoding of a character: 0xF0 0x9F 0x98 0x82 (11110000_10011111_10011000_10000010).

What is the code point of this character in binary? Write using the minimum number of bits.

What is the code point in hexadecimal? U+

(Fun fact: This is "Face with Tears of Joy," the most-used emoji in 2021.)

String Length

Consider the string 你好 (U+4F60, U+597D, U+0000 where U+0000 is the null terminator) which has UTF-8 encoding in hexadecimal: 0xE4BDA0E5A5BD00.

Try running the following program in your own terminal.

#include <string.h>

char str[] = {0xE4, 0xBD, 0xA0, 0xE5, 0xA5, 0xBD, 0x00};
// can also be equivalently written as one of the following
// char str[] = "\xE4\xBD\xA0\xE5\xA5\xBD";
// char str[] = "\u4F60\u597D";
// char str[] = "你好";

int main(void) {
  printf("%d", strlen(str)); // outputs: 6
  return 0;
}

Notice that C believes the string str has length 6, because there are exactly 6 bytes starting from str to the first null terminator and C does not recognize Unicode by default. We want the answer to be 2.

Your task is to write the function utf8_strlen that takes in a string (a pointer to char array) and returns the number of characters according to UTF-8 encoding described above.

You may assume the string contains valid UTF-8 characters and is terminated with null character. You are not expected to handle compound characters like 🇺🇸 specially.

Exercise 3

Complete the code for the function utf8_strlen.

Hint: You can figure out how many bytes a character takes by looking at the top bits of the first byte. Use bitmasks.

Back to Exercises