Unicode Escape and Unescape

Turn accents, symbols and emoji into \uXXXX escapes that survive any ASCII-only system, or make an escaped string from a log or a JSON file readable again.

Escapes accents, symbols and emoji as \uXXXX and leaves plain ASCII readable. Valid in JSON, JavaScript, Java and C#.

42 characters, 47 bytes
63 characters

Unicode escapes examples

TextEncoded
café

caf\u00e9

Non-ASCII: Only the é is escaped

£5 → €6

\u00a35 \u2192 \u20ac6

Non-ASCII: ASCII stays readable

😀

\ud83d\ude00

Non-ASCII: Beyond U+FFFF, so a surrogate pair

😀

\u{1f600}

\u{…}: The ES2015 form holds it in one escape

Hi

\u0048\u0069

Everything: ASCII is escaped too

Encoding guides

How each one works, with worked examples.

Everything happens in your browser. What you type or paste is never uploaded, stored or added to the share link, so it is safe to use with tokens, keys and customer data.

What a Unicode escape is

A Unicode escape writes a character as a backslash, a u and the four hex digits of its code: é is \u00e9. The same syntax is understood by JSON, JavaScript, Java, C#, Python and many configuration formats, which makes it the standard way to carry non-ASCII text through something that only handles ASCII reliably: old source files, .properties files, logs and some APIs.

Four hex digits only reach U+FFFF. Characters beyond that, which includes all emoji, are written as two escapes called a surrogate pair: 😀 is \ud83d\ude00. This is the form JSON requires. JavaScript since ES2015, along with Rust, Swift and others, also has a code point form in braces, \u{1f600}, which needs no pair.

Unescaping

Decoding understands \uXXXX (pairing surrogates back into one character), \u{…}, two-digit \xXX, and the common single-letter escapes \n, \t, \r, \\, \" and \'. Anything that is not a complete escape is left exactly as written rather than rejected, so you can paste a whole log line or JSON document and only the escapes change.

When encoding, a backslash already in your text is doubled (\\), so that C:\new is not read back as a line break and encoding then decoding always returns the original. Quotes are left alone, so add your own escaping for those if you are pasting the result between quotes.

Frequently asked questions

Why is an emoji written as two \u escapes?
\uXXXX has room for four hex digits, which covers code points up to U+FFFF. Emoji sit above that, so they are split into a surrogate pair of two escapes, exactly as they are stored in UTF-16.
Does JSON allow \u{1f600}?
No. JSON only has the four-digit \uXXXX form, so characters above U+FFFF must be written as a surrogate pair. The braces form is JavaScript (ES2015 and later) syntax.
What is the difference between \u00e9 and %C3%A9?
\u00e9 is the character's Unicode code point, used in source code and JSON. %C3%A9 is URL encoding of the two bytes that é takes in UTF-8. They describe the same character in different systems.

More bits and bobs