Unicode Escape and Unescape
Turn accents, symbols and emoji into \uXXXX escapes that survive any ASCII-only system, or make an escaped string from a log or a JSON file readable again.
Unicode escapes examples
| Text | Encoded |
|---|---|
| café | caf\u00e9 Non-ASCII: Only the é is escaped |
| £5 → €6 | \u00a35 \u2192 \u20ac6 Non-ASCII: ASCII stays readable |
| 😀 | \ud83d\ude00 Non-ASCII: Beyond U+FFFF, so a surrogate pair |
| 😀 | \u{1f600} \u{…}: The ES2015 form holds it in one escape |
| Hi | \u0048\u0069 Everything: ASCII is escaped too |
Encoding guides
How each one works, with worked examples.
Everything happens in your browser. What you type or paste is never uploaded, stored or added to the share link, so it is safe to use with tokens, keys and customer data.
What a Unicode escape is
A Unicode escape writes a character as a backslash, a u and the four hex digits of its code: é is \u00e9. The same syntax is understood by JSON, JavaScript, Java, C#, Python and many configuration formats, which makes it the standard way to carry non-ASCII text through something that only handles ASCII reliably: old source files, .properties files, logs and some APIs.
Four hex digits only reach U+FFFF. Characters beyond that, which includes all emoji, are written as two escapes called a surrogate pair: 😀 is \ud83d\ude00. This is the form JSON requires. JavaScript since ES2015, along with Rust, Swift and others, also has a code point form in braces, \u{1f600}, which needs no pair.
Unescaping
Decoding understands \uXXXX (pairing surrogates back into one character), \u{…}, two-digit \xXX, and the common single-letter escapes \n, \t, \r, \\, \" and \'. Anything that is not a complete escape is left exactly as written rather than rejected, so you can paste a whole log line or JSON document and only the escapes change.
When encoding, a backslash already in your text is doubled (\\), so that C:\new is not read back as a line break and encoding then decoding always returns the original. Quotes are left alone, so add your own escaping for those if you are pasting the result between quotes.
Frequently asked questions
- Why is an emoji written as two \u escapes?
- \uXXXX has room for four hex digits, which covers code points up to U+FFFF. Emoji sit above that, so they are split into a surrogate pair of two escapes, exactly as they are stored in UTF-16.
- Does JSON allow \u{1f600}?
- No. JSON only has the four-digit \uXXXX form, so characters above U+FFFF must be written as a surrogate pair. The braces form is JavaScript (ES2015 and later) syntax.
- What is the difference between \u00e9 and %C3%A9?
- \u00e9 is the character's Unicode code point, used in source code and JSON. %C3%A9 is URL encoding of the two bytes that é takes in UTF-8. They describe the same character in different systems.
More bits and bobs
- JWT DecoderDecode a JSON Web Token to read its header, claims and expiry in plain English, and check an HMAC signature, without the token leaving your browser.
- Cron Expression BuilderBuild or decode a crontab schedule field by field, read it in plain English and see the next five runs in any time zone.
- Regex TesterTest a JavaScript regular expression on your own text with live highlighting, groups, replace, a plain-English explanation and code for six languages.
- UK Salary CalculatorTake-home pay after income tax, National Insurance, pension and student loans, with Scottish rates, tax codes, bonuses and overtime.