Millet Porridge

English version of https://corvo.myseu.cn

0%

Handling JSON Strings Containing Chinese in Python

Please note this continues the previous blog post: Some Understanding of Python Encoding Issues. Therefore Python2 code uniformly uses unicode. If you don’t understand this, please take some time to read the first post.

Why Can’t Data from json.dumps Be Printed?

Take Python2 as an example:

1
2
3
4
5
6
7
8
9
# In python2
import json
info = {'err_msg': u'中文字符'}

json.dumps(info) # printed with print or other means, this form cannot display Chinese properly
# Out[37]: '{"err_msg": "\\u4e2d\\u6587\\u5b57\\u7b26"}'

print(json.dumps(info)) # you want to print it, and the result is this
# {"err_msg": "\u4e2d\u6587\u5b57\u7b26"}

I searched online — of course there are solutions:

1
2
3
4
5
json.dumps(info, ensure_ascii=False)  # the result after adding ensure_ascii
# Out[36]: u'{"err_msg": "\u4e2d\u6587\u5b57\u7b26"}'

print(json.dumps(info, ensure_ascii=False))
# {"err_msg": "中文字符"}

I believe three questions arise. With these three questions in mind, let me introduce them one by one:

  • Why does adding ensure_ascii=False solve it?
  • Can the different data from these two dumps forms still be json.loads-ed back?
  • Since there are two dumps forms, which should we use?

What Does ensure_ascii Do?

The Python2 official documentation has this sentence:

If ensure_ascii is true (the default), all non-ASCII characters in the output are escaped with \uXXXX sequences, and the result is a str instance consisting of ASCII characters only. If ensure_ascii is false, some chunks written to fp may be unicode instances. This usually happens because the input contains unicode strings or the encoding parameter is used. Unless fp.write() explicitly understands unicode (as in codecs.getwriter()) this is likely to cause an error.

A rough translation:

If ensure_ascii is true, all non-ASCII characters will be printed as \uXXXX sequences, and the result will be composed as str (bytes in Python2). If ensure_ascii is false, what you get will be a unicode object. This form occurs because the input contains unicode characters or the parameter was specified during encoding (json.loads). If fp.write cannot explicitly understand unicode, this conversion is likely to cause an error.

There are the following 2 differences:

  1. The return value’s type differs: ensure_ascii true returns bytes; false returns unicode
  2. The return values differ: when ensure_ascii is true, unicode characters take the form \\u7b26 — unicode is fully decomposed into ascii; when ensure_ascii is false, \u7b26 still represents one unicode character

Can Both Forms Be Re-decoded into Dictionaries?

Since two kinds of strings can be produced, naturally someone asks whether both strings can be decoded back to the original string. The answer is yes — both forms can be traced back with json.loads:

1
2
3
4
5
6
json.loads('{"err_msg": "\\u4e2d\\u6587\\u5b57\\u7b26"}')
# Out[15]: {u'err_msg': u'\u4e2d\u6587\u5b57\u7b26'}


json.loads(u'{"err_msg": "中文字符"}')
# Out[14]: {u'err_msg': u'\u4e2d\u6587\u5b57\u7b26'}

The craziest part: after re-decoding, both forms give the same result.

How Should We Choose

Based on my own situation, two suggestions:

  1. For log printing and communication with other libraries, use the ensure_ascii=False form — this guarantees the final string is definitely in unicode form, keeping uniformity throughout the program
  2. For data exchange inside a class, you can use the default behavior, because loads can turn it back.

Appendix: Similarities and Differences Between Python2 and Python3

Conclusion first: as long as your code ensures unicode is used (in Python2), you’ll get basically consistent results. Please ensure all simple data in dictionaries uses the unicode form.

The Two dump Forms

1
2
3
4
5
6
7
8
9
# In python2
import json
info = {'err_msg': u'中文字符'}

json.dumps(info)
# Out[37]: '{"err_msg": "\\u4e2d\\u6587\\u5b57\\u7b26"}'

json.dumps(info, ensure_ascii=False)
# Out[36]: u'{"err_msg": "\u4e2d\u6587\u5b57\u7b26"}'
1
2
3
4
5
6
7
# In python3 — programs no longer have the unicode form, but plain dumps data still can't be printed

json.dumps(info)
# Out[17]: '{"err_msg": "\\u4e2d\\u6587\\u5b57\\u7b26"}'

json.dumps(info, ensure_ascii=False)
# Out[18]: '{"err_msg": "中文字符"}'

Re-loading the Two String Forms

1
2
3
4
5
6
7
# In python2

json.loads('{"err_msg": "\\u4e2d\\u6587\\u5b57\\u7b26"}')
# Out[3]: {u'err_msg': u'\u4e2d\u6587\u5b57\u7b26'}

json.loads(u'{"err_msg": "\u4e2d\u6587\u5b57\u7b26"}')
# Out[4]: {u'err_msg': u'\u4e2d\u6587\u5b57\u7b26'}
1
2
3
4
5
6
# In python3 — behavior consistent with Python2; still recoverable
json.loads('{"err_msg": "\\u4e2d\\u6587\\u5b57\\u7b26"}')
# Out[2]: {'err_msg': '中文字符'}

json.loads('{"err_msg": "中文字符"}')
# Out[4]: {'err_msg': '中文字符'}