Please note this continues the previous blog post: Some Understanding of Python Encoding Issues. Therefore Python2 code uniformly uses unicode. If you don’t understand this, please take some time to read the first post.
Why Can’t Data from json.dumps Be Printed?
Take Python2 as an example:
1 | # In python2 |
I searched online — of course there are solutions:
1 | json.dumps(info, ensure_ascii=False) # the result after adding ensure_ascii |
I believe three questions arise. With these three questions in mind, let me introduce them one by one:
- Why does adding
ensure_ascii=Falsesolve it? - Can the different data from these two dumps forms still be json.loads-ed back?
- Since there are two dumps forms, which should we use?
What Does ensure_ascii Do?
The Python2 official documentation has this sentence:
If ensure_ascii is true (the default), all non-ASCII characters in the output are escaped with \uXXXX sequences, and the result is a str instance consisting of ASCII characters only. If ensure_ascii is false, some chunks written to fp may be unicode instances. This usually happens because the input contains unicode strings or the encoding parameter is used. Unless fp.write() explicitly understands unicode (as in codecs.getwriter()) this is likely to cause an error.
A rough translation:
If ensure_ascii is true, all non-ASCII characters will be printed as \uXXXX sequences, and the result will be composed as str (bytes in Python2). If ensure_ascii is false, what you get will be a unicode object. This form occurs because the input contains unicode characters or the parameter was specified during encoding (json.loads). If fp.write cannot explicitly understand unicode, this conversion is likely to cause an error.
There are the following 2 differences:
- The return value’s type differs: ensure_ascii true returns bytes; false returns unicode
- The return values differ: when ensure_ascii is true, unicode characters take the form
\\u7b26— unicode is fully decomposed into ascii; when ensure_ascii is false,\u7b26still represents one unicode character
Can Both Forms Be Re-decoded into Dictionaries?
Since two kinds of strings can be produced, naturally someone asks whether both strings can be decoded back
to the original string. The answer is yes — both forms can be traced back with json.loads:
1 | json.loads('{"err_msg": "\\u4e2d\\u6587\\u5b57\\u7b26"}') |
The craziest part: after re-decoding, both forms give the same result.
How Should We Choose
Based on my own situation, two suggestions:
- For log printing and communication with other libraries, use the
ensure_ascii=Falseform — this guarantees the final string is definitely inunicodeform, keeping uniformity throughout the program - For data exchange inside a class, you can use the default behavior, because loads can turn it back.
Appendix: Similarities and Differences Between Python2 and Python3
Conclusion first: as long as your code ensures unicode is used (in Python2), you’ll get basically consistent results. Please ensure all simple data in dictionaries uses the unicode form.
The Two dump Forms
1 | # In python2 |
1 | # In python3 — programs no longer have the unicode form, but plain dumps data still can't be printed |
Re-loading the Two String Forms
1 | # In python2 |
1 | # In python3 — behavior consistent with Python2; still recoverable |