Millet Porridge

English version of https://corvo.myseu.cn

0%

Some Understanding of Python Encoding Issues

I encountered many Python2 programs at work, where str, unicode and bytes caused me a lot of trouble. After working hard to understand these encoding issues, let me briefly share my understanding.

In this article:

  1. Nothing about the origin of utf8 or its encoding scheme will be covered (I neither know it nor have I dug deeply into it).
  2. I’ll introduce some problems existing in Python2 programs (when handling Chinese or special characters).
  3. How programs should cope with the upgrade from Python2 to Python3.

Common String Operations in Code

The snippets in the following sections are all code that runs in Python2 (more or less problematic in Python3):

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
# 1
ret = subprocess.check_output(cmd)
if PY3:
ret = ret.decode('utf-8')

# 2
if PY3:
value = json.dumps(local_conf[key]).encode('utf-8')
else:
value = json.dumps(local_conf[key])

# 3
line = p.stdout.readline()

# 4
with open(filepath, 'w') as fp:
fp.write(content)

# 5
raw = base64.b64decode(raw)

# 6
msg = '设置成功'

# 7
__name_mapping__ = {
Disabled: u'关闭',
Enable: u'开启',
}

If a program contains this many string-related tricks, I can almost certainly say this Python2 program was absolutely not written by one person: some noticed the unicode problem, some didn’t, and some made efforts for Python3 compatibility. But I must say the program above is disorganized. For subsequent maintenance, or when you want to upgrade the program bit by bit, you’ll find encoding issues must be handled everywhere in the program — encode and decode may appear anywhere.

encode and decode

I’ve said the blog won’t consider too many low-level encoding issues; let me briefly describe Python’s two ways of handling characters:

encode decode

The processing form above matches intuitive understanding: in Python, conversion happens through the two operations encode and decode, from string => character array or character array => string.

In Python2, the conversion is like this (here str and bytes are the same thing):

1
2
3
       +----->  unicode  +-----+
decode | | encode
+-----+ str/bytes <-----+

In Python3, the conversion is like this (here str and bytes differ, and unicode was removed):

1
2
3
       +----->    str    +-----+
decode | | encode
+-----+ bytes <-----+

Back to the previous section: if encode and decode appear all over the program, it means you’re mixing these two data structures in the program.

Perhaps the program was originally consistent, but Python3’s arrival broke that consistency; the root cause is the non-universality of encodings.

Problems Caused by Inconsistency

For a weakly typed language like Python, you can hardly know a variable’s type when writing code. The types of all objects are only determined when execution reaches a specific statement. That is to say, you’ll fall into a kind of mystical debugging: in a statement with an encoding error, adding encode or decode makes the program run again. The error is fixed — but what about the next person taking over? They can only curse under their breath while silently enduring it.

Why Can’t We Stay Consistent?

I sincerely recommend everyone read “Code Complete”, which emphasizes more than once: keep uniformity inside the program. There are two kinds of data in total:

  1. Ordinary strings on which encode can be performed
  2. Character arrays on which decode can be performed

The ideal case is that the program keeps only one of them, so no encode or decode is needed at all. But unfortunately, even if our own program can, many libraries we reference may not support it. They may produce or consume different kinds of data.

At such times, what we should do is not let the two kinds of strings mix freely, but ensure that inside the program there is only one kind of data. After receiving data produced by other libraries, convert it immediately; and when calling other libraries, convert our data into the type they require. Remember, the only principle is: only one of the data types appears in the program you write.

Maintaining Consistency

Of course, I made many attempts before writing this blog. Below I present my own solution; you can also explore based on your existing programs. The only principle is to stay consistent.

  1. In my programs, all strings are stored as unicode (this is Python2 phrasing), i.e. the kind of data on which encode can be performed.

  2. All output produced by other libraries is filtered (using the to_unicode function)

  3. All calls made to other libraries are processed according to the library’s requirements

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
# 1
ret = to_unicode(subprocess.check_output(cmd))

# 2
value = to_unicode(json.dumps(local_conf[key]))

# 3
line = to_unicode(p.stdout.readline())

# 4 output files are special; the Python2/3-compatible way is to output bytes
with open(filepath, 'wb') as fp:
fp.write(to_utf8(content))

# 5 base64 needs bytes input
raw = to_unicode(base64.b64decode(to_utf8(raw)))

# 6 this is a Python2/3-compatible form
msg = u'设置成功'

Below are my to_unicode and to_utf8 functions.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
if sys.version_info > (3, 0):
unicode = str

def to_utf8(string, errors="replace"):
# type: (any) -> bytes
"""
Convert string "测试字符串" to b"测试字符串"

Use only when the program interacts with other libraries.
"""
if not isinstance(string, unicode):
return string
# Be quiet by default
logger.debug("Encoding string with: %r" % string)
try:
return string.encode("UTF-8", errors)
except UnicodeEncodeError:
raise UnicodeEncodeError("Conversion from unicode failed: %r" % string)

def to_unicode(string, errors="replace"):
# type (any) -> str
"""
Convert b"测试字符串" to "测试字符串".

All input strings must be converted, ensuring only str is used in the program
"""
if isinstance(string, unicode):
return string
# Be quiet by default
logger.debug("Decoding string with: %r" % string)
try:
return unicode(string, "UTF-8", errors)
except UnicodeDecodeError:
raise UnicodeDecodeError("Conversion to unicode failed: %r" % string)

You can copy this code directly and use it as part of your base library, ensuring only one kind of variable exists in the program. For the Python2-to-3 conversion, you hardly need to modify a single line of code, and the program won’t be drowned in encode and decode.

Also, about efficiency: you’re already using Python — do you really care about this tiny bit of format-conversion efficiency? This small efficiency sacrifice makes the program structure much clearer; I think it’s well worth it.

Summary

The original intent of this article is not to introduce a solution to encoding problems; I hope that after reading it you have your own understanding of program consistency, and can adopt suitable approaches according to the concrete code of your current program, optimizing and iterating step by step. That is what I really want to express.

References

String, Unicode and bytes in Python