First, I Hope Everyone Understands
json has never been the first choice for serialization in Python; json is used only to communicate with other applications. For internal Python communication and object serialization, use pickle — just as with C/C++ you basically wouldn’t consider json; for internal communication only, you should use ProtoBuf.
A Problem I Encountered
In Python you inevitably define some classes, and serializing custom classes becomes a troublesome matter. For the class below, suppose we want to format an object.
1 | class Inner(): |
So we consider adding a function to this Inner class, making it look like this.
1 | class Inner(): |
Then you think the world has become normal again, everything satisfied. Or is it?
The wish is excellent, but… normal life is not composed of a single class. The class above is named Inner — might there be a class named Outer containing it? Seems quite likely.
1 |
|
So you decide to add a function to Outer too; something like the following seems to achieve the effect.
1 | class Outer(): |
The World Surely Has More Than Two Classes
Now we face a problem: when similar situations appear again, we have to wrap them layer by layer,
purely by hand. Can you imagine how special this process is? And every object needs to call
to_json-style functions — truly hard.
At this point you should consider Google and StackOverflow; maybe someone has encountered this problem before.
I searched with keywords like python json dumps custome object.
How to make a class JSON serializable was the first result; it introduces some methods, for example:
1 | class CustomJsonEncoder(json.JSONEncoder): # define a class inheriting from json.JSONEncoder |
With this approach we can update the original Outer and Inner classes and combine them with CustomJsonEncoder:
1 | class Inner(): |
Believe it or not, it’s quite successful. This way, classes you define yourself can also implement the to_json method.
The downside is that every json.dumps now needs the trailing cls=CustomJsonEncoder.
Can It Be More Elegant
json.dumps(o, cls=CustomJsonEncoder) is really quite long; it would be nice to simplify it
a little — how good it was originally: json.dumps(o).
I was fairly lucky and found How to make a class JSON serializable again. It provides
an idea: since during json.dumps the default JSONEncoder can be found by us,
replacing it is also possible. This is the code from the original article (the article has other methods worth reading too).
1 | from json import JSONEncoder |
In my own project I switched to another form: I want the original default to be called first, and only then
our custom to_json. Also, for Decimal objects in the project the original function doesn’t seem to work,
so I had to pull that out separately.
1 | def _default(self, obj): |
** << Dive Into Python >> ** There are no constants in Python; everything can be changed if you try hard enough. This satisfies one of Python’s core principles: bad behavior should be overcome, not outlawed.
What Do You Think of This Method?
The benefit is that you can drop the cls=CustomJsonEncoder.
The accompanying downside is that your program has modified the original default in JSONEncoder. You see, in my
program above I did my best to let the program call the original default first — even so, Decimal objects
still had problems serializing and had to be pulled out separately.
Another problem: every class you define means you need to implement a to_json method.
This kind of implicit convention is always a place where problems can arise.
Fortunately, what we should normally use is pickle, not json.
json should only be used in this one situation: when you want to communicate with programs in other languages.