Early in the contest I didn’t quite understand the evaluation method; afterwards I read the benchmark source code and got a rough understanding. Programs I write myself will also need performance testing in the future, so I took this chance to read the evaluation tool’s code — it may come in handy.
So the focus of this post is not discussing how to optimize programs; the focus is sharing what I learned after reading the benchmark program. At the end I’ll also share my own optimizations of these days, to provide some ideas.
I lack friends to form a team; if you’re interested, why not join (though the experts have probably all teamed up already).
Contest Overview
Because the deployment forms of the programs differ, you can’t just read the source code — you must look at it together with the real architecture. Below is the architecture diagram:

This diagram only contains our own program; external requests are made through the
interface of the Spring Cloud Consumer Service.
The 4 Agents in the diagram are for coordination. The left Agent chooses the suitable service provider;
if the Consumer Service side is not the bottleneck, throughput can be quickly increased by adding machines on the right.
In the contest, what we need to do is rewrite the Agent so it can guarantee high concurrency.
The benchmark Program
The benchmark program mainly uses Python and Shell scripts, with roughly these functions:
- Deploy and run etcd, and the service (using docker)
- Run wrk to benchmark performance
- Collect and package log files
The program running benchmark is not the actual evaluation machine; the real evaluation machine requires ssh login
and remote operation.
1 | def __run_remote_script(self, script): # remote command execution, worth learning from |
Whether docker pull or docker run, the operations are all invoked on the remote server.
benchmark Program Startup Flow
Read the config file => request the mock server to get the image of the team being evaluated =>
docker pull to build etcd and service => start etcd and service => stress test with wrk
What Does docker run Actually Do
This starts with the image we need to submit. TL;DR — if you’re not interested, skip to the next section.
How is the image we submit built? Take the Dockerfile of agent-demo as an example:
1 | # Builder container |
What this Dockerfile builds is actually a single image containing consumer, provider, and agent-demo,
started with different parameters passed in. For example, docker run <image> provider-small starts the
provider-small service and the agent service corresponding to provider-small.
During evaluation, we start the programs via consumer, provider-small, provider-medium,
and provider-large.
The Principle and Use of Wrk
I remember in February this year my sister came asking whether there was any stress-testing software; I found wrk then. It doesn’t have much code — suitable for reading and then bragging to interviewers about, haha.
The main logic code is in src/wrk.c: it initializes multiple threads, and each thread uses one epoll or
select to listen on multiple socket connections.
The request scripts are written in lua and called from C; example scripts are in the scripts directory.
Calling Lua scripts from C: Simple Lua Api Example
Using epoll for stress testing can effectively increase concurrency
Performance optimization of a government website
This is a simple example; I feel its testing plan and metrics are worth referencing.
Some Optimizations I Made
Using Nginx as the consumer agent
Functionally, the consumer-agent is positioned as a simple reverse proxy tool. It can be completely replaced by nginx.
In the end I set different upstream weights — a simple reverse proxy.
1 | upstream cluster { # the real IPs are not fixed; they must be filled in before nginx starts |
Using golang to increase the provider agent’s concurrency
This optimization is nothing special. Leveraging golang’s high-concurrency features, concurrency improved,
but I had to write the dubbo interfaces myself. Currently my dubbo interface uses short connections — every service call
must establish a connection, which is very slow. My current score is just 200-something (hope the experts won’t laugh).
Out of respect for the contest, I’ll open-source my code after it ends.
6-3: Using long connections for dubbo interface calls
The score reached 3000; I’m basically in the parameter-tuning stage now.

6-10: Using long connections in the Nginx reverse proxy
Mainly added the use of long connections in the Nginx reverse proxy, referring to how to keep connections during Nginx reverse proxying. The score reached 4500 — a bit of progress at least.
