Millet Porridge

English version of https://corvo.myseu.cn

0%

Stress Testing with wrk and Fine Control of Concurrent Requests

In the earlier article Completing simple custom dynamic requests with wrk, I introduced how to make random requests with wrk and gave the usage of lua scripts. This blog mainly introduces how to finely control concurrent requests with wrk during stress testing.

wrk’s Parameters

wrk has no qps control option; it can only control the number of connections, and the specified connection count is evenly distributed across threads.

1
2
3
4
5
6
7
8
9
10
11
Usage: wrk <options> <url>                            
Options:
-c, --connections <N> Connections to keep open
-d, --duration <T> Duration of test
-t, --threads <N> Number of threads to use

-s, --script <S> Load Lua script file
-H, --header <H> Add header to request
--latency Print latency statistics
--timeout <T> Socket/request timeout
-v, --version Print version details

For example ./wrk -t8 -c1000

means starting 8 threads, each maintaining 125 connections.

The Relationship Between Connection Count and Concurrency

The number of requests processable within 1s is called qps.

Consider this problem: suppose you want the effect of 1000qps, and your processing capacity is one request per 0.2s. Assuming enough servers, how can you have 1000 requests within this 1s?

You can divide 1s into 5 portions (0.2s each), and the 1000 requests occur evenly across the 5 portions — 200 requests per portion, so we need 200 concurrent connections.

Practice

This is a rough estimation. In actual use, note one thing: if your request response is itself very fast, say 0.05s, the concurrency estimate may not be so accurate — mainly because other time is consumed along the request chain. If we use 200 concurrent connections, 200/0.05 theoretically gives 4000 qps, but other overheads making the concurrency lower than 4000 is quite normal.

When using wrk, I don’t raise the request count very high directly, because we don’t necessarily have enough backend machines; adding a large volume of requests at once may reduce service availability. You can increase the request count gradually. Here’s what I do: specify the response time of the tested content as 0.1s, then specify the concurrent connection count; when using it, run the script several times, and only continue increasing after seeing the qps stabilize.

1
2
3
4
5
6
7
8
9
10
11
-- ./wrk -t10 -c400 -d3600s -T2s -s 4k.lua
-- 400 concurrent requests can reach 4k qps
request = function()
local path = "/test/wait"
local body = "wait=0.1"

local headers = {}
headers["Content-Type"] = "application/x-www-form-urlencoded"
headers["Host"] = "XXX"
return wrk.format('GET', path, headers, body)
end

For instance now, the qps has stabilized:

Response time is also fairly stable:

Maybe some don’t understand what P50 means: the bottom green line indicates over 50% of requests complete within 0.2s.

This way you get data in a stable request state.

Summary

I hope readers understand wrk’s parameter settings and their actual effects in Nginx — maybe one day you’ll use them when stress testing.

Appendix – My Ingress Stress Testing Process

The recent Ingress stress test was mainly because a big application would join our system, possibly with more traffic than all existing applications combined. Without stress testing, users wouldn’t have enough confidence.

The stress test script used the 0.1s data above. The machine configuration:

1
2
3
4
5
6
Kernel version: 4.9.0-15-amd64
OS: Debian 9.13
Memory: 131072 M
CPU cores: 24
CPU count: 2
CPU: Intel(R) Xeon(R) Silver 4214R CPU @ 2.40GHz

The Background uwsgi Program

Each pod uses 40 workers with gevent enabled, CPU limited to 10 cores (utilization under 5 in tests). The pod count was kept sufficient during testing — 15 pods.

The requested interface is below; during testing the wait time was unified at 0.1s:

1
2
3
4
5
6
7
8
9
10
11
@index_handler.route('/test/wait', methods=['POST', 'GET'])
@response_process
def test_wait():
wait = float(request.values.get('wait', 5))
import gevent
gevent.sleep(wait)
result = {
'status': 'ok',
'wait': wait,
}
return jsonify(result=result)

Ingress Behavior at Various qps Levels

I stress-tested the Kubernetes Ingress machines, mainly testing Ingress metrics under various request volumes. Test time: 11:30~12:00

Concurrency CPU utilization mem usage Application response
0 20% 2.4G no requests
3.6k 160% 2.3G program response time stable
7.1k 291% 2.35G program response time stable
10.4k 413% 3.36 program response basically stable
11.8k 470 2.37G program response already slower: P50-180ms P90 244ms P99 470ms; adding machines didn’t help

After reaching the uwsgi bottleneck, I started adding static resource requests, trying larger concurrency

Concurrency CPU utilization mem usage Application response
13.2K 490 % 2.4G static resource requests respond normally

This program hit its bottleneck after reaching 13~14k. At this point I could only keep this program’s request volume and add another program for stress testing.

For the other program I no longer specified the wait time, hoping it returns as fast as possible so I get the highest possible concurrency.

Under concurrent requests, the new application’s qps was around 25k:

At this point the Ingress began approaching its performance bottleneck; you can see CPU utilization dropping at request peaks:

Conclusions of This Stress Test

14kqps (first application) + 25k (second application): a single-machine Ingress can roughly allow around 40K concurrency. After reaching this stage, the main bottleneck is CPU. With a better CPU, I think the concurrency could be higher.

If you think my stress testing method is unscientific or have anything else to say, leave a comment — I’ll check whether the process has problems.