In the earlier article Completing simple custom dynamic requests with wrk, I introduced how to make random requests with wrk and gave the usage of lua scripts. This blog mainly introduces how to finely control concurrent requests with wrk during stress testing.
wrk’s Parameters
wrk has no qps control option; it can only control the number of connections, and the specified connection count is evenly distributed across threads.
1 | Usage: wrk <options> <url> |
For example ./wrk -t8 -c1000
means starting 8 threads, each maintaining 125 connections.
The Relationship Between Connection Count and Concurrency
The number of requests processable within 1s is called qps.
Consider this problem: suppose you want the effect of 1000qps, and your processing capacity is one request per 0.2s. Assuming enough servers, how can you have 1000 requests within this 1s?
You can divide 1s into 5 portions (0.2s each), and the 1000 requests occur evenly across the 5 portions — 200 requests per portion, so we need 200 concurrent connections.
Practice
This is a rough estimation. In actual use, note one thing: if your request response is itself very fast, say 0.05s, the concurrency estimate may not be so accurate — mainly because other time is consumed along the request chain. If we use 200 concurrent connections, 200/0.05 theoretically gives 4000 qps, but other overheads making the concurrency lower than 4000 is quite normal.
When using wrk, I don’t raise the request count very high directly, because we don’t necessarily have enough backend machines; adding a large volume of requests at once may reduce service availability. You can increase the request count gradually. Here’s what I do: specify the response time of the tested content as 0.1s, then specify the concurrent connection count; when using it, run the script several times, and only continue increasing after seeing the qps stabilize.
1 | -- ./wrk -t10 -c400 -d3600s -T2s -s 4k.lua |
For instance now, the qps has stabilized:

Response time is also fairly stable:

Maybe some don’t understand what P50 means: the bottom green line indicates over 50% of requests complete within 0.2s.
This way you get data in a stable request state.
Summary
I hope readers understand wrk’s parameter settings and their actual effects in Nginx — maybe one day you’ll use them when stress testing.
Appendix – My Ingress Stress Testing Process
The recent Ingress stress test was mainly because a big application would join our system, possibly with more traffic than all existing applications combined. Without stress testing, users wouldn’t have enough confidence.
The stress test script used the 0.1s data above. The machine configuration:
1 | Kernel version: 4.9.0-15-amd64 |
The Background uwsgi Program
Each pod uses 40 workers with gevent enabled, CPU limited to 10 cores (utilization under 5 in tests). The pod count was kept sufficient during testing — 15 pods.
The requested interface is below; during testing the wait time was unified at 0.1s:
1 |
|
Ingress Behavior at Various qps Levels
I stress-tested the Kubernetes Ingress machines, mainly testing Ingress metrics under various request volumes. Test time: 11:30~12:00
| Concurrency | CPU utilization | mem usage | Application response |
|---|---|---|---|
| 0 | 20% | 2.4G | no requests |
| 3.6k | 160% | 2.3G | program response time stable |
| 7.1k | 291% | 2.35G | program response time stable |
| 10.4k | 413% | 3.36 | program response basically stable |
| 11.8k | 470 | 2.37G | program response already slower: P50-180ms P90 244ms P99 470ms; adding machines didn’t help |
After reaching the uwsgi bottleneck, I started adding static resource requests, trying larger concurrency
| Concurrency | CPU utilization | mem usage | Application response |
|---|---|---|---|
| 13.2K | 490 % | 2.4G | static resource requests respond normally |

This program hit its bottleneck after reaching 13~14k. At this point I could only keep this program’s request volume and add another program for stress testing.
For the other program I no longer specified the wait time, hoping it returns as fast as possible so I get the highest possible concurrency.
Under concurrent requests, the new application’s qps was around 25k:

At this point the Ingress began approaching its performance bottleneck; you can see CPU utilization dropping at request peaks:

Conclusions of This Stress Test
14kqps (first application) + 25k (second application): a single-machine Ingress can roughly allow around 40K concurrency. After reaching this stage, the main bottleneck is CPU. With a better CPU, I think the concurrency could be higher.
If you think my stress testing method is unscientific or have anything else to say, leave a comment — I’ll check whether the process has problems.