This investigation happened in February 2021.
The Specific Phenomenon
After the application migrated to our PaaS platform, sporadic 502 problems appeared; see the picture for the error:

Compared with the program’s request volume the errors are certainly few, but the errors keep happening and affect the caller’s code, so the cause needs checking.
Why Did We Only See POST Requests
Readers will surely say: your elk filter field says POST, so of course there are only POST requests, haha. Actually that’s not it — GET requests also get 502s, it’s just that Nginx retries GET requests, producing logs like the following:

The retry mechanism is Nginx’s default: proxy_next_upstream
http://nginx.org/en/docs/http/ngx_http_proxy_module.html#proxy_next_upstream
Because the GET method is considered idempotent, when an upstream returns 502 nginx tries again. For our problem, we mainly wanted to confirm why there were 502s; looking only at POST requests was enough — the causes of both should be the same.
Network Topology
When network requests flow into the cluster, for our cluster’s structure:
1 | user request => Nginx => Ingress => uwsgi |
Don’t ask why there’s Nginx as well as Ingress. Historical reasons — some work temporarily needs to be borne by Nginx.
Statistical Investigation
Based on our error request statistics for Nginx and Ingress, we found the 502 errors in the two were equal,
meaning the problem must occur between Ingress<=>uwsgi.
Packet Capture
This wasn’t the first solution we thought of — we used many statistical methods and found no pattern, so in the end we could only hope for packet capture…
The request log was as follows:

The capture result was as follows:

From the capture, the current tcp connection was reused. Since HTTP1.1 is used in the Ingress, it tries to send a second HTTP request within one tcp connection, but uwsgi doesn’t support http1.1, so the second request is not processed at all — directly rejected. So from the Ingress’s perspective this request failed, hence the 502. Since GET requests retry but POST requests cannot, POST-request 502s appeared in the access statistics.
Learning Ingress Configuration
In Ingress, the default http version for upstreams is 1.1, but our uwsgi uses http-socket, not http11-socket. Our Ingress used a different protocol from the backend, producing unexpected 502 errors.

Because Nginx defaults to 1.0, but switching to Ingress the default became 1.1, and we hadn’t unified them. The solution is forcing the http protocol version used in the Ingress.
{% if keepalive_enable is sameas true %}
nginx.ingress.kubernetes.io/proxy-http-version: "1.1"
{% else %}
nginx.ingress.kubernetes.io/proxy-http-version: "1.0"
{% endif %}
If any expert sees this, please briefly explain when Ingress reuses http1.1 connections, and why Ingress doesn’t reuse every connection — that way problems would surface sooner. I didn’t dig further into these questions. After all, switch to another language like Golang and this problem doesn’t exist; it’s a uwsgi-specific error.
Summary
Regarding this 502 investigation, I personally feel that finally solving the problem in one shot via packet capture wasn’t anything special: capture and you find the problem; don’t capture and you don’t.
What I hope is to give everyone an idea of how to troubleshoot a whole request chain: here we used both Nginx and Ingress. When troubleshooting, first check the error counts of both; if the errors are confirmed basically equal, that means the errors have nothing to do with Nginx and you need to check errors on the Ingress. For requests with multiple relays, this kind of investigation can fairly quickly locate the error position in the chain.