Millet Porridge

English version of https://corvo.myseu.cn

0%

OpenSSH Series (Extra 2) - One Reason ssh Disconnects Right After Login, and the Fix

Following the last post How a Kubernetes-based PaaS platform provides ssh services (research): after I implemented this feature in our production environment, users reported that they disconnect immediately after logging into ssh, with no shell available.

This post records the cause of this situation and the solution, hoping everyone who meets related problems in Kubernetes environments can avoid them.

The Specific Phenomenon

It manifests as: after connecting to the server and printing the welcome message, it exits:

The Investigation Process

At first I thought the user’s mac system and client ssh had problems; later I tried with my own computer and the same phenomenon appeared — preliminarily determining it was a server-side problem.

ssh -vvv only shows logs like this; the log is too ordinary to reveal what the error is.

1
2
3
4
5
6
7
8
9
10
11
12
13
debug3: send packet: type 97
debug2: channel 0: is dead
debug2: channel 0: garbage collecting
debug1: channel 0: free: client-session, nchannels 1
debug3: channel 0: status: The following connections are open:
#0 client-session (t4 r0 i3/0 o3/0 e[write]/0 fd -1/-1/7 sock -1 cc -1)

debug3: send packet: type 1
debug3: fd 1 is not O_NONBLOCK
Connection to 10.96.115.41 closed.
Transferred: sent 3912, received 3292 bytes, in 0.1 seconds
Bytes per second: sent 33161.8, received 27906.1
debug1: Exit status 255

I found some posts, but they were basically useless:

https://stackoverflow.com/questions/23357926/ssh-closes-connection-immediately-after-login

https://unix.stackexchange.com/questions/148714/cant-ssh-connection-terminates-immediately-with-exit-status-254

Finally my testing found that ssh could connect in the test environment but absolutely couldn’t in the production environment. Recalling the operations I had done, there was this line:

1
/usr/bin/env | grep _ >> /etc/environment

Then I discovered that in the test environment /etc/environment was about 700 lines, while production had about 1400 more lines. I deleted half of the production text directly, confirmed it could connect, and finally found the problem: too many environment variables caused it. Annoying or what — first time I knew an ssh server can’t have too many environment variables.

In Kubernetes environment containers, the environment variables recording service information are especially numerous; several thousand is on the low side.

The Solution

My final solution: keep only important environment variables in /etc/environment, and save Kubernetes’ other environment variables to a specific file that users can find. The line above was split into the two below:

1
2
3
4
5
# grep -v means don't take these values — i.e. remove variables like this:
# STAGE_XXX_STATIC_PORT_80_TCP_PROTO=tcp
/usr/bin/env | grep -Ev "SERVICE_HOST|SERVICE_PORT|TCP_ADDR|TCP_PORT|tcp" >> /etc/environment
# save all environment variables to a certain file
/usr/bin/env > $env_file

Extension

Why does having too many environment variables cause an immediate exit?

I did a quick search of the OpenSSH code and found these lines, excerpted from the OpenSSH project:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
static void
read_environment_file(char ***env, u_int *envsize,
const char *filename, const char *whitelist)
{
FILE *f;
u_int lineno = 0;

f = fopen(filename, "r");
while (getline(&line, &linesize, f) != -1) {
## just look at this line — if the environment file has too many lines, it exits directly !!!
if (++lineno > 1000)
fatal("Too many lines in environment file %s", filename);
// ...
}
free(line);
fclose(f);
}

// The code also has this line, which basically confirms the cause:
read_environment_file(&env, &envsize, "/etc/environment",
options.permit_user_env_whitelist);

Summary

This blog mainly shares a specific problem I met when using an ssh server. If other Kubernetes users meet it, I hope this helps. Don’t panic when problems arise — just look slowly. Because it was an ssh service inside a container, no logs were printed at first; actually printing logs properly should have located this problem faster.

From my personal perspective, a connection that drops immediately after being established is generally a server-side problem; client problems would fail during the connection process, like the netcat packet-splitting problem in the earlier research report.