Introduction: A Way to Accelerate Image Building
I’ve seen quite a few Dockerfile build optimizations; I also introduced one before: Docker Series 7: Dockerfile advanced. But besides these methods, there’s an alternative scheme — one that optimizes image build time even more extremely and extends the concept of layers.
Problems Caused by Oversized Code Repositories
Our code uses Python. For such interpreted languages, running in containers means the code must be copied into the image — a common scenario:
1 | # this is our prepared code structure |
1 | # install dependencies |
This is a fairly common scheme, presumably what readers expect: we cached the code’s dependencies but not the code itself.
This raises a problem: if the src/sample repository itself is very large, then with each build, image upload, and even runtime image pull,
this layer stays huge — the time slowed during building is actually amplified 3x in this process.
Two-Stage Code Cloning
For this problem we introduced a new scheme: clone the code directly during the Dockerfile build, but in two stages. The first stage has only the basic clone statement — that layer gets cached directly. The subsequent fetch statement adds the use of git checkout, forcibly changing the current HEAD to the needed commit hash. The modified code:
1 | COPY src/sample/requirements.txt requirements.txt |
The code given above has run in our production environment for over a year with no anomalies reported.
The Concrete Effect

The first clone layer is about 180M; the second fetch is only 90M. With this modification, image push and pull also save at least 90M. Even at 10M/s network speed, we compressed the overall time by roughly 20s.
Scheme Summary
For a certain class of projects with lots of history but not especially large changes each time, the time and space saved are more obvious. But there’s one problem: the Dockerfile needs a directly clonable address, so I only recommend it for private repositories or open source code — otherwise there’s still a code-leak risk.
Cloning Code with the ssh Protocol
The two-stage clone above has two main problems:
- Our users all use GitLab; the clonable address must be manually assembled by users — poor experience
- That address goes directly into the Dockerfile — not especially safe
Based on these problems we plan to introduce a new scheme: change the git clone address to the ssh protocol. The Dockerfile then contains only the ssh address, and the final product contains no private key — no need to worry about code-leak risk.
stackoverflow has many discussions of this problem, e.g.:
- Two-stage build: although the final container has no ssh files, we still copied the user’s code in full
1 | # https://stackoverflow.com/a/66648529/5563477 |
- Using
--squashwhen building: after deleting the private key, squash compresses the layers into one, erasing the key’s existence in earlier layers. It achieves the function, but compressing to one layer means no caching at all
1 | # docker build -t example --build-arg ssh_prv_key="$(cat ~/.ssh/id_rsa)" --build-arg ssh_pub_key="$(cat ~/.ssh/id_rsa.pub)" --squash . |
- Using Docker buildkit — a new feature Docker supports; so far the most elegant scheme
From: https://stackoverflow.com/a/58883743/5563477
1 | export DOCKER_BUILDKIT=1 |
1 | # syntax=docker/dockerfile:experimental |
Future Image Build Methods
There’s already more than one docker image packaging tool on the market. I’ll just briefly introduce them here and look at their support for cloning code via the ssh protocol. Our production environment hasn’t used them yet; of course readers shouldn’t replace their currently stable build/packaging tools just for this feature.
kaniko
1 | docker run \ |
I couldn’t get the cache-dir form working — the local /cache directory never wrote data — so I only tested the cache-repo form, which lets kaniko put the cache in a repository, checking the repository’s cache at each image build.
I briefly tested kaniko’s caching scheme: for the corvofeng:develop repository, it defaults to caching at the /corvofeng/develop/cache address; you can also specify your own cache repository.

In GitLab runner, building images this way couldn’t be more suitable: https://docs.gitlab.com/ee/ci/docker/using_kaniko.html

kaniko also supports cloning code via the ssh protocol — just mount the ssh-agent’s corresponding unix socket into the container
1 | docker run \ |
1 | FROM python:3.8-alpine |
The corresponding effect:

buildah
buildah’s usage is fairly close to docker buildkit: unix sockets can be mounted during image building. Here’s a simple example:
1 | sudo buildah build --build-arg=SSH_AUTH_SOCK=$SSH_AUTH_SOCK --volume $SSH_AUTH_SOCK:$SSH_AUTH_SOCK . |
1 | FROM alpine |
The effect:

Summary
First, the optimization introduced in this blog doesn’t suit all projects — it mainly targets large project image builds. If your project is small, there’s no need to consider this approach.
Also, my borrowing of SSH_AUTH_SOCk is meant to show that existing tools already support mounting files or unix sockets during builds.
For passwords or private keys needed only at build time, file mounting can completely implement this.
Although this blog revolves around private project building and deployment to explain the optimizations, kaniko and buildah can absolutely be applied to open source projects. When creating containers for some command-line tools, I also choose these two tools first.
I haven’t deeply tested image builds for arm64 and the like; if you have such needs, I suggest confirming feasibility yourself.