I was responsible for the project’s early migration work; later, due to compliance issues I couldn’t handle the US environment, so other colleagues continued and completed the operations. This blog records the problems we met during migration and some solutions, hoping to give readers a line of thinking for cloud-native schemes in multi-cloud environments. I assume readers have Docker and Kubernetes experience and have deployed some applications with them.

Project Timeline
- Sep ‘22: When I joined, my boss told me one task was migrating our project from the Mesos cluster to a Kubernetes cluster
- Apr ‘23: Started sorting out all alarm rules, organizing logs and other data-reporting chains
- Jun ‘23: Pre-migration planning and preparation, and attempting to build services from scratch
- Oct ‘23: The world’s first region migrated successfully; multiple regions migrated in parallel afterwards
- Jan ‘24: All overseas regions except the US migrated successfully. I took a very long annual leave until the end of the Spring Festival holiday — very happy
- Mar ‘24: The migration work was taken over by other colleagues; the US region migration began
- Jul ‘24: The US portion completed migration; all overseas environments had migrated to Kubernetes clusters
Prerequisite Work

The following contents are actually rather abstract, yet very necessary. The necessity: you must tell your boss, tell colleagues, tell partner departments that the whole migration’s success rate is above 90%, and that fallback solutions exist when problems occur.
Sorting Out the Existing Environment
Content to sort out mainly includes: monitoring alarms, migration scope.
We generally distinguish alarms by whether they have Player Impact, to guarantee users feel nothing during the whole migration. The alarm sorting work took me quite a while — mainly confirming data sources and data-reporting chains; sometimes I went directly to the backend code to confirm the concrete alarm logic.
Demo Environment Setup and Verification
This part directly determined our choice of migration scheme. Although the partner provided a Helm Chart, in concrete landing you need to step through the possible pitfalls beforehand to be confident in making the plan and countermeasures.
Department Coordination
This part of the work isn’t only for reporting to the boss — it’s also to make partner colleagues trust us more, to make them feel this work is assured too.
A rough migration timeline must be given, with risk assessment and rollback plans done well.
For example, our system distinguishes major versions from hotfixes — one major version every 3 months. So when planning the timeline we must consider the major-version timeline; best to finish between two major versions, avoiding incidents during major version releases.
Problems and Solutions
Here I record some problems we once met. Much time has passed, so I’ll tell it from memory.
Pipeline Transformation

Our original pipeline was built on the BK GUI; every modification required clicking around on pages, and version control wasn’t good either. Based on this, we switched the whole pipeline to Jenkins.
I previously wrote Programming with Groovy: Flow Control Techniques and Insights; the concurrency and retry techniques used there are basically the core of the whole pipeline.
So far the pipeline has run stably for a year — basically very stable, with no major changes like functional refactoring.
I should write another blog on doing pipeline unit testing and project planning well in Jenkins.
Stateful Nodes

Although our business already ran in containers on Mesos before migration, machines still had several directories shared by all containers on that machine.
For this we used two DaemonSets to sync these directory files on machines, but risks still existed.
Later a colleague proposed a solution: newly created Nodes carry a taint mark; if we haven’t manually operated on this Node,
it won’t be put into use. We still haven’t fully solved this problem — the consequence being automatic scaling has been slow to land.
It taught me a sufficient lesson: the business must not depend on stateful nodes from the start. In plain words: mount cloud disks or volumes in yourself — don’t use folders on the machine.
AWS CNI and Multi-NIC Problems

On some US servers, a separate acceleration IP is bound to reduce players’ game latency — originally meant to improve user experience. But colleagues found that in the EKS cluster, NICs we added ourselves were automatically deleted by vpc-cni. They had no ideas at the time, so I said I’d go read vpc-cni’s code.
For the deletion reason, I had ChatGPT analyze the code:
https://chatgpt.com/share/8e96aedc-b297-4f97-b765-267d108229f9
Then I saw this field, which can make vpc-cni ignore our NICs.
1 | // eniNoManageTagKey is the tag that may be set on an ENI to indicate ipamd |
https://docs.aws.amazon.com/eks/latest/userguide/pod-multiple-network-interfaces.html https://docs.aws.amazon.com/whitepapers/latest/ec2-networking-for-telecom/multus-container-network-interface-cni.html
1 | // https://github.com/aws/amazon-vpc-cni-k8s/blob/06828cee09446fd9e501984727ed807254385cb8/pkg/ipamd/ipamd.go#L1348 |
Of course one could say we hadn’t read enough documentation. Still, getting the chance to briefly read base-component code and solve the problem felt great.
Hybrid Cloud Services
Among the cloud services our business uses, for policy reasons part is neither AWS nor Tencent Cloud — we merely leased several physical machines, networked via dedicated lines. These machines also needed adding to Tencent Cloud’s TKE to complete the overall K8S upgrade. Here we leveraged Tencent Cloud’s registered nodes. Only the network side differed slightly — we used the hostNetwork approach directly to avoid the problem.
https://www.tencentcloud.com/document/product/457/60282
Some Thoughts

Terraform and IaC
What do we use Terraform for? I think it must be thought through before using it.
There’s a concept called SSoT (Single Source of Truth). Mapped to our system architecture, its meaning is:
what functions our current machines have, and what tasks they respectively carry.
Before Terraform, this data might be stored in a CMDB — also a kind of SSoT — but CMDB maintenance requires manual work.
That is, from modifying a machine’s attribute to it finally being reflected in the CMDB, there’s a process.
Terraform’s existence simplifies this process: when we make infrastructure changes, Terraform’s state is an SSoT,
and the original CMDB can degenerate into merely a display platform.
Here I recommend the Teragrunt+Atlantis approach — truly combining infrastructure changes with approval flows. From an engineer’s perspective, this flow is beautiful.
Some Thoughts on Cloud-Native Problems
Cloud Native or Cloud Provider Native?
After the migration ended, I received a task to migrate part of the original logic into the Jenkins pipeline — logic that goes into a certain machine and executes some scripts. I don’t know what schemes readers can think of; I’ll directly list these, and I’m sure you’ve thought of some:
- Use ssh to log into the machine and execute the corresponding script — simplest, but low extensibility, and our machines have basically canceled ssh service
- Use the SSM/TAT services provided by AWS or Tencent Cloud, calling the cloud vendor’s script execution API
- I already have a Kubernetes cluster, so I can create a new pod and execute
kubectl exec -ti xxxx run.shin Jenkins During exec, stdout and stderr logs are retained — equivalent to using Jenkins to preserve execution logs
I finally chose the third and persuaded my boss; it’s still running this way now. These were my considerations:
- Distinguish well between cloud-native and cloud-provider-native. When we use vendor-specific APIs, we should ask ourselves: is this necessary?
- Distinguish business requirements: which requirements come from the business, which from infrastructure — is it worth writing a multi-cloud-vendor compatibility layer for this requirement?
Global Collaboration
Being in a globalized team is actually very happy. I can think of 2 points:
- You needn’t worry much about PageDuty (oncall) — US colleagues can handle problems during our early morning.
- For US-East colleagues, we’re exactly 12 hours apart: CN’s daytime I handle things, leaving good notes on work results; CN’s night they continue — work efficiency is very high.
Calling people up for oncall in the early morning is inhumane; ops globalization is definitely the trend.
Organization and Summary
Actually I wanted to write this blog in July; recently I finally had time to organize my thoughts. Solving real problems is still very fulfilling, but you must keep good records, haha. Otherwise you can only painfully search your memory when writing the blog.
Image by Freepik