Educational Background
Bachelor’s degree or above in Computer Science, Electronic Engineering, Automation, or a related field.
Skills Required
1、Proficient in Linux, Kubernetes, containers, and cluster scheduling.
2、Experienced in networking, distributed systems, or cloud-native infrastructure.
3、Proficient in Python, Go, Shell, or C++.
Key Responsibilities
1、Build GPU clusters, compute, storage, networking, and resource orchestration systems.
2、Optimize the stability, resource utilization, and system performance of AI/HPC clusters.
3、Support cluster monitoring, capacity planning, fault recovery, and continuous delivery.
Preferred Qualifications
Experience with GPU clusters, Slurm, RDMA, or AI infrastructure is preferred.