Job description
Responsibilities
* Oversee the lifecycle planning, construction operation, and architecture optimization of enterprise-level hybrid cloud infrastructure, Kubernetes container clusters, physical data centers, and underlying networks.
* Lead the implementation of multi-cloud strategies (AWS, Alibaba Cloud, Tencent Cloud, etc.), responsible for network interconnection in multi-cloud environments, convergence of permission boundaries, resource governance, cost measurement, and the construction of multi-active disaster recovery systems.
* Plan and implement modern software delivery pipelines (CI/CD), creating a closed-loop from code hosting, automated pipelines, artifact version control, security compliance scanning to gray release and second-level rollback.
* Responsible for capacity scaling, version evolution, core gateway maintenance, node fault self-healing, performance bottleneck breakthroughs, and security container protection of large-scale, production-grade Kubernetes clusters.
* Deeply participate in the architectural planning of global CDN networks and core dynamic-static separation businesses, tackling challenges such as intelligent scheduling, cache hit rates, origin station return, DNS resolution, HTTPS encrypted transmission, and cross-network link jitter.
* Establish a unified management mechanism for cross-regional, multi-data center, and large-scale physical servers, achieving standardized server delivery, automated configuration, panoramic monitoring, and capacity self-adaptation.
* Build a highly scalable, integrated observability engineering, connecting metric aggregation, distributed log indexing, call chain tracing, intelligent alert noise reduction, and root cause automatic positioning capabilities.
* Strengthen system resilience engineering, lead high-availability architecture design, second-level elastic scaling strategies, remote disaster recovery, and emergency response specifications, regularly participate in major online fault reviews and mechanism iterations.
* Deeply apply AI programming assistants (AI Coding) to collaboratively carry out automated script writing, IaC code generation, online fault analysis, configuration security audits, and technical documentation accumulation.
* Continuously track cutting-edge developments in cloud-native, container security, automated operations, and AI-assisted R&D, driving technological evolution and deep integration with business scenarios.
Requirements
* Bachelor’s degree or above in Computer Science, Telecommunications, Software Engineering, or related fields.
* More than 5 years of practical experience in DevOps, SRE, large cloud platform architecture, or system stability operations.
* Proficient in mainstream public clouds such as AWS, Alibaba Cloud, or Tencent Cloud, with the ability to design cloud-native resource architecture, business deployment, and complex fault troubleshooting from scratch.
* Mastery of core cloud components, including elastic computing, virtual private cloud, load balancing systems, object storage, distributed databases, DNS, CDN, IAM, and security boundary division.
* Proficient in CI/CD delivery methodologies, skilled in using tools like Jenkins, GitLab CI, Argo CD, with the ability to design and implement complex pipelines.
* Proficient in Kubernetes container orchestration and cloud-native ecosystems, with a deep understanding of Docker, Helm, Ingress, CNI networking, CSI storage, and cluster scheduling mechanisms.
* Strong foundation in operating large-scale production clusters, capable of independently troubleshooting and resolving issues such as cluster network congestion, node downtime, resource preemption, and Pod anomalies.
* Familiar with the underlying logic of CDN operations, with experience in edge node performance tuning, bandwidth cost control, and cross-operator network link quality analysis.
* Proficient in the Linux operating system kernel, with a deep understanding of process scheduling, virtual memory, file systems, network protocol stacks, and system parameter tuning.
* Experience in managing large-scale physical servers, familiar with out-of-band management, bulk out-of-band upgrades, system initialization, and automated asset inventory.
* Proficient in infrastructure as code (IaC) tools such as Terraform, Ansible, Pulumi.
* Solid understanding of the TCP/IP protocol stack, DNS recursive resolution, BGP routing, HTTPS handshake, NAT conversion, firewall policies, and network packet capture troubleshooting methods.
* Proficient in monitoring and observability technology stacks such as Prometheus, Grafana, ELK/OpenSearch, Loki, SkyWalking.
* Engineering development capability in at least one programming language (Shell, Python, or Go), able to independently write operational tools and platform services.
* Strong ability to withstand pressure and make decisions on-site during failures, capable of quickly blocking risks, controlling impact, and driving root cause resolution in complex emergencies.
AI-assisted R&D (AI Coding) Hard Requirements
* Practical background in deeply internalizing AI programming tools into daily operations and development flows, not just at the trial stage.
* Proficient in advanced usage techniques of mainstream AI-assisted tools (such as GitHub Copilot, Cursor, Claude Code, etc.).
* Able to efficiently leverage AI to support the following scenarios:
* Writing and refactoring Shell, Python, Go operational scripts;
* Writing Terraform, Ansible, Kubernetes YAML configuration files;
* CI/CD pipeline logic orchestration;
* Deep analysis of massive logs and fault anomaly localization;
* Security compliance configuration review;
* Writing automated test cases and operational manuals.
* Strictly adhere to the
