Job description
Responsibilities
1. Cloud Platform and Infrastructure Management
- Responsible for the design, deployment, and daily operations of production environments on AWS / Alibaba Cloud / Huawei Cloud (at least two), ensuring core business high availability (over 99.95%).
- Lead hybrid cloud or multi-cloud network architecture planning, including VPC interconnection, dedicated line access, and cross-availability zone disaster recovery design.
- Continuously promote cloud resource cost optimization, achieving quantifiable cost reduction through reserved instances, elastic scaling, resource governance, and other means.
2. Kubernetes and Container Platform
- Experience in deploying self-built or managed K8s clusters from 0 to 1, familiar with production-level operations such as control plane high availability, etcd backup and recovery, and certificate rotation.
- Responsible for technical support and promotion of business containerization transformation, formulating container resource specifications (Request/Limit, Pod anti-affinity, security context, etc.).
- Maintain cluster multi-tenant isolation policies (Namespace, ResourceQuota, NetworkPolicy, OPA/Gatekeeper).
3. Observability System (Prometheus + ELK)
- Design and maintain a monitoring and alerting system based on Prometheus + Grafana + Alertmanager, writing high-value alert rules to reduce false positive rates.
- Responsible for the cluster deployment of the ELK / EFK log platform, index lifecycle management, and the implementation and optimization of log alerts (ElastAlert or Kibana Alerting).
4. Middleware Operations
- Responsible for the cluster deployment, tuning, and troubleshooting of core middleware such as Kafka / Zookeeper / etcd / Redis.
- Familiar with middleware security hardening, including but not limited to TLS encrypted transmission, SASL/SCRAM authentication, and ACL permission control.
5. CI/CD and Release Engineering
- Lead or deeply participate in CI/CD pipeline design, familiar with tools such as Jenkins / GitLab CI / GitHub Actions.
- Promote GitOps release model (ArgoCD / FluxCD), achieving automated deployment and rollback across multiple environments (development / testing / production).
6. Collaboration and Documentation
- Write high-quality technical documentation, incident review reports, and operational SOPs, and promote team knowledge accumulation.
- Efficiently communicate with development, testing, and product teams during incident handling and architecture reviews, using data to drive decisions.
Requirements
Hard Requirements
- Full-time bachelor's degree or above in a computer-related field.
- Over 3 years of production environment operations or SRE experience, with at least 1 year focused on cloud-native technology stack.
- Familiar with at least two mainstream public clouds (AWS, Alibaba Cloud, Huawei Cloud), understanding the core IaaS/PaaS product differences.
- Practical experience in building production-grade Kubernetes clusters from 0 to 1, able to clearly describe the reasons for architectural choices.
- Proficient in deploying and maintaining Prometheus monitoring systems, having written custom Exporters or complex PromQL alert rules is a plus.
- Experience in building, tuning, and log alerting of ELK clusters, capable of handling TB-level log volumes.
- Familiar with the cluster deployment and troubleshooting of at least three middleware such as Kafka, Redis, etcd, with practical experience in security encryption preferred.
- Proficient in using at least one CI/CD tool and participating in pipeline design and implementation.
Soft Requirements
- Good communication skills, able to explain technical issues to non-technical parties regarding impact and action plans.
- Habit of documenting knowledge, understanding BPD (Blameless Postmortem) culture.
- Strong problem-driven and self-motivated, able to proactively drive improvements from operational pain points.
Bonus Points
- Holder of Hong Kong identity or eligible to apply for Hong Kong entry visa.
- Holder of CKA / CKAD / CKS certification.
- Holder of one of the following certificates: AWS SAA / SAP, Alibaba Cloud ACP / ACE, Huawei Cloud HCIE.
- Experience contributing to open-source projects or technical blogs/sharing.
- Familiar with either Golang or Python, able to develop simple automation tools or Operators.
