We are seeking a Kubernetes Global Topic Lead (GTL) to design, build, operate, and continuously improve enterprise Kubernetes platforms supporting both traditional application workloads and emerging AI/GPU-based workloads. The ideal candidate will possess strong Linux, containerization, networking, and Kubernetes administration skills, with hands-on experience operating managed Kubernetes services, particularly Azure Kubernetes Service (AKS).
This role will be responsible for platform reliability, automation, security, observability, CI/CD integration, and supporting development teams consuming Kubernetes services at scale.
Responsibilities
- Strong experience with Docker containers and Linux administration
- Solid understanding of DNS, networking, ports/protocols, and load balancing.
- Hands-on experience managing Kubernetes clusters, workloads, storage, networking, security, and operations.
- Experience integrating Kubernetes with CI/CD and GitOps processes.
- Knowledge of monitoring, observability, performance tuning, and troubleshooting.
- Experience with Rancher for Kubernetes management.
- Strong preference for AKS and EKS experience; GKE experience is also valuable.
- Understanding of container image management, vulnerability scanning, and workload security.
- Basic knowledge of supporting GPU-enabled AI/ML workloads in Kubernetes environments.
Preferred: Longhorn, Cilium, Kube-VIP, Fleet, Harbor, and Zabbix.
Education/Experience/Skills
- 3–7+ years of experience in Infrastructure, Cloud, Platform Engineering, DevOps, or Site Reliability Engineering (SRE) roles.
- 2–5+ years of hands-on experience administering and operating production Kubernetes environments.
- Experience deploying and supporting containerized applications using Docker and Kubernetes.
- Strong experience administering Linux-based systems and troubleshooting via command line.
- Experience with managed Kubernetes platforms, preferably:
- Azure Kubernetes Service (AKS), Amazon EKS (strong preference)
- Google GKE (Good to have)
- Experience with Kubernetes networking, storage, security, RBAC, secrets management, and cluster operations.
- Experience integrating Kubernetes platforms with CI/CD pipelines, GitOps workflows, and automation tools.
- Experience monitoring and troubleshooting production environments using observability and logging tools.
- Experience working in enterprise environments with multiple teams including infrastructure, security, networking, and application development.
- Exposure to Kubernetes-based AI/ML workloads and GPU-enabled clusters is advantageous.
- Strong communication skills with the ability to explain design choices and trade-offs to both technical and non-technical stakeholders.
- Collaborative working style, comfortable facilitating cross team discussions and design workshops.
- Self-driven and structured individual
