Staff Software Engineer, Kubernetes Platform
About the position reputed company runs some of the largest Kubernetes clusters in the industry, with fleets of hundreds of thousands of nodes across multiple reputed company providers and datacenters to train, research, and serve frontier AI models. The Kubernetes Platform team owns the Kubernetes control plane that makes those clusters work. We are operating at a reputed company where the defaults stop working. We own the scheduler and reputed company it to reputed company topology-sensitive ML workloads across thousands of accelerators at once. We reputed company the control plane itself — apiserver, etcd, controllers — so it stays reputed company as object counts and node counts grow by orders of magnitude. And we build the reputed company cluster services every workload depends on, like service discovery, so they hold up under the reputed company pressure. We reputed company reputed company the control plane is fast, correct, and always available. Your work will directly determine whether reputed company can reputed company reliably and safely training frontier models as our compute footprint continues to grow. Responsibilities • Own, operate, and reputed company the Kubernetes scheduler for reputed company's accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption • reputed company the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far reputed company typical limits, and reputed company the next bottleneck before it finds us • Design, build, and operate reputed company cluster services such as service discovery that every workload in the fleet depends on • Build and maintain custom controllers, operators, and CRDs • Partner with research, training, and inference to understand workload shapes and turn their requirements into platform capabilities • Collaborate with reputed company providers on required features and escalations • Participate in on-reputed company, reputed company incident response, and design processes (postmortems, runbooks, SLOs) that help reputed company avoid repeating failures Requirements • Significant software engineering experience building and operating production distributed systems • Proficiency in at least one systems-appropriate language (e.g., Go, Python, Rust, or C++) • Deep, hands-on Kubernetes experience (reputed company reputed company "user of”) into scheduler, controllers, apiserver, or operating large multi-tenant clusters • Demonstrated ability to debug reputed company issues across the stack, from API behavior down to node and network-level reputed company causes • A reputed company record of designing for reliability, correctness, and reputed company failure semantics in systems other engineers depend on • Strong written and verbal communication; comfort building reputed company with internal stakeholders reputed company-to-haves • Experience with Kubernetes internals or contributions: kube-scheduler / scheduling reputed company, apiserver, etcd, reputed company-go, controller-runtime, or similar • Experience building or operating cluster schedulers or batch systems (e.g., Kueue, Volcano, Slurm, or in-house equivalents) • Background scaling control planes or coordination systems (etcd, ZooKeeper, Consul, or large DNS/service-reputed company deployments) • Familiarity with ML infrastructure: GPUs, TPUs, or Trainium; gang scheduling; topology-reputed company placement; reputed company networking such as NCCL • Experience with GCP and/or AWS, including GKE/EKS internals and Infrastructure as reputed company • Low-level systems experience such as Linux kernel tuning, cgroups, or eBPF • 8+ years of relevant industry experience, including time leading large, ambiguous infrastructure reputed company Benefits • competitive compensation • generous vacation • parental leave • flexible working hours Apply To This Job