Software Development Engineer II Aws Eks positions focus on delivering results in their domain. This page aggregates open Software Development Engineer II Aws Eks roles and what employers typically expect.
We are looking for a Software Development Engineer II (SDE-2) to join the EKS Node Runtime team. In this role, you will design, build, and operate systems that power the compute layer for Amazon EKS, working on critical infrastructure that enables customers to run containerized workloads reliably and securely at scale. You will own the end-to-end AMI lifecycle including designing optimized Amazon Machine Images for EKS workloads across Amazon Linux distributions, implementing automated build and release pipelines with integration testing, and ensuring compliance with 21-day CVE patching SLAs through automated tooling. Your work will involve qualifying and certifying new GPU and accelerator instance types such as P6 (B200/B300), G7, and Trainium2, managing NVIDIA driver updates and multi-version support strategies, and integrating Dynamic Resource Allocation (DRA) drivers for GPU, EFA, and Neuron workloads. You will develop and operate the Node Monitoring Agent that runs as a DaemonSet on customer nodes, implementing health checks, automated log collection, and EFA monitoring capabilities to provide fleet-wide observability. A significant portion of your work will focus on enabling AI and ML workloads by implementing VM isolation runtimes with Nitro partition support for secure container startup, optimizing node startup sequences for Auto Mode, and building support for AI agent workload patterns including pod pause, resume, snapshot, and restore capabilities. You will also update and maintain container runtimes including containerd and runc, configure SOCI snapshotters for improved cold-start performance, implement advanced kubelet configurations for huge pages and CPU topology management, and ensure GPU Operator compatibility across the fleet. As an SDE-2, you will drive technical design decisions for node runtime architecture, collaborate with service teams across AWS to integrate new capabilities, provide 24/7 oncall coverage for operational support, respond to SEVs and customer escalations, and mentor junior engineers while raising the bar on engineering and operational excellence. This is an opportunity to work on foundational infrastructure that directly impacts millions of customers running containerized workloads on AWS, with particular focus on enabling the next generation of GPU-accelerated AI and ML applications. Key job responsibilities Design & Build: Architect and implement the node-level runtime infrastructure that powers EKS compute. Design optimized Amazon Machine Images (AMIs), integrate GPU drivers and accelerator support, implement VM isolation runtimes, and build node monitoring systems that operate reliably across millions of customer nodes. Operate at Scale: Own the operational health of the EKS node runtime layer handling millions of customer workloads across diverse instance types and accelerators. Participate in on-call rotations, respond to SEVs, resolve customer escalations, and drive operational improvements that reduce MTTR and improve fleet stability. Technical Leadership: Lead the design and implementation of complex node runtime features end-to-end, from requirements through AMI build pipelines, testing frameworks, deployment, and production validation. Drive technical decisions on AMI architecture, container runtime strategies, and GPU/accelerator integration. Kubernetes & Runtime Expertise: Work deeply with node-level Kubernetes components including kubelet configuration, device plugins, Dynamic Resource Allocation (DRA) drivers, and container runtimes (containerd, runc). Understand Linux internals, systemd, GPU drivers, and isolation technologies to build robust node capabilities. Cross-Team Collaboration: Partner with EKS control plane teams, EC2 instance teams, NVIDIA, AWS AI/ML services (SageMaker, Trainium, Inferentia), and open-source communities to deliver integrated solutions that enable customer workloads from traditional containers to cutting edge AI training and inference. Mentorsh…