Senior Site Reliability Engineer positions focus on delivering results in their domain. This page aggregates open Senior Site Reliability Engineer roles and what employers typically expect.
As a **Senior DevOps Engineer**, you will work closely with Product, Engineering, and AI teams to shape our infrastructure strategy, design resilient cloud architectures, and ensure our platforms are secure, scalable, and high-performing. You will play a key role in bringing AI systems into production, enabling reliable delivery, strong observability, and operational excellence across our products and internal systems. **Your Responsibilities** - Design and operate secure, scalable, and high-quality infrastructure that supports modern applications and advanced AI workloads - Build and maintain robust automation across CI/CD pipelines, infrastructure provisioning, and operational processes to improve reliability and minimize manual effort - Integrate AI-driven solutions into operational workflows to enhance efficiency, detect anomalies, and accelerate delivery - Apply strong systems engineering practices, including monitoring, incident management, performance optimization, and capacity planning - Establish and uphold DevOps best practices, ensuring reproducibility, testing, documentation, and operational excellence - Communicate technical decisions clearly and collaborate cross-functionally to support predictable delivery and effective problem-solving - Provide mentorship and technical leadership, raising the level of platform engineering, DevOps maturity, and overall engineering quality across the organization **Requirements** - **Experience & Discipline**: 6+ years of progressive experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Engineering. - **Cloud Expertise**: Strong, hands-on experience across multi-cloud environments (AWS, GCP, Azure), including expertise in networking, compute, storage, security, and cost optimization. - **Core Platform Stack**: Deep expertise in containerization and orchestration and extensive experience with Infrastructure as Code (IaC) (e.g., Terraform, Pulumi, CloudFormation). - **AI/ML Infrastructure**: Experience supporting or deploying AI/ML workloads (e.g., model inference, vector databases, GPU workloads), or strong familiarity with the infrastructure requirements for these systems. - **System Reliability**: Proven ability to design, build, and operate highly reliable, scalable production systems utilizing advanced Zero-Downtime Deployment Patterns (e.g., Blue/Green, Canary, progressive delivery, Preview Environments). - **Modern Delivery & Tooling**: Expertise in modernizing deployments via GitOps practices (e.g., ArgoCD, Flux) and building Self-Service Developer Platforms that enable engineering efficiency (e.g., environment automation, internal tooling). - **Networking & Edge Routing**: Experience implementing and managing Multi-Cloud API Gateways and Edge Routing solutions. - **Security & Hardening**: Strong background in platform security, including secrets management, Identity and Access Control (IAM), and Runtime/Security Hardening - **Observability**: Solid understanding and practical experience with modern observability stacks. - **Mentorship & Communication**: Excellent communication and collaboration skills with a proven ability to describe complex infrastructure decisions clearly and a background in mentoring engineers and driving improvements in engineering practices. - **Development Expertise**: Familiarity with modern programming languages like Node.js, NestJS, and Python is highly desirable for extending DevOps capabilities or integrating tooling.