Senior Site Reliability Engineer positions focus on delivering results in their domain. This page aggregates open Senior Site Reliability Engineer roles and what employers typically expect.
**This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in Canada.** This role offers the opportunity to design, improve, and operate highly scalable cloud infrastructure supporting a global technology platform. You will work at the intersection of software engineering and infrastructure, helping build reliable systems that handle billions of daily requests. The position requires strong technical expertise, a problem-solving mindset, and a passion for automation, performance, and operational excellence. You will collaborate with engineering teams to strengthen platform reliability, improve observability, and develop solutions that support long-term scalability. The ideal candidate thrives in complex environments, enjoys solving challenging infrastructure problems, and takes ownership of critical systems. This is a hands-on engineering role where your contributions will directly impact platform stability, efficiency, and customer experience. ### Accountabilities: The Senior Site Reliability Engineer will be responsible for improving the reliability, scalability, and performance of complex cloud-based systems. This role combines infrastructure engineering, software development, automation, and operational leadership to ensure highly available services and continuous platform improvement. - Develop and improve infrastructure tooling and automation to reduce manual processes, increase efficiency, and minimize operational risks. - Build, maintain, and support critical applications and platform services. - Monitor systems for capacity, availability, and performance while proactively identifying and resolving technical issues. - Collaborate with engineering and SRE teams to ensure reliable delivery of services to customers. - Analyze system weaknesses, identify improvement opportunities, and implement solutions to strengthen platform resilience. - Design and maintain observability solutions, monitoring frameworks, and performance tracking systems. - Improve monitoring, alerting, incident response processes, and operational documentation. - Participate in on-call rotations, troubleshoot production issues, and create effective runbooks for incident management. - Contribute to technical strategies that support platform growth, scalability, and long-term reliability. - Support best practices around cloud infrastructure, container orchestration, and infrastructure-as-code. ## Requirements: The ideal candidate is an experienced infrastructure and software engineer with a strong background in building reliable, scalable systems. You should have hands-on experience with cloud platforms, automation, observability, and modern infrastructure technologies, along with the ability to operate effectively in a collaborative engineering environment. - 4+ years of experience as a Site Reliability Engineer, Software Engineer, or similar role focused on scalable and resilient services. - Strong experience designing, building, and maintaining cloud-based infrastructure using platforms such as AWS or GCP. - Experience creating automation tools and improving operational efficiency through engineering solutions. - Solid understanding of observability platforms, monitoring strategies, and service-level objectives (SLOs). - Experience with tools such as Prometheus, Grafana, Loki, Tempo, Thanos, or similar monitoring technologies. - Experience operating production systems in an on-call environment and improving incident response processes. - Strong infrastructure-as-code experience, with Terraform knowledge considered a major advantage. - Hands-on Kubernetes experience, including cluster operations, multi-tenancy strategies, and container orchestration best practices. - Proficiency with one or more programming languages such as Node.js, Go, Ruby, Python, and experience with shell scripting. - Strong Linux administration and troubleshooting skills…