Jobedly Post a Job

Staff Software Engineer - Reliability (US Citizen Only)

Rubrik Job Board · Palo Alto, CA
Full-timeEngineeringTechnology$179,000–$241,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Staff Software Engineer Reliability role

Staff Software Engineer Reliability positions focus on delivering results in their domain. This page aggregates open Staff Software Engineer Reliability roles and what employers typically expect.

About Team & About Role The Site Reliability Engineering (SRE) team at Rubrik ensures the absolute reliability, availability, performance, and security of our enterprise infrastructure services, spanning both global SaaS platforms and government-compliant environments. We operate at the intersection of software development and systems engineering, prioritizing hyperscale platform automation, self-healing architectures, and structural resiliency. As a Staff Site Reliability Engineer, you will serve as a primary technical leader and architect across our broader distributed cloud systems. You will drive long-term technical roadmaps, establish cross-organizational reliability standards, and solve complex distributed systems challenges that safeguard both enterprise and public sector environments. Beyond the core SRE charter, this Staff role also leads the Application-SRE team — a US-based group that partners closely with engineering, Sales, and Support to unblock POCs, drive complex customer escalations to resolution, and convert recurring field signals into engineering and reliability roadmap items. You will be the technical leader and project owner for Application-SRE: setting direction, tracking commitments, and ensuring the team operates as a high-leverage bridge between the field and the broader engineering org. What You'll Do As a Staff Site Reliability Engineer, you will possess engineering-wide influence and take ownership of the following critical areas: Infrastructure Strategy & Architecture: Formulate and execute the architectural vision for Rubrik's Cloud Platform, optimizing backend infrastructure systems like Kubernetes, MySQL, and cloud-native services for performance, security, and multi-region scale. Hyperscale Automation & Platform Tooling: Build, scale, and maintain sophisticated custom internal tools, platform controllers, and automation frameworks in Go or Python to systematically eliminate operational toil. AI Infrastructure for SaaS: Deploy, scale, and operate the AI infrastructure that powers Rubrik's SaaS offerings, owning the reliability, performance, cost, and security controls required to run AI workloads in multi-tenant, compliance-bound environments. AI for SRE & Engineering Productivity : Drive the adoption of AI-driven solutions across the SRE charter to compress toil and multiply the org - applying agentic and LLM-based approaches to automated triage, incident response, operational analysis, and developer productivity. AI Adoption Guardrails for SaaS Reliability : Build the guardrails, controls, and platform patterns that keep Rubrik's SaaS reliable as AI adoption accelerates across product and engineering, ensuring new AI capabilities ship without eroding availability, performance, security, or cost posture. Cross-Functional Leadership: Wield engineering-wide influence to create technical consensus among component, platform, and security engineering teams, effectively "shifting left" to embed structural resilience, capacity guards, and compliance from initial feature designs. Reliability Governance: Define, audit, and enforce robust Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets across all critical enterprise platform services, translating telemetry insights into actionable product roadmaps during executive reviews. Incident Command & Operations Review: Serve as a primary Incident Commander for high-severity cloud outages, establishing roles, directing mitigation vectors under pressure, and orchestrating comprehensive, blameless post-mortems that drive durable systemic fixes. Cost Governance & Capacity Modeling: Architect cost-observability tools and attribution frameworks, leading cloud infrastructure capacity forecasting, resource quota optimization, and vendor SLA management. Application-SRE Leadership : Set the technical direction for the Application-SRE team, raising the bar on how the team diagnoses, mitigates, and durably resolves the most complex custo…

Salary estimate

$179,000 – $241,000/yr
Provided by the employer.

Skills for this role

PythonGOMysqlKubernetesLLMSalesLeadershipSecurityAutomation

Resume tips for Staff Software Engineer Reliability applicants

Interview preparation

Prepare concrete STAR-format stories that show Staff Software Engineer Reliability outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Staff Software Engineer Reliability problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About Rubrik Job Board

Rubrik Job Board is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles