Jobedly Post a Job

Principal Site Reliability Engineer

Jobgether · Remote
RemoteFull-timeFinance & InsuranceLead$153,000–$207,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Principal Site Reliability Engineer role

Principal Site Reliability Engineer positions focus on delivering results in their domain. This page aggregates open Principal Site Reliability Engineer roles and what employers typically expect.

**This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Principal Site Reliability Engineer based in Spain.** This is an exceptional opportunity to shape reliability engineering practices within a fast-growing, technology-driven environment operating at the intersection of finance and digital assets. You will play a strategic role in defining how reliability, observability, and operational excellence are embedded across engineering teams. The position combines hands-on technical leadership with broad organizational influence, enabling you to drive scalable automation and resilient infrastructure practices. Working closely with engineering, product, and platform teams, you will help build highly available systems that support mission-critical services. This role is ideal for an experienced engineer who enjoys solving complex distributed systems challenges while mentoring others and influencing engineering culture in a collaborative, globally distributed setting. ### Accountabilities: The Principal Site Reliability Engineer will lead initiatives that improve system resilience, operational efficiency, and engineering excellence across the organization. - Define and promote Site Reliability Engineering principles, establishing frameworks for reliability, observability, service level indicators (SLIs), service level objectives (SLOs), and error budgets. - Drive the adoption of operational excellence practices and ensure reliability metrics are measurable and continuously improved. - Design and implement automation solutions that enhance system scalability, reliability, and deployment efficiency. - Conduct production readiness assessments and provide architectural guidance to engineering teams to ensure services are built for scale and resilience. - Lead initiatives to improve the lifecycle of distributed systems and microservices, from development and deployment to monitoring and optimization. - Identify performance bottlenecks, capacity challenges, and operational risks while implementing sustainable solutions. - Partner with engineering and product leadership to embed reliability considerations into product development processes. - Lead incident management improvements through blameless postmortems, root cause analysis, and systemic remediation initiatives. - Mentor engineers and advocate for best practices in reliability engineering, fostering ownership and operational maturity across teams. - Contribute to the future vision and strategic direction of the Site Reliability Engineering function. ## Requirements: The ideal candidate brings deep expertise in distributed systems, reliability engineering, and organizational leadership, combined with a passion for building scalable and resilient platforms. - Proven experience designing, operating, and troubleshooting distributed systems and microservices architectures. - Strong expertise in observability, monitoring strategies, incident management, and operational excellence frameworks. - Demonstrated ability to drive organizational change and influence engineering practices across multiple teams. - Extensive experience implementing reliability frameworks, including SLIs, SLOs, error budgets, and production readiness processes. - Strong problem-solving skills with a structured and analytical approach to complex technical challenges. - Excellent communication and stakeholder management abilities, with experience collaborating across engineering and leadership teams. - Experience working with cloud environments, particularly AWS, is highly desirable. - Previous exposure to financial services, regulated industries, or mission-critical platforms is considered an advantage. - Interest in blockchain technologies, digital assets, or decentralized finance ecosystems is a plus. - Master's degree in Computer Science, Engineering, or a related field is advantageous. ## Benefits: - Fully remote opportunity within a global…

Salary estimate

$153,000 – $207,000/yr
Provided by the employer.

Skills for this role

AWSCommunicationLeadershipAutomation

Resume tips for Principal Site Reliability Engineer applicants

Interview preparation

Prepare concrete STAR-format stories that show Principal Site Reliability Engineer outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Principal Site Reliability Engineer problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About Jobgether

Jobgether is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles