Jobedly Post a Job

Senior Site Reliability Engineer

Akamai Technologies · Remote
RemoteFull-timeTransportationSenior$136,000–$184,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Senior Site Reliability Engineer role

Senior Site Reliability Engineer positions focus on delivering results in their domain. This page aggregates open Senior Site Reliability Engineer roles and what employers typically expect.

In this role, you will balance deep system diagnostics with progressive automation. You will serve as a premier technical authority for our platform and a champion for reducing systemic operational toil. **Do you like collaborating across teams to solve complex problems?** **Do you enjoy solving large scale distributed content delivery challenges?** **Join our critical Edge Reliability Engineering Team!** Site Reliability Engineers at Akamai leverage software engineering, systems expertise, and operational skills to deliver reliable global services. The ERE team ensures performance, resilience, and availability of Akamai's media and web delivery platform while addressing distributed systems challenges. This role acts as the top technical escalation point for critical customer-impacting issues, connecting engineering, operations, and global support teams effectively. **Partner with the best** In this role, you will balance deep system diagnostics with progressive automation. You will serve as a premier technical authority for our platform and a champion for reducing systemic operational toil. As a Senior Site Reliability Engineer, you will be responsible for: - Leading complex reliability and performance investigations across Akamai's global edge, media delivery, and web delivery platforms. - Troubleshooting critical distributed systems issues spanning application, platform, network, and operating system layers, serving as the highest technical escalation point. - Partnering with Engineering, Product, Support, and Network teams to identify root causes and deliver scalable, long-term solutions that improve platform reliability. - Designing and improving observability through SLIs, SLOs, KPIs, telemetry, dashboards, and alerts to identify and address customer-impacting issues. - Analyzing platform performance, traffic patterns, and system bottlenecks to improve scalability, resilience, and overall service reliability. - Developing automation, internal tools, AI-assisted diagnostics, and self-service workflows to streamline operations, reduce manual effort, and accelerate incident response. - Enhancing operational excellence through reliability-centered architecture reviews, post-incident analysis, continuous improvements, and offering off-hours support during critical incidents as needed. **Do what you love** To be successful in this role you will: - Possess Bachelors in CS/Engineering or a related field with 6 years of industry experience in large-scale SRE/Systems Infrastructure roles. - Have logical reasoning skills diagnosing complex performance bottlenecks, data integrity anomalies, and system failure modes in distributed environments. - Have understanding of internet technologies and foundational networking concepts, including caching, proxies, TLS, TCP/IP, DNS, and HTTP/HTTPS architectures. - Have foundation in Linux/Unix administration, diagnostic tools, and low-level environment troubleshooting. - Be able to retrieve data, analyze telemetry streams, and troubleshoot platform data integrity issues through SQL queries. - Have experience developing automation tools using languages like Python, Bash, or Go. - Demonstrate expertise in AI models and focus on implementing agentic workflows to reduce operational inefficiencies effectively. **About us** At Akamai, we make life better for billions of people, trillions of times a day. Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone. Our focus is simple: **Cloud and Edge:** Running apps closer to users for instant performance. **Security**: Neutralizing threats before they ever reach your data. **Content Delivery**: Scaling the world's biggest moments without a glitch. **AI**: Enabling…

Salary estimate

$136,000 – $184,000/yr
Provided by the employer.

Skills for this role

PythonGOSQLSecurityAutomation

Resume tips for Senior Site Reliability Engineer applicants

Interview preparation

Prepare concrete STAR-format stories that show Senior Site Reliability Engineer outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Senior Site Reliability Engineer problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About Akamai Technologies

Akamai Technologies is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles