Jobedly Post a Job

Site Reliability Engineer II

Zuora · Costa Rica
Full-timeEngineering & TechOpsGeneral$187,000–$253,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Site Reliability Engineer II role

Site Reliability Engineer II positions focus on delivering results in their domain. This page aggregates open Site Reliability Engineer II roles and what employers typically expect.

Costa Rica About Zuora At Zuora, we help businesses grow smarter and adapt faster. Our platform powers modern business models — from subscriptions and usage-based pricing to AI-driven and outcome-based offerings — helping companies launch new products, automate complex billing, and unlock predictable, recurring revenue. We’ve led the Subscription Economy for more than a decade. Now we’re evolving again by building the definitive platform for quote to cash and helping companies monetize their products and services with an adaptable, AI-ready foundation. This is a location-specific position that requires you to come into the office regularly to be most effective. * Zuora Costa Rica office (Heredia): hybrid model with 3 days in office and 2 days remote. The Opportunity Join Zuora's high-impact Operations team and help power the backbone of our industry-leading SaaS platform. In this role, you'll help ensure the reliability, scalability, and performance of Zuora's global production environment while building the next generation of intelligent operations. We're looking for an engineer who enjoys solving complex infrastructure challenges, embraces an automation-first mindset, and is excited about applying AI and modern cloud technologies to improve operational excellence. You'll have the opportunity to: Design and implement intelligent automation for infrastructure lifecycle management, including self-healing, anomaly detection, and automated remediation using Infrastructure as Code (IaC) and AI-driven tooling. Apply AI/ML techniques for predictive monitoring and proactive performance optimization to identify issues before they impact customers. Lead complex incident response efforts and root cause analyses, embedding automation and continuous learning into operational processes. Improve system reliability through dynamic scaling, telemetry instrumentation, and automated performance tuning. Enhance operational runbooks and playbooks by eliminating manual processes through automation. Evaluate and adopt emerging AIOps, cloud-native, and distributed systems technologies to continuously improve our platform. Partner cross-functionally with Product Engineering, Customer Support, Global Services, Deal Desk, and Sales to deliver exceptional customer experiences. Our technology stack includes Linux, Python, Docker, Kubernetes, AWS, Kafka, ActiveMQ, MySQL, Oracle, Redis, Tomcat, Jenkins, Terraform, GitOps, Ansible, Puppet, Prometheus, Grafana, OpenTelemetry, Debezium, Web Application Firewalls, and Load Balancers. About You You're passionate about building reliable systems, automating repetitive work, and continuously improving how infrastructure operates. You enjoy troubleshooting complex technical problems, collaborating across teams, and learning new technologies. We're looking for someone with: 2–4 years of experience in Linux systems administration and/or Python development in production environments. Strong Linux administration skills, including troubleshooting, service management, performance tuning, and networking fundamentals. Experience developing Python scripts or lightweight applications to automate operational workflows and system management. Hands-on experience with Docker and familiarity with Kubernetes concepts, including deployments, services, and scaling. At least one year of experience supporting SaaS or cloud-native production environments. Working knowledge of messaging platforms and databases such as Kafka, Redis, MySQL, or similar technologies. Experience contributing to CI/CD pipelines and deployment automation. Hands-on experience with monitoring and observability platforms such as Prometheus, Grafana, or similar tools. Experience participating in incident response, post-incident reviews, and root cause analysis. A demonstrated passion for automation and improving operational efficiency. Nice to have: Experience with Jenkins, Terraform, GitOps, or advanced Infrastructure as Code practices. Exposure to AI/ML technol…

Salary estimate

$187,000 – $253,000/yr
Provided by the employer.

Skills for this role

PythonMysqlRedisAWSDockerKubernetesTerraformCi/CdKafkaSalesAutomation

Resume tips for Site Reliability Engineer II applicants

Interview preparation

Prepare concrete STAR-format stories that show Site Reliability Engineer II outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Site Reliability Engineer II problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About Zuora

Zuora is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles