Jobedly Post a Job

Site Reliability Engineer

Axle · Frederick, MD
Full-timeInformation TechnologyHealthcare$179,000–$241,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Site Reliability Engineer role

Site Reliability Engineer positions focus on delivering results in their domain. This page aggregates open Site Reliability Engineer roles and what employers typically expect.

(ID: 2025-1135) Axle is a bioscience and information technology company that offers advancements in translational research, biomedical informatics, and data science applications to research centers and healthcare organizations nationally and abroad. With experts in biomedical science, software engineering, and program management, we focus on developing and applying research tools and techniques to empower decision-making and accelerate research discoveries. We work with some of the top research organizations and facilities in the country including multiple institutes at the National Institutes of Health (NIH). Benefits We Offer: 100% Medical, Dental & Vision Coverage for Employees Paid Time Off and Paid Holidays 401K match up to 5% Educational Benefits for Career Growth Employee Referral Bonus Flexible Spending Accounts: Healthcare (FSA) Parking Reimbursement Account (PRK) Dependent Care Assistant Program (DCAP) Transportation Reimbursement Account (TRN) The Site Reliability Engineer role centers on modernizing and consolidating a complex multi-cloud environment across AWS, Azure, and GCP, building a scalable, secure, and observable platform from the ground up using Kubernetes, AI/ML infrastructure, and zero-trust principles. You'll combine DevOps and SRE practices to support mission-driven scientific and clinical programs, emphasizing automation, reliability, compliance, and proactive monitoring while enabling innovation through AI-driven tooling. The team culture is highly collaborative and growth-oriented, valuing experimentation, continuous learning, and cross-functional leadership, with opportunities to shape future multi-cloud and platform engineering solutions. Responsibilities: Design and implement enterprise-grade monitoring and observability frameworks (metrics, logs, traces) across distributed systems using enterprise Splunk, Grafana and Open-telemetry tools Establish and manage SLIs, SLOs, and error budgets to drive reliability improvements Develop and maintain real-time asset inventory systems across cloud, on-prem, and hybrid environments Automate workload onboarding and offboarding processes, ensuring standardization and governance Track system ownership, dependencies, and lifecycle states for operational transparency Build proactive detection mechanisms using AIOps and intelligent alerting to minimize incident impact Design and operate scalable, resilient, and secure infrastructure platforms across cloud and hybrid environments Implement automated compliance tracking and enforcement aligned with organizational and regulatory standards (e.g., NIST, FISMA, FedRAMP) Embed ITIL processes (incident, change, problem, configuration management) into SRE workflows Build and maintain automated deployment environments and pipelines that enforce security, compliance, and operational standards Develop “golden paths” and standardized platform templates for consistent workload deployment Automate provisioning, patching, configuration management, and environment lifecycle Leverage AI/ML coding assistants and vibe coding practices to rapidly develop automation scripts, tools, and internal platforms Integrate AI-driven tooling into DevOps pipelines for code quality, security scanning, and operational insights Lead adoption of AI-enhanced SRE practices, including intelligent remediation and predictive operations Champion DevOps and SRE practices including Infrastructure as Code, CI/CD, observability, and reliability engineering Build developer-friendly platforms (“golden paths”) that simplify deployments, reduce friction, and improve velocity Enable and optimize infrastructure for AI/ML workloads, including data pipelines, storage systems, and inference environments, GPU-enabled and high-performance compute workloads Build and manage containerized and orchestrated platforms (Docker, Kubernetes) Support cloud migration, modernization, and platform standardization initiatives Ensure systems meet security, compliance, backup, and d…

Salary estimate

$179,000 – $241,000/yr
Provided by the employer.

Skills for this role

AWSAzureGCPDockerKubernetesCi/CdLeadershipData ScienceDevopsSecurityAutomation

Resume tips for Site Reliability Engineer applicants

Interview preparation

Prepare concrete STAR-format stories that show Site Reliability Engineer outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Site Reliability Engineer problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About Axle

Axle is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles