Jobedly Post a Job

Vice President Site Reliability Engineering (Data Centers)

Galaxy · Remote
RemoteFull-timeData CentersGeneral$179,000–$241,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Vice President Site Reliability Engineering role

Vice President Site Reliability Engineering positions focus on delivering results in their domain. This page aggregates open Vice President Site Reliability Engineering roles and what employers typically expect.

Who We Are: Galaxy Digital Inc. (Nasdaq: GLXY) is a global leader in digital assets and data center infrastructure, growing the economy that runs on code. Galaxy delivers the onchain infrastructure that connects institutions to digital assets, including trading, advisory, asset management, staking, self-custody, and tokenization. Galaxy also develops and operates data center infrastructure to power AI and HPC workloads. Anchored by its Helios campus in Texas, Galaxy is building a multi-gigawatt pipeline of more than 5.7 GW of potential capacity, positioning it among the largest and fastest-growing data center developers in North America. The Company is headquartered in New York City, with offices across North America, Europe, the Middle East, and Asia. Additional information about Galaxy's businesses and products is available on www.galaxy.com . What We Value: We are a diverse team of free thinkers, and fast movers united to help investors and creators energize the global economy. We are looking for individuals who thrive in a culture of builders and overachievers and embrace high performance, transparent feedback, and a mission-first approach. Our culture shapes our way of working and gets us where we want to be. Seek Excellence. Be Selective To Be Effective. Be Highly Aligned, Loosely Coupled. Disagree Transparently. Encourage Independent Decision-Making. Build Dream Teams. Who You Are A collaborative and strategic leader with deep hands-on experience in Site Reliability Engineering (SRE) and infrastructure Automation. You are comfortable steering the vision for an enterprise automation roadmap while remaining technical enough to dive into the code. You treat infrastructure as a product, ensuring that your automation workflows are as reliable as the services they deploy. You have a proven track record of managing complex hybrid environments and are proactive in building self-service platforms that enhance engineering velocity and system stability. Responsibilities Automation Platform Leadership : Oversee a specialized SRE team focused on the design, deployment, and maintenance of automation toolsets as well as the systems they interact with. Infrastructure as Code (IaC) Governance : Establish and enforce standards for IaC to ensure consistent, repeatable, and secure deployments across an entire infrastructure ecosystem. Strong proficiency in Terraform is required. Configuration Management : Lead the strategy for automated configuration and state management, ensuring Ansible playbooks and Packer image pipelines are optimized for both Windows, Linux, and ESXi Platforms. Monitoring & Observability : Manage the monitoring and health of the automation platforms themselves. Implement SLIs/SLOs to ensure the "tools that build the servers" are highly available and performant. Lifecycle Management : Drive the automated lifecycle of both physical and virtual assets, from initial template creation/deployment to automated patching, scaling, and decommissioning. Custom Tooling & Scripting : Lead the development of custom scripts and internal providers (Python, Go, PowerShell, Bash) to provide better insights and tooling for our systems. Collaboration: Outside of the automation team you will need to be able to collaborate and foster workflows alongside the rest of the Datacenter team and be able to facilitate needs for the team as a whole. Capacity & Performance : Analyze system behavior and resource utilization in virtual environments to optimize the performance of automated deployments. Mentorship & Growth : Provide technical guidance and career mentorship to SREs, fostering a culture of "automate-first" and continuous improvement. Requirements 6-10 years’ experience in Infrastructure, SRE or DevOps, specifically focused on infrastructure automation at scale. Deep proficiency with Terraform (providers, modules, state management) and Ansible (roles, playbooks, Tower/AWX). Hands-on experience with Image Creation (i.e. Packer, Ansible, SC…

Salary estimate

$179,000 – $241,000/yr
Provided by the employer.

Skills for this role

PythonGORESTTerraformLeadershipDevopsAutomation

Resume tips for Vice President Site Reliability Engineering applicants

Interview preparation

Prepare concrete STAR-format stories that show Vice President Site Reliability Engineering outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Vice President Site Reliability Engineering problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About Galaxy

Galaxy is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles