Senior Manager AI Infrastructure Operation positions focus on delivering results in their domain. This page aggregates open Senior Manager AI Infrastructure Operation roles and what employers typically expect.
## Who We Are Vultr is on a mission to make high-performance cloud infrastructure easy to use, affordable, and locally accessible for enterprises and AI innovators around the world. With 33 global cloud data center locations, Vultr is trusted by hundreds of thousands of active customers across 185 countries for its flexible, scalable, global Cloud Compute, Cloud GPU, Bare Metal, and Cloud Storage solutions. In December 2024 Vultr announced an equity financing at a $3.5 billion valuation. Founded by David Aninowsky and self-funded for over a decade, Vultr has grown to become the world’s largest privately-held cloud infrastructure company. ## Vultr Cares - Excellent Medical Benefits w/ 100% company-paid premiums for employee only plan + 100% company-paid dental & vision premiums - 401(k) plan that matches 100% up to 4% with immediate vesting - Professional Development Reimbursement of $2,500 each year - 11 Holidays + Paid Time Off Accrual + Rollover Plan + take your birthday off - Commitment matters to Vultr! Increased PTO at 3 year & 10 year anniversary + 1 month paid sabbatical every 5 years + Anniversary Bonus each year - $500 first year remote office setup + $400 each following year for new equipment - Internet reimbursement up to $75 per month - Gym membership reimbursement up to $50 per month - Company-paid Wellable subscription **Join Vultr** We are seeking a **Senior Manager, AI Infrastructure** to lead the engineers responsible for deploying, operating, and optimizing Vultr’s AI compute clusters. In this role, you will execute the technical roadmap developed by Infrastructure leadership and ensure reliable, high-performance cluster operations at scale. This position requires a strong technical foundation, hands-on operational leadership, and the ability to coordinate complex engineering workflows. You will guide engineers through cluster deployments, provisioning, configuration management, and GPU fleet operations, ensuring the infrastructure supporting AI workloads is reliable, performant, and continuously improving. This role is ideal for someone who thrives in a fast-growing environment, enjoys solving hard infrastructure problems, and excels at driving consistent execution across a high-performing engineering team. **Key Responsibilities** - Lead the engineering team responsible for the day-to-day implementation, scaling, and operation of AI compute clusters. - Translate engineering roadmaps and technical requirements from the Director of AI Infrastructure into detailed project plans and execution milestones. - Drive delivery of cluster deployments, hardware bring-up, node configuration, and integration with orchestration and scheduling systems. - Ensure cluster reliability, uptime, and performance through monitoring, automation, and continuous operational improvements. - Oversee lifecycle operations for bare metal and GPU fleets, including provisioning, configuration management, firmware/driver updates, and hardware validation. - Manage incident response for GPU and cluster infrastructure, ensuring timely resolution and root-cause analysis. - Work closely with AI/ML, SRE, Networking, and Hardware Engineering teams to ensure cluster capabilities meet training and inference needs. - Coordinate with Product to confirm technical requirements, feature readiness, and delivery timelines. - Support integrations across networking, storage, scheduler, and resource orchestration components. - Improve tooling and automation for cluster provisioning, observability, configuration management, and large-scale fleet operations. - Contribute to the development and refinement of multi-tenant scheduling, workload management, and orchestration systems in partnership with senior technical staff. - Identify performance bottlenecks and propose engineering-level optimizations. - Coach and mentor engineers, fostering a high-performance, detail-oriented engineering culture. - Support career development, expectations, and performance managem…