Engineering Manager Devops positions focus on delivering results in their domain. This page aggregates open Engineering Manager Devops roles and what employers typically expect.
Maven AGI is an enterprise AI platform founded in July 2023 by executives from HubSpot, Google, and Stripe. We build conversational AI agents for autonomous customer support at scale. Our platform unifies fragmented systems, integrates knowledge sources, and enables intelligent actions without costly infrastructure changes. Our team includes talent from Google, Meta, Amazon, Microsoft, and Stripe, with advisors from OpenAI, Google, HubSpot, and Stripe. The Role We’re looking for a DevOps Manager to lead and evolve the infrastructure powering Maven AGI’s AI platform. You will manage and scale a high-performing infrastructure team while helping ensure our systems remain reliable, secure, and scalable across cloud and on-premises environments. This is a technical leadership role that combines people management, operational ownership, and strong infrastructure judgment. You will partner closely with engineering leaders and technical leads to translate business and customer requirements into clear infrastructure priorities and execution plans. You will also work directly with enterprise customers to understand complex deployment requirements, particularly for private-cloud and on-premises environments, and coordinate stakeholders across the organization to deliver sustainable solutions. Leadership and Management Responsibilities: - Manage, coach, and develop a team of DevOps and infrastructure engineers - Establish clear expectations, ownership, and accountability across the team - Partner with technical leads to align technical strategy, architecture, and execution - Hire and onboard engineers as the team grows - Lead performance management, career development, and regular feedback - Own team planning, prioritization, capacity management, and delivery - Balance reliability, security, customer commitments, and long-term platform investments - Communicate infrastructure risks, trade-offs, and progress to technical and non-technical stakeholders - Build strong partnerships across Engineering, Product, Security, and Customer Success - Improve operational processes while avoiding unnecessary overhead and reducing team toil Technical and Operational Responsibilities: - Guide the design, implementation, and operation of cloud and on-premises infrastructure across Azure, AWS, and customer-managed environments - Oversee infrastructure-as-code practices using Pulumi, Bicep, Terraform, or similar tools - Own the reliability and operation of production Kubernetes environments, including deployments, scaling, monitoring, and incident response - Drive the development and improvement of CI/CD pipelines for a large-scale monorepo - Establish consistent observability practices across metrics, logs, traces, and alerting - Advance reliability practices, including SLOs, capacity planning, disaster recovery, and runbook development - Support and scale enterprise AI deployments, including GPU infrastructure, model-serving workloads, and high-concurrency systems - Partner with engineering teams to improve developer experience, platform usability, and deployment velocity - Strengthen secrets management, access controls, and infrastructure security - Evaluate and adopt tools that improve reliability, scalability, and operational efficiency - Participate in incident response and ensure incidents lead to durable improvements Required Qualifications: - 7+ years of professional DevOps/SRE/Infrastructure experience - 3+ years of experience managing teams - Deep expertise with Kubernetes in production (AKS, EKS, or GKE) - Strong infrastructure-as-code skills (Pulumi, Terraform, or Bicep) - Experience operating CI/CD systems (GitHub Actions, ArgoCD, or Jenkins) - Proficiency in at least one scripting/programming language (Python, Go, TypeScript, or Bash) - Solid understanding of IaaS providers, networking, DNS, load balancing, and TLS - Experience with monitoring and observability stacks (Datadog, Prometheus, Grafana, or similar) - Experience with multi-cloud o…