Jobedly Post a Job

Director, Engineering, Global ML Scheduling Infrastructure

Google · Sunnyvale, CA
Full-timeTechnologyExecutive$307,000–$428,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Director Engineering Global ML Scheduling Infrastructure role

Director Engineering Global ML Scheduling Infrastructure positions focus on delivering results in their domain. This page aggregates open Director Engineering Global ML Scheduling Infrastructure roles and what employers typically expect.

Minimum qualifications: Bachelor’s degree in Computer Science or equivalent practical experience. 15 years of experience in software engineering. 10 years of experience managing and leading large-scale distributed engineering teams. Experience managing international teams and driving cross-site organizational alignment. Experience leading infrastructure engineering organizations, specifically managing control planes, cluster management systems, or distributed job scheduling platforms. Preferred qualifications: Experience leading enterprise-level AI transformations, including scaling TPU/GPU accelerator infrastructure, and accelerating the transition of complex ML research innovations into high-performance, production-ready developer platforms. Ability to navigate complex matrixed organizations, influence technical strategy at the industry level, and drive convergence across legacy and modernized stacks. Domain expertise in distributed resource management, machine learning training infrastructure, hardware accelerator orchestration (GPUs/TPUs), and large-scale cloud computing platforms. Technical expertise in designing, building, and operating global-scale scheduling and orchestration systems, specifically specializing in multi-cell/multi-tenant scheduling ecosystems, throughput-oriented batch workloads, and resource optimization. About the job A core suite of systems and infrastructure manages Google's global orchestration for throughput-oriented workloads across various fleet locations, maximizing resource efficiency on a massive scale. Specializing in accelerator scheduling and massive-scale Machine Learning (ML) training, this infrastructure serves both internal Google fleets and the Google Cloud Platform (GCP). Its capabilities are continuously expanding to encompass the entire traditional compute fleet alongside the accelerator fleet. By offering workload flexibility across spatial, platform, and quota dimensions, these systems achieve exceptionally high fleet occupancy while maintaining robust usability and reliability for third-party customers and all major Google product areas. As the Director of Engineering for the Global ML Scheduling Infrastructure team, you will lead the strategic direction, engineering execution, and operational excellence of Google's global orchestration layer. Leading a distributed organization of approximately 90 engineers across the US and Poland, you will oversee the mission-critical multi-cell scheduling ecosystem that powers Google's large-scale ML training, inference, and general throughput-oriented batch workloads. Collaborating with platform, storage, data center, networking, and resource management teams, you will drive new capabilities and support the growth and efficient usage of Google's fleet. You will partner with leads from Google product areas, such as Deepmind, Search, Ads, and YouTube, to accelerate the transition of research innovations to production, with focus on developer experience and acceleration of experimentation and productionization time. You will also contribute to delivering GPUs and Google’s advanced internal technology, TPUs, to external customers via Google’s Cloud Compute Platform. You will advocate for architectural innovation, drive efficiency initiatives that directly impact Google's infrastructure footprint, and foster a high-performance culture across international sites. Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems. Individual pay is determined by factors including job-related skills, experience, and relevant education or training. US: $307000 - $428000 (USD) + 30% bonus target + equity + benefits…

Salary estimate

$307,000 – $428,000/yr
Provided by the employer.

Skills for this role

GCPMachine Learning

Resume tips for Director Engineering Global ML Scheduling Infrastructure applicants

Interview preparation

Prepare concrete STAR-format stories that show Director Engineering Global ML Scheduling Infrastructure outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Director Engineering Global ML Scheduling Infrastructure problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About Google

Google is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles