Jobedly Post a Job

ML Compute Efficiency Automation Engineer, Infrastructure & Planning

Apple · Cupertino, CA
Full-timeTechnologyMid Level$179,000–$241,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the ML Compute Efficiency Automation Engineer Infrastructure role

ML Compute Efficiency Automation Engineer Infrastructure positions focus on delivering results in their domain. This page aggregates open ML Compute Efficiency Automation Engineer Infrastructure roles and what employers typically expect.

Apple’s Platform Acceleration & Compute Efficiency (PACE) is a high-leverage team operating at the intersection of our ML organizations, underlying compute infrastructure, and core platform tooling. Our mission is to empower Apple’s software engineering teams with efficient, scalable compute. By driving out operational friction and optimizing the broader machine learning ecosystem, we directly accelerate the pace of development for our Software and AIML organization. Foundation models are central to Apple's user experiences and maximizing the efficiency of our ML compute is paramount. Compute efficiency sits at the center of this role, ensuring that Apple’s models run as fast, reliably, and cost-effectively as possible. In this role you will tackle optimization challenges, from maximizing hardware utilization across GPUs, TPUs, and custom Apple Silicon, to shaping workload scheduling and capacity allocation for large model serving. We are looking for a particular kind of builder, an exceptional engineer who can think through hard problems and code their way past them, especially the ones involving scale and the slow manual work that quietly drains a high-leverage team. The ideal candidate treats every repeated process as a system waiting to be automated, every manual escalation as a system not yet built, and every prioritization request as a problem the right tooling can solve faster. The resulting data forms a foundation for the rapid, high-quality decisions that empower Apple's technical and business leaders. This is a founding role. The majority of your time goes to AI automation, building the systems that turn manual operations into tooling that runs and corrects itself. Your remaining time will go towards hands-on ML compute efficiency, working along side senior ML efficiency engineers directly on the optimization problems behind the numbers. You will share ownership of PACE's governance and operations with our tools team who is actively building solutions with AI. The work a traditional operations team would grind through by hand, things like resource requests, allocation tracking, escalations, and efficiency reporting, you will turn into systems that run themselves and watch themselves. When you have done it well, the busywork is gone and PACE moves faster than its size says it should. Your challenging work will result in high development velocity and efficient compute, accelerating not only Apple, but also your career as well. ## Description - Govern compute as code. Build the systems of record for resource requests, allocations, and utilization, accurate and at scale, so leadership can trust the numbers. - Hunt down ML inefficiency. Dig into inference and training workloads across GPUs, TPUs, and custom Apple Silicon, find where compute is wasted, trace it to a cause, and drive the fix. - Work the real optimization problems: scheduling, capacity allocation, and serving cost, alongside the engineers who own those systems. - Get rid of the toil. Replace the time-sink workflows, triage, reporting, reconciliation, with systems that handle the routine and pull a person in only when judgment matters. Drive manual escalations toward zero instead of standing up a tiered on call org. - Make the data useful. Build the telemetry, schemas, and anomaly detection that surface efficiency and cost opportunities, then wire them into tooling that acts rather than just files a report. - Rebuild what breaks at scale. When a process buckles under Apple scale ML demand, re-architect it so it grows with usage instead of headcount. - Make a lasting impact. Turn what you build into reusable tooling so the rest of the team benefits without coming back to you each time. ## Minimum qualifications BS in Computer Science, Computer Engineering, or equivalent practical experience. 6 or more years building production software, automation, tooling, or data and infrastructure systems. A problem solver who builds first. You have designed things from sc…

Salary estimate

$179,000 – $241,000/yr
Provided by the employer.

Skills for this role

GORESTMachine LearningLeadershipAutomation

Resume tips for ML Compute Efficiency Automation Engineer Infrastructure applicants

Interview preparation

Prepare concrete STAR-format stories that show ML Compute Efficiency Automation Engineer Infrastructure outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical ML Compute Efficiency Automation Engineer Infrastructure problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About Apple

Apple is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles