ML Operations Lead positions focus on delivering results in their domain. This page aggregates open ML Operations Lead roles and what employers typically expect.
Staff / Principal MLOps Engineer Contract (6 months, potential to convert) or Full-Time | Remote (US or Canada) Come join our Data team! High velocity, high intensity, high trust, high bar, high impact, and a will to win. If those words resonate deeply with you, this could be your next career move. We're seeking someone who leads with humility, pursues audacious goals, and is motivated by meaningful impact on people and the world. At FutureFit AI, our core mission is to help more people get to better jobs faster and cheaper, with a specific focus on those facing barriers to opportunity. Our work helps resolve the growing issue of economic inequality, ensuring that no one is left behind in the future of work. Our AI-powered platform brings efficiency and insight to workforce development, replacing outdated systems and unlocking human potential at scale. Ready to make an impact? Apply today. Important note: Data shows that men typically apply when meeting 3/10 requirements, while women often wait until it's 10/10. We encourage you to apply if you see a strong (not necessarily perfect) fit. YOUR ROLE We're seeking a Staff / Principal MLOps Engineer to join our team. Our ML footprint has grown quickly alongside the business: batch models that process records in the backend, real-time models that serve recommendations to job seekers, and daily pipelines that process every available job across the US and Canada. The layer we have not yet built is the observability and traceability around all of it. Today, when a model regresses, a job fails, or a recommendation looks wrong, especially where LLMs are involved, tracing the cause and reproducing it takes far longer than it should. You will own that problem: assess our ML pipelines and data architecture with clear eyes, decide what to build and in what order, and then build it. This is a greenfield mandate, influencing production models and users. We are open to running this as a six-month contract or as a full-time hire, depending on fit and what you are looking for. WHAT YOU'LL OWN - Assessment and plan: Evaluate our current pipelines, data architecture, and ML workflows, and produce a prioritized, opinionated plan for what needs to change and why. - AI/ML observability: Architect our AI/ML observability and traceability from the ground up: model and data monitoring, regression detection, lineage, and the ability to reproduce a questionable recommendation on demand, including for LLM-based systems. - Systems design: Design data and ML systems that are anchored in customer needs and built to last, with clear tradeoffs documented so the team can build on them. - Implementation: Rebuild and harden pipelines, upgrade the data architecture, and ship the improvements. - Reliability and standards: Raise the bar on reliability and data quality, establishing the patterns and practices the rest of the team can run with. - Dependency and security hygiene: Keep the stack current and secure: framework and package upgrades across services and model images, and vulnerability remediation carried out without destabilizing production. REQUIRED EXPERIENCE - Staff or principal-level experience in MLOps, ML platform, or ML infrastructure - Experience standing up MLOps practice: CI/CD for models, experiment tracking, feature stores, and model monitoring - Experience building AI/ML observability and traceability in production: detecting regressions, diagnosing failures, tracing a prediction back to the inputs that produced it, and reproducing issues after the fact - Experience operating models in both batch and real-time serving contexts, with an understanding of how the reliability requirements differ - A track record of walking into complex, fast-grown systems, diagnosing the real problems, and materially improving them - Strong systems design ability: you can translate customer and product needs into durable, scalable architecture, write it down clearly, and stay close enough to the code to implement it…