Jobedly Post a Job

Principal ML Ops Engineer

Jobgether · Remote
RemoteFull-timeManufacturingLead$179,000–$241,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Principal ML Operations Engineer role

Principal ML Operations Engineer positions focus on delivering results in their domain. This page aggregates open Principal ML Operations Engineer roles and what employers typically expect.

**This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Principal ML Ops Engineer based in Italy.** This is a highly technical leadership opportunity to build and scale the infrastructure powering next-generation AI applications. The role focuses on designing reliable, efficient, and production-grade machine learning platforms capable of supporting large-scale AI workloads. You will take ownership of model serving systems, GPU-powered infrastructure, deployment pipelines, and operational excellence. Working closely with infrastructure, platform, and AI teams, you will help shape the foundations of a modern AI-native cloud ecosystem. This position offers the chance to solve complex distributed systems challenges while improving performance, scalability, and cost efficiency. You will play a key role in defining engineering standards and building critical ML infrastructure from the ground up. ### Accountabilities: - Design, build, and operate production-grade ML inference infrastructure using modern model serving frameworks such as vLLM, TGI, Triton, or equivalent solutions. - Develop scalable deployment pipelines supporting reliable model releases through strategies such as blue/green deployments and canary rollouts. - Build and maintain auto-scaling systems, multi-model serving architectures, and intelligent request routing mechanisms. - Optimize GPU utilization, memory efficiency, network performance, and model artifact storage to improve system reliability and cost effectiveness. - Implement observability solutions to monitor inference latency, throughput, GPU usage, operational health, and infrastructure costs. - Manage model registries, CI/CD workflows, and automation processes to enable reproducible and efficient model deployments. - Own the complete lifecycle of ML systems, from development and deployment through production operations and ongoing support. - Establish engineering best practices and contribute to platform architecture decisions in a fast-moving, remote-first environment. - Collaborate with infrastructure, platform, and applied AI teams to deliver scalable and reliable AI systems. ## Requirements: - 4+ years of experience in ML Ops, Platform Engineering, SRE, or similar infrastructure-focused roles supporting machine learning systems. - Strong hands-on experience with production model serving frameworks such as vLLM, TGI, Triton, or comparable technologies. - Proven experience operating GPU-based workloads and managing containerized environments in production. - Strong understanding of MLOps practices, including model registries, experiment tracking, automated deployment pipelines, and lifecycle management. - Proficiency in Python and infrastructure-as-code tools such as Terraform, Helm, or similar technologies. - Solid knowledge of distributed systems, performance optimization, scalability, and reliability engineering principles. - Experience using AI coding assistants to accelerate software development, troubleshooting, and debugging workflows. - Ability to work independently with strong ownership and accountability in a remote-first environment. - Experience with ML platforms such as Kubeflow, MLflow, or KubeAI is a plus. - Knowledge of GPU scheduling, CUDA/ROCm optimization, multi-tenant inference systems, and infrastructure cost optimization is advantageous. - Previous experience building greenfield infrastructure projects or working in early-stage technology environments is highly valued. ## Benefits: - Opportunity to own and shape critical ML infrastructure for a rapidly scaling AI-focused technology platform. - Fully remote working environment with flexibility to work from Romania. - Chance to build foundational systems from the ground up rather than maintaining legacy infrastructure. - Exposure to cutting-edge technologies across distributed systems, GPU computing, and large-scale AI model serving. - High level of ownership…

Salary estimate

$179,000 – $241,000/yr
Provided by the employer.

Skills for this role

PythonTerraformCi/CdMachine LearningLeadershipAutomation

Resume tips for Principal ML Operations Engineer applicants

Interview preparation

Prepare concrete STAR-format stories that show Principal ML Operations Engineer outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Principal ML Operations Engineer problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About Jobgether

Jobgether is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles