Jobedly Post a Job

Senior AI Reliability Engineer (Platform)

Flatiron Health · London office
Full-timeEngineeringGeneral$136,000–$184,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Senior AI Reliability Engineer role

Senior AI Reliability Engineer positions focus on delivering results in their domain. This page aggregates open Senior AI Reliability Engineer roles and what employers typically expect.

We’re looking for a Senior AI Reliability Engineer (Platform) to help us accomplish our mission to improve and extend lives by learning from the experience of every person with cancer. Are you ready to be the next changemaker in cancer care? Flatiron Health is a healthtech company using data for good to power smarter care for every person with cancer, around the world. Flatiron partners with cancer centers in the US, Europe and Asia to transform patients’ real-life experiences into real-world evidence and create a more modern, connected oncology ecosystem. Our multidisciplinary teams include oncologists, data scientists, software engineers, epidemiologists, product experts and more. Flatiron Health is an independent affiliate of the Roche Group. What You’ll Do We’re seeking a Senior AI Reliability Engineer (Platform) to help Flatiron safely and effectively scale AI-enabled workflows across our engineering, product, and business teams. This role sits at the intersection of data science, platform engineering, AI evaluation, and production reliability. As Flatiron’s use of AI grows, we need to move beyond experimentation and build the systems, standards, and feedback loops that allow AI workflows to be evaluated, monitored, trusted, and improved over time. Platform’s strategy is to enable AI adoption without becoming a gatekeeper: building reusable patterns, evaluation infrastructure, observability, and guardrails that help teams move quickly while managing reliability, safety, and cost. As a Senior AI Reliability Engineer (Platform), you will: Design, build, and continuously improve evaluation frameworks, benchmarks, and automated testing pipelines for AI, LLM-powered, and agentic workflows. Define and monitor quality, reliability, safety, performance, and cost metrics for AI systems, including observability, drift detection, hallucination risk, retrieval quality, and end-to-end workflow behaviour. Develop reliability engineering practices for AI-enabled systems, including SLOs, SLIs, monitoring, alerting, incident response, runbooks, and root-cause analysis of AI failure modes. Design orchestration, governance, and guardrails for multi-agent AI systems, including agent coordination, permissions, auditability, human oversight, and secure deployment patterns. Partner with platform, product, security, engineering, and data science teams to evaluate AI solutions, establish reusable standards, and guide build-vs-buy, model selection, and AI adoption decisions. Support experimentation with emerging AI technologies while helping the organisation make pragmatic, scalable decisions in a rapidly evolving landscape, collaborating across global teams and participating in on-call rotations. Who You Are You’re a senior technical practitioner with experience working across data science, machine learning, software engineering, platform engineering, or reliability engineering. You are comfortable operating in ambiguous spaces where the right answer is not always obvious, and you are motivated by turning emerging AI capabilities into production-ready systems that teams can actually trust. You understand that AI systems fail differently from traditional software. A model may not crash, but it may silently degrade, become less accurate, respond inconsistently, produce poor outputs, or create business risk in ways that are hard to detect without the right evaluation and observability patterns. This role is focused on that production behaviour and system health, not on pure model research or training. You likely have: 5+ years of experience in platform engineering, SRE, machine learning, MLOps or a related technical field, with strong Python skills and experience building production-quality systems. Solid experience with the following AWS, Bedrock, AgentCore, ML Flow, Databricks, Cloudwatch, Gitlab Experience designing experiments, evaluation frameworks, statistical analyses, and quality metrics for ML or AI systems, with familiarity in LLMs, RAG,…

Salary estimate

$136,000 – $184,000/yr
Provided by the employer.

Skills for this role

PythonAWSMachine LearningLLMData ScienceSecurity

Resume tips for Senior AI Reliability Engineer applicants

Interview preparation

Prepare concrete STAR-format stories that show Senior AI Reliability Engineer outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Senior AI Reliability Engineer problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About Flatiron Health

Flatiron Health is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles