Staff Site Reliability Engineer positions focus on delivering results in their domain. This page aggregates open Staff Site Reliability Engineer roles and what employers typically expect.
Location: Oxford or London (Hybrid) THE MISSION: WHY WE EXIST Genomics is a science-led transatlantic TechBio combining large-scale genetic and health data with proprietary analytics to accelerate drug discovery and advance predictive, preventative healthcare. We are united by a single vision to help people live longer healthier lives, using the power of genomics. Genomics aims to help people live longer, healthier lives in two ways: super-charging drug discovery and development for novel treatments with our AI-enabled advanced genetic analytics platform, and by helping people understand their personal risk of common chronic diseases through polygenic risk scores - giving doctors and health systems the chance to get the right people into the right prevention, screening and treatment programmes at the right time. ROLE PURPOSE Genomics runs on multi-trillion-row, petabyte-scale genetic and health data. Mystra is the analytical data platform that makes that usable, and MystraAI is the agentic layer we are building on top of it. This is a Staff-level role that owns the reliability, performance, security and integrity of that infrastructure end-to-end — and sets the technical direction that other teams build on. You will lead the analytical data platform and our Data-as-a-Service offering, shape the data foundations behind MystraAI, and line-manage a small team of engineers while staying deeply hands-on. If you want your work to sit at the intersection of large-scale infrastructure and genuinely meaningful science, this is it. A DAY IN THE LIFE The role spans the full depth of the platform — from low-level infrastructure reliability and database internals up to the systems that serve data across the product. On any given day you might be: - Architecting, tuning and operating multi-tenant ClickHouse or equivalent OLAP database clusters at multi-trillion-row, petabyte scale — sharding and replication, materialised views, merge and query optimisation, tenant isolation and cost/performance trade-offs. - Owning SLOs, observability, capacity planning and incident response for data-intensive systems and pipelines, alongside the orchestration and job execution that power them. - Shaping Data-as-a-Service as a real product — a secure, multi-tenant, immutable release architecture, with data made discoverable at scale through catalogue and ontology-driven search, in partnership with the data team. - Powering agentic AI from the data platform — scaling the data foundations behind MystraAI and owning secure, tenant-scoped access for agents, partnering closely with the ML team on how our AI systems perform. - Treating security as a first-class architectural concern — tenant isolation, secrets and credential management, and least-privilege access to systems, tools and data. - Setting technical direction, mentoring and line-managing a small team of IC2/IC3 engineers, and influencing decisions well beyond your own team. WHO YOU ARE You are a recognised specialist who leads through expertise — principal-level individual-contributor depth combined with the judgement to set standards others follow. You will bring: - Large-scale OLAP / ClickHouse. Deep experience designing and tuning multi-tenant ClickHouse at extreme scale. You engage with the engine at source level rather than as a black box — and ideally have contributed code upstream. - Reliability engineering for data platforms. You bring true SRE discipline — SLOs, observability, capacity planning and incident response — to analytical data systems and pipelines. - Data-as-a-Service productisation. You think in terms of data as a product: secure, multi-tenant, immutable releases, and discoverability at scale. - Agent-ready data. You understand what it takes to expose platform data to agents safely and at scale — access patterns, tenant isolation, sensible guardrails — and have scaled data infrastructure under demanding workloads. - Multi-tenant security. You design for isolation, least-privilege…