Cloud Platform Engineer Data Reliability Backing positions focus on delivering results in their domain. This page aggregates open Cloud Platform Engineer Data Reliability Backing roles and what employers typically expect.
**About Mozn** MOZN is a leading Enterprise AI company enabling organizations to make informed decisions in two critical domains: Financial Crime Prevention and Enterprise Knowledge Intelligence. We’re a diverse, collaborative team of innovators united by a shared purpose: to build AI that delivers tangible business value, builds trust, and empowers people and organizations with augmented intelligence. Our culture is built on the relentless pursuit of excellence and meaningful impact. If you’re passionate about working alongside exceptional talent on world-class AI, and you want the autonomy and runway to do the best work of your career, join us in shaping the future of intelligent enterprises. **About the role** We are looking for a highly motivated Cloud Platform Engineer III to join our Cloud Engineering team. The ideal candidate is passionate about data reliability, performance, and scalable backing services. This role focuses on the reliability, performance, operation, automation, and continuous improvement of critical backing services such as MySQL, PostgreSQL, MongoDB, Elasticsearch/OpenSearch, Kafka, analytical databases (e.g., StarRocks, ClickHouse), and other database and messaging technologies across cloud-native and hybrid environments. This is not a traditional DBA role. The ideal candidate understands distributed systems, Kubernetes, cloud platforms, automation, IaC, observability, and AI-assisted workflows, and can help product engineering teams use backing services safely and effectively. **What you'll do** **Data Reliability & Backing Services Operations** - Own reliability, performance, scalability, and operational health of MySQL, PostgreSQL, MongoDB, Elasticsearch/OpenSearch, Kafka, StarRocks, ClickHouse, and similar platforms. - Define best practices for how product engineering teams use transactional, document, search, messaging, and analytical platforms. - Design and maintain highly available, scalable, and resilient platform services, including replication, backup, recovery, failover, and disaster recovery capabilities. - Perform capacity planning, performance tuning, workload reviews, upgrades, patching, and lifecycle management for platform services. - Identify and resolve risks such as slow queries, hot partitions, consumer lag, replication lag, index growth, retention issues, and storage saturation. - Troubleshoot and resolve complex production issues related to databases, messaging systems, search platforms, and distributed data platforms. **Kubernetes & Cloud Platform Engineering** - Hands-on experience deploying, operating, and troubleshooting stateful workloads in Kubernetes-based environments. - Strong understanding of Kubernetes fundamentals, including networking, storage, workload lifecycle management, scalability, and reliability concepts. - Enable and support Kubernetes-based deployments of database, messaging, search, and analytical platforms using cloud-native patterns and operational best practices. **Automation & Platform Enablement** - Use automation, Infrastructure as Code, GitOps, and CI/CD to make backing services repeatable, reliable, and easier to operate. - Contribute to self-service platform capabilities, guardrails, dashboards, alerts, runbooks, and production readiness checks. - Use AI-assisted workflows where appropriate for incident triage, root cause analysis, query analysis, capacity forecasting, documentation, and developer support. - Collaborate with Product Engineering, SRE, Security, Data Engineering, and Cloud Platform teams to improve reliability, performance, availability, and security posture. **Qualifications** - 4-7 years of experience in Platform Engineering, SRE, Database Reliability Engineering, Data Platform Engineering, DevOps, or related roles. - Strong hands-on experience with MySQL, PostgreSQL, Kafka, and at least one of MongoDB or Elasticsearch/OpenSearch in production environments. - Experience with analytical or distributed data platforms such as Star…