Jobedly Post a Job

Senior Data Engineer, Data Lakehouse Infrastructure

trm-labs · Remote
RemoteFull-timeR&DConstruction$179,000–$241,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Senior Data Engineer Lakehouse Infrastructure role

Senior Data Engineer Lakehouse Infrastructure positions focus on delivering results in their domain. This page aggregates open Senior Data Engineer Lakehouse Infrastructure roles and what employers typically expect.

BUILD A SAFER WORLD. TRM Labs provides AI-powered intelligence solutions that help public and private sector agencies investigate and disrupt crime. TRM's platforms enable investigators to trace illicit activity, build cases, and construct operating pictures of threat networks. Leading agencies and businesses worldwide rely on TRM to make the world safer and more secure. We’re building the foundational data infrastructure powering next-generation analytics at scale. As part of our mission, we’re architecting a modern data lakehouse to support complex workloads, real-time data pipelines, and secure data governance—at petabyte scale. We are looking for a Senior Data Engineer to help us design, implement, and scale core components of our lakehouse architecture. You will have ownership over data modeling, ingestion, query performance optimization, and metadata management using cutting-edge tools and frameworks like Apache Spark, Trino, Hudi, Iceberg, and Snowflake. We’re looking for engineers with deep expertise in at least one area and a solid understanding of the trade-offs among different technologies. The impact you’ll have here: - Architect and scale a high-performance data lakehouse on GCP, leveraging technologies like StarRocks, Apache Iceberg, GCS, BigQuery, Dataproc, and Kafka. - Design, build, and optimize distributed query engines such as Trino, Spark, or Snowflake to support complex analytical workloads. - Implement metadata management in open table formats like Iceberg and data discovery frameworks for governance and observability using Iceberg compatible catalogs. - Develop and orchestrate robust ETL/ELT pipelines using Apache Airflow, Spark, and GCP-native tools (e.g., Dataflow, Composer). - Collaborate across departments, partnering with data scientists, backend engineers, and product managers to design and implement What we’re looking for: - 5+ years of experience in data or software engineering, with a focus on distributed data systems and cloud-native architectures. - Proven experience building and scaling data platforms on GCP, including storage, compute, orchestration, and monitoring. - Strong command of one or more query engines such as Trino, Presto, Spark, or Snowflake. - Experience with modern table formats like Apache Hudi, Iceberg, or Delta Lake. - Exceptional programming skills in Python, as well as adeptness in SQL or SparkSQL. - Hands-on experience orchestrating workflows with Airflow and building streaming/batch pipelines using GCP-native services. About the Team: - The Data Platform team is the funnel between all of TRM's data world and product world. We care about all layers of stack including petabyte of data stores, pipelines, and processes. - We have quite a big scope as a the team with new and exciting projects every quarter. As a result, we collaborate across the board with most teams at TRM. - We believe in async communication and are also not afraid to jump on a quick huddle if that helps to move things faster. We are both scrappy when the situation demands and also process-oriented when we need to achieve our OKRs. - We are always looking for people who can elevate the quality our tech and our execution. If you enjoy a remote-first and async friendly environment to achieve efficacy and efficiency at petabyte scale, our team could be a great pick for you! - Team members are based in the US across almost all timezones! - We do try to reserve some overlap in the day for meetings. Our north star - no IC spends more than 3-4 hours/week in meetings. Learn about TRM Speed in this position: - Build scalable engines to optimize routine scaling and maintenance tasks like create self-serve automation for creating new pgbouncer, scaling disks, scaling/updating of clusters, etc. - Enable tasks to be faster next time and reducing dependency on a single person. - Identify ways to compress timelines using 80/20 principle. For instance, what does it take to be operational in a new environment? Identify the…

Salary estimate

$179,000 – $241,000/yr
Provided by the employer.

Skills for this role

PythonSQLGCPSnowflakeSparkKafkaAirflowCommunicationETLAutomation

Resume tips for Senior Data Engineer Lakehouse Infrastructure applicants

Interview preparation

Prepare concrete STAR-format stories that show Senior Data Engineer Lakehouse Infrastructure outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Senior Data Engineer Lakehouse Infrastructure problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About trm-labs

TRM Labs is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles