Data Warehouse Software Engineer Iii positions focus on delivering results in their domain. This page aggregates open Data Warehouse Software Engineer Iii roles and what employers typically expect.
Babel Street is the trusted technology partner for the world’s most advanced identity intelligence and risk operations. We deliver advanced AI and data analytics solutions providing unmatched, analysis-ready data regardless of language, proactive risk identification, 360-degree insights, high-speed automation, and seamless integration into existing systems. Babel Street empower s government and commercial organizations to transform high-stakes identity and risk operations into a strategic advantage . The actionable insights we deliver safeguard lives and protect critical assets around the world . Babel Street is headquartered in Reston, Virginia , with regional offices in Boston, MA and Cleveland, OH, and international offices in Australia, Canada, Israel, Japan, and the U.K. For more information, visit www. babelstreet.com . We are looking for an experienced Data Warehouse Software Engineer III to own, scale, and optimize our core Google Cloud Platform (GCP) data warehousing and ingestion ecosystem! In this role, you will be a key technical contributor responsible for architecting high-volume data ingestion pipelines into Google BigQuery, orchestrating complex workflow topologies in Apache Airflow (Cloud Composer), and maintaining the performance and reliability of our enterprise data warehouse. You will collaborate closely with cross-functional software, data, and analytics teams to integrate disparate commercial and government data sources into clean, highly available, and query-optimized data models. This is an experienced engineering role requiring strong Python skills, deep expertise across the entire GCP analytics stack, and hands-on experience managing and upgrading production Airflow environments. Key Responsibilities: Data Warehouse & Pipeline Engineering: Architect, build, and maintain production-grade ELT/ETL data pipelines that stream and batch high-volume, complex datasets into Google BigQuery. Airflow Management & Platform Upgrades: Manage, monitor, and upgrade our Apache Airflow / Cloud Composer platform. Write clean, modular, and maintainable DAGs in Python, implement custom operators/hooks, and optimize cluster performance. GCP Architecture & Optimization: Maintain and optimize core GCP data services (BigQuery, Cloud Storage, Cloud Functions/Run, Pub/Sub, IAM). Implement advanced BigQuery performance tuning (partitioning, clustering, slot management, and query cost control). Python Engineering & Integrations: Develop high-performance Python scripts, API connectors, and automation tools for data extraction, schema validation, and third-party data integrations. Data Quality & Integrity: Enforce strict data quality frameworks, monitoring, and automated alerting across all ingestion pipelines to ensure zero data loss and minimal latency. Technical Leadership & CI/CD: Maintain infrastructure-as-code and CI/CD deployment pipelines for data workloads using Git. Assist in architectural reviews and mentor junior engineers on data warehousing best practices. Job Requirements: U.S. Citizenship: Must be a U.S. Citizen due to federal government contract requirements. Software Engineering Experience: 4–7+ years of professional software development experience, with at least 3+ years focused heavily on data warehousing, cloud data pipelines, and data platform operations. Advanced Python Mastery: Production-grade Python development skills for data processing, custom pipeline tooling, and API integrations. Airflow Platform Operations: Hands-on experience with Apache Airflow (Cloud Composer preferred), including DAG architecture, troubleshooting, environment configuration, and platform version upgrades. GCP Ecosystem Expertise: Deep experience with the Google Cloud analytics suite—specifically BigQuery, Cloud Storage (GCS), Cloud Composer, Pub/Sub, and Cloud Functions/Run. BigQuery Optimization: Advanced understanding of BigQuery architecture, DDL/DML, partitioning, clustering, slot allocation, and query cost-control strategie…