Jobedly Post a Job

Senior Software Engineer, Data Platform

matterworks · Remote
RemoteFull-timeSoftware EngineeringTechnology$179,000–$241,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Senior Software Engineer Data Platform role

Senior Software Engineer Data Platform positions focus on delivering results in their domain. This page aggregates open Senior Software Engineer Data Platform roles and what employers typically expect.

ABOUT US Most of the molecules driving human biology are invisible to us. Mass spectrometers already detect metabolites, lipids, and peptides, but the vast majority of those signals never get identified. A typical experiment names a small fraction of its features and discards the rest. We call this biology's dark matter. It’s signal-rich and mechanism-defining, yet almost entirely opaque. Matterworks is building the foundation models that make that dark matter legible. Our Large Spectral Models do for biochemical biology what AlphaFold and ESM did for proteins: turning a library-bound discipline into something predictable and generative, and embedding it at every stage of R&D. Come build the future of biological discovery with us. Position Overview As a Senior Software Engineer you’ll work to build the connective tissue of our data platform, in both what you build and how you build it. Design, build and scale systems to enrich data from raw samples and information into readily usable datasets enriched with biological context. The data produced by you and the team will serve our customers through both our ML research and model development activities as well as our product. You will report to the Head of Engineering and work daily with our machine learning researchers, scientists, and product team. Key Responsibilities - Build and Scale Data Contracts: Own the pipelines and system other teams consume from. Design and implement systems that scale to multiple petabytes of data in an effective way. - Serving and Cost at Scale: You will build and scale systems to acquire, store and serve data in a fast and affordable as it grows: data layout, featurization throughput, Kubernetes-native orchestration, and cost surfaced before it adds up. - Labels and Enrichment: Turn raw data into datasets people can use, with consistent schemas, trustworthy metadata, and documented definitions. Scale scientific labels from studies down to their spectra and underlying features. - Quality Gates: Automate quality checks that enable increasing capability without regression and promote only on a pass. - Interfaces People Use: Own the surfaces AI, chemistry, product, and agents call, from the SDK used to build datasets to the tools that expose platform capabilities. - Operations and Data Rights: Ensure effective operations of our data needs meeting our designed service level agreements, while providing high quality, provenance, and secure data processing in line with our customer needs. About You - Significant professional experience building production data systems and pipelines. We level on scope and judgment rather than years. - Proficient in Python and SQL for large-scale data processing. - Proficient in Kubernetes-native batch orchestration and modern data lake technologies (Argo Workflows, Metaflow, EKS, Glue, Athena, Apache Iceberg, Parquet, DuckDB, Terraform). Airflow or Dagster experience transfers fine. - Demonstrated experience designing stable identifiers for a large, changing corpus, and building validation that gates a publish rather than reporting on it after the fact. - Experience putting an LLM or agent component into a production data path, including the eval loop, the gold set, and cost per record. - Daily use of AI coding tools, paired with healthy skepticism about their output on questions of production data correctness. - A track record of owning work through to a running, validated system, including the unglamorous parts: fixing the malformed dataset, writing the backfill, debugging last night's bad publish. - Comfort with messy scientific formats and toolchains (mzML, RDKit, ProteoWizard or similar). Engineering depth is the requirement. - A passion for contributing to an early-stage startup where autonomy, eagerness to learn, and enthusiasm for solving novel scientific challenges prevail over rigid processes and egos. WORKING AT MATTERWORKS Given the cross-disciplinary and innovative nature of our work, effective collaboration an…

Salary estimate

$179,000 – $241,000/yr
Provided by the employer.

Skills for this role

PythonSQLRESTKubernetesTerraformAirflowMachine LearningLLM

Resume tips for Senior Software Engineer Data Platform applicants

Interview preparation

Prepare concrete STAR-format stories that show Senior Software Engineer Data Platform outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Senior Software Engineer Data Platform problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About matterworks

matterworks is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles