Jobedly Post a Job

Forward Deployed Engineer - SRE

andromeda · Remote
RemoteFull-timeEngineeringGeneral$179,000–$241,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Forward Deployed Engineer Sre role

Forward Deployed Engineer Sre positions focus on delivering results in their domain. This page aggregates open Forward Deployed Engineer Sre roles and what employers typically expect.

FORWARD DEPLOYED ENGINEER - SRE LOCATION: NORTH AMERICA REMOTE/SF-HYBRID · FULL-TIME About Andromeda Andromeda gives AI companies access to the kind of scaled compute once reserved for hyperscalers. Our platform connects 100+ AI customers to 50+ global providers, with billions of GPU-hours supported, and those numbers are all rapidly growing. We combine enterprise-grade reliability with the speed and economics of an open market, serving teams running everything from large-scale training to production inference. Nat Friedman (former CEO of GitHub) and Daniel Gross (former head of AI at Apple, YC partner) started Andromeda in 2023 with a single GPU cluster. It filled almost immediately. Three years later, we're a $1.5B company, profitable since day one, with a Series A from Paradigm to scale the platform globally. The global flow of compute is already a multi-trillion dollar market, and our team is building the infrastructure that enables it to continue to scale. The problem is deceptively hard. Not all compute is equal: interconnect, networking, OEM, firmware, and cluster age all vary across providers, and the differences matter at scale. Our platform benchmarks and validates capacity, takes positions, structures contracts, and operates clusters globally, delivering a consistent product regardless of where it runs. No one else has built this layer, and the AI industry can't scale without it. The Role This is not a generalist SRE role, and it is not a support role. You will embed directly with the teams running large-scale training and inference on our clusters. You are responsible for onboarding them, tuning their jobs, and debugging their failures alongside them, while owning the infrastructure and automation that makes those clusters reliable in the first place. Forward deployed means you spend real time inside customer environments: reading their training code, sitting in their Slack channels, watching their runs, and shipping fixes that land in our platform. When a multi-hundred-GPU run stalls, you are the person who figures out whether it's the fabric, the driver, the scheduler, or their dataloader, then you make sure it can't happen the same way twice. We're looking for engineers who have personally run GPU clusters in production, understand the failure modes of distributed training, and can reason about performance from network fabric → kernel → framework. Equally important: you can explain what you found to someone else's engineering team without condescension, and turn that conversation into a product improvement. What You’ll Do - Serve as the primary technical point of contact for teams running large-scale training and inference workloads. Own onboarding end to end; environment setup, orchestration choice (Slurm, Kubernetes, or direct SSH), storage layout, first successful run at scale. You will continue to stay engaged as their workloads grow. - Work inside customer environments to diagnose real failures: NCCL timeouts, stragglers, checkpoint I/O stalls, degraded links, OOM patterns, container and driver mismatches. Read their code when you need to. Reproduce, isolate, fix, and write it down. - Profile and improve distributed training performance on live workloads. Improving MFU, cutting idle GPU time, and reducing time-to-first-successful-run for new deployments. - Own reliability outcomes for the accounts you're deployed on. - Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) that underpin distributed training. Diagnose and resolve fabric-level issues that degrade collective operations. - Build deep visibility into GPU utilization, memory pressure, interconnect throughput, job performance, and hardware health. - Turn every repeated deployment problem into automation: cluster provisioning, GPU health checks and burn-in, preflight validation, self-healing, firmware/driver lifecycle management, and reusable reference configurations for common training and serving stacks. - Lead…

Salary estimate

$179,000 – $241,000/yr
Provided by the employer.

Skills for this role

KubernetesAutomation

Resume tips for Forward Deployed Engineer Sre applicants

Interview preparation

Prepare concrete STAR-format stories that show Forward Deployed Engineer Sre outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Forward Deployed Engineer Sre problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About andromeda

andromeda is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles