Staff Platform Engineer Service Infrastructure positions focus on delivering results in their domain. This page aggregates open Staff Platform Engineer Service Infrastructure roles and what employers typically expect.
**About the Role** Together AI is hiring a Staff Platform Engineer to join the Product Foundations engineering organization and drive its service infrastructure strategy. Product Foundations builds and operates Together’s mission-critical product platforms that support all cloud products, including API Platform (non-Inference), web UI Platform, Billing, and customer-facing IAM. These services sit on the critical path for customers and internal systems. This is a hands-on Staff role focused on evolving Product Foundations’ core infrastructure strategy from the inside: understanding service team needs, turning repeated infrastructure problems into reusable patterns, and coordinating across platform owners so Product Foundations services are reliable, repeatable, and built on the right company-wide foundations. ## **Responsibilities** - Own the technical direction for service infrastructure within Product Foundations, including Kubernetes, AWS, Terraform, CDNs, ALBs, DNS, IAM, service networking, and related operational patterns. - Up-level existing Product Foundations services by improving reliability, operability, deployment safety, infrastructure consistency, and production readiness. - Partner deeply with API Platform and UI Platform on networking, DNS, CDN, load balancing, delivery, and gateway patterns for critical customer-facing interfaces. - Work closely with Infrastructure, Networking, and Security teams to bring company-wide platform standards into Product Foundations and contribute PF requirements back into shared frameworks. - Help drive cross-company infrastructure initiatives that Product Foundations depend on or help maintain, including Terraform CI/CD, Kubernetes networking, zero-trust service communication, policy-as-code, and cross-DC/provider networking. - Build and evolve reusable service infrastructure primitives, including Helm charts, Terraform modules, GitHub Actions/GitOps workflows, service scaffolding, runbooks, and documentation. - Establish durable technical standards through design docs, architecture reviews, mentorship, and hands-on implementation that help Together scale services across teams, regions, and cloud environments. ## **Requirements** - 7+ years of professional experience in platform engineering, service infrastructure, SRE, distributed systems, cloud infrastructure, or related roles. - Deep production experience with Kubernetes, including EKS, Helm, ArgoCD/Argo Rollouts, ingress, autoscaling, secrets, service identity, networking, and progressive delivery. - Strong Terraform experience, including module design, infrastructure CI/CD, policy enforcement, production applies, and safe self-service workflows. - Experience operating networking and edge infrastructure such as CDNs, ALBs/NLBs, DNS, TLS, ingress/egress controls, and traffic management. - Proficiency in one or more programming languages used for infrastructure tooling and automation, such as Go, Python, TypeScript, or similar. - AWS experience, ideally including EKS, IAM, VPC networking, load balancing, Route 53, CloudFront, ECR, and related service infrastructure. - Direct experience with observability systems, including metrics, logs, traces, dashboards, alerting, SLOs, and incident response. - Proven ability to lead cross-functional technical initiatives across product engineering, infrastructure, networking, and security teams. - Strong written communication skills, with experience producing clear design docs, migration plans, operational guidance, and technical standards. - Staff-level judgment: you can define ambiguous problems, make pragmatic tradeoffs, influence without authority, and leave both systems and teams better than you found them. ## **Nice to Have** - Experience building internal developer platforms or paved-path service frameworks used by many engineering teams. - Experience embedding infrastructure best practices into product engineering teams at scale. - Experience with service mesh or zero-trust infrastru…