Software Engineer Observability positions focus on delivering results in their domain. This page aggregates open Software Engineer Observability roles and what employers typically expect.
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com . We're proud to be a Living Wage accredited Employer. What You'll Do: The Observability team builds, scales, and operates the core telemetry infrastructure at CoreWeave, including our logging, metrics, and tracing platforms. We provide high-throughput ingestion pipelines and visualization systems that empower internal teams and customers to understand, troubleshoot, and optimize high-performance GPU infrastructure for the world's most demanding AI workloads. About the role: As a Software Engineer on the Observability team, you will play a pivotal role in designing, building, and owning core observability infrastructure and high-throughput telemetry pipelines. You will tackle challenges at extreme scale—supporting clusters of thousands of GPUs, petabyte-scale telemetry, and high-cardinality workloads. Applying a platform-as-a-product mindset, you will write resilient Go code, manage deployments using Kubernetes and Helm, and continually optimize service performance, security, and durability. Additionally, you will participate in an on-call rotation to support critical production environments and collaborate across engineering teams to embed telemetry best practices across our entire platform. Who You Are: 5+ years of experience in software or infrastructure engineering, with a proven track record of designing, building, and operating large-scale distributed systems in production. Proficient in Go (our primary language) or Python, with the ability to write clean, resilient, and testable production code. Hands-on production experience with Kubernetes, containerization, and microservices architectures, alongside a strong understanding of their unique observability challenges. Demonstrated experience delivering robust, scalable systems with a commitment to operational excellence, automated testing, and progressive release strategies. Ability to analyze and decompose complex problems in elastic, distributed architectures into clear, well-scoped engineering work. Comfortable working with Helm and YAML-based configuration management, templating, and infrastructure-as-code practices for service deployments. Customer-obsessed, platform-minded engineer with experience participating in on-call rotations for critical production systems. Preferred: Direct, hands-on experience designing, operating, or scaling core logging, tracing, or metrics platforms (e.g., Loki, ClickHouse, Elasticsearch, Prometheus, VictoriaMetrics, Grafana, Thanos). Familiarity with high-throughput data streaming systems (e.g., Kafka, Kafka Connect) for telemetry pipelines. Experience automating and provisioning infrastructure as part of the SDLC using tools like Terraform. Hands-on experience with OpenTelemetry for unified telemetry collection and application instrumentation. Exposure to modern AI workloads, GPU-based infrastructure, or MLOps tooling. Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk. You love to build high-throughput telemetry infrastructure and treat internal developers as key product customers. You're curious about solving petabyte-scale logging and metrics ingestion challenges across massive multi-tenant GPU clusters. You're an expert in writing resilient systems…