Jobedly Post a Job

Software Engineer, GPU Infrastructure- ChatGPT Engineering

openai · Remote
RemoteFull-timeApplied AITechnology$102,000–$138,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Software Engineer Gpu Infrastructure Chatgpt Engineering role

Software Engineer Gpu Infrastructure Chatgpt Engineering positions focus on delivering results in their domain. This page aggregates open Software Engineer Gpu Infrastructure Chatgpt Engineering roles and what employers typically expect.

About the Team ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance. As our GPU fleet continues to grow, we're investing in the infrastructure that operates it. Our team builds the tooling, automation, and intelligent systems that make GPU infrastructure scalable, observable, and increasingly autonomous. We work across production engineering, distributed systems, capacity management, and AI-powered operational tooling to help researchers and product teams move faster while maximizing the efficiency of every GPU. This is a unique opportunity to work on infrastructure at the frontier of AI, where small improvements in fleet efficiency, reliability, and automation have an outsized impact on the development and deployment of AGI. About the Role We're looking for a Software Engineer with deep experience operating large-scale GPU or compute infrastructure. You'll design and build the systems that manage GPU clusters at scale—from fleet health and capacity planning to operational automation and intelligent agents that reduce manual intervention. You'll partner closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and overall compute utilization. This role is ideal for engineers who enjoy solving complex operational challenges, building internal platforms, and working on infrastructure that directly powers frontier AI. In This Role, You Will - Design, build, and operate software that manages large-scale GPU infrastructure supporting ChatGPT inference. - Build internal platforms, tooling, and AI-powered agents that automate fleet operations and reduce operational overhead. - Improve observability, reliability, and operational efficiency across thousands of GPUs. - Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response. - Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance. - Partner closely with research, platform, networking, and systems teams to continuously improve our compute platform. - Help establish engineering best practices around operational excellence, automation, and infrastructure reliability. You Might Thrive in This Role If You - Have experience operating large-scale production infrastructure, preferably GPU clusters or other compute-intensive distributed systems. - Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering. - Have built software that automates operational workflows rather than relying on manual processes. - Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure. - Understand infrastructure observability, monitoring, capacity planning, and incident management. - Enjoy identifying cross-team pain points and building reusable platforms that improve developer productivity. - Are comfortable working across software engineering and systems operations, owning problems end-to-end. - Thrive in fast-moving environments with significant technical ambiguity. Qualifications - 5+ years of software engineering experience building production infrastructure. - Strong programming skills in Go, Python, C++, Rust, or similar systems languages. - Experience designing and operating highly available distributed systems. - Experience with GPU infrastructure, high-performance computing, ML infrastructure, or large-scale compute platforms. - Experience with Kubernetes, cloud infrastructure, Linux, networking, and observability tooling. - Excellent debugging, systems design, and operational problem-solving skills. - Strong communication skills and experience collaborating across engineering organizations. About OpenAI OpenAI is an AI research and deploy…

Salary estimate

$102,000 – $138,000/yr
Provided by the employer.

Skills for this role

PythonC++GORUSTKubernetesCommunicationAutomation

Resume tips for Software Engineer Gpu Infrastructure Chatgpt Engineering applicants

Interview preparation

Prepare concrete STAR-format stories that show Software Engineer Gpu Infrastructure Chatgpt Engineering outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Software Engineer Gpu Infrastructure Chatgpt Engineering problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About openai

openai is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles