Jobedly Post a Job

Inference Engineer

designworkstalent · Remote
RemoteFull-timeEngineeringTechnology$187,000–$253,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Inference Engineer role

Inference Engineer positions focus on delivering results in their domain. This page aggregates open Inference Engineer roles and what employers typically expect.

INFERENCE ENGINEER Location: Hybrid | Bellevue, WA Area Titles: Senior and Staff (multiple roles available) BUILD THE INFERENCE PLATFORM POWERING NEXT-GENERATION AI APPLICATIONS ABOUT THE OPPORTUNITY A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads—including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI applications. Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging modern tooling and automation to build infrastructure capable of supporting the industry's most demanding AI workloads. We're seeking Inference Engineers to build and operate the model-serving systems behind a next-generation AI inference platform. This team focuses on delivering high-throughput, low-latency, reliable inference experiences that enable customers to consume advanced AI capabilities through production-scale APIs. THE OPPORTUNITY This is a foundational engineering role focused on building the systems that bring AI models from research environments into reliable production services. You'll work on the infrastructure layer responsible for serving large models efficiently, optimizing performance, and ensuring reliability as usage scales. You'll collaborate closely with GPU performance, AI training infrastructure, platform engineering, and operations teams to solve complex challenges around model serving, latency optimization, resource efficiency, and production reliability. This opportunity is ideal for engineers who enjoy working at the intersection of distributed systems, machine learning infrastructure, GPU computing, and large-scale production systems. WHAT YOU'LL DO - Build and operate production-grade model-serving and inference systems supporting high-throughput, low-latency AI workloads. - Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency across different model architectures and workloads. - Design systems that maximize GPU utilization while maintaining predictable performance and reliability. - Improve the scalability and operational maturity of inference platforms as customer demand grows. - Partner with AI training, GPU performance, orchestration, and infrastructure teams to ensure smooth transitions from model development to production serving. - Develop monitoring, alerting, and operational practices to maintain reliable inference services. - Investigate and resolve performance, reliability, and capacity challenges across inference workloads. - Contribute to architecture decisions and engineering standards as the platform evolves. WHAT WE'RE LOOKING FOR - Experience building and operating production machine learning inference or model-serving systems at scale. - Strong understanding of the performance trade-offs involved in serving large AI models, including latency, throughput, memory utilization, and cost efficiency. - Experience designing reliable distributed systems or production infrastructure. - Understanding of GPU-backed AI workloads and the challenges of scaling inference systems. - Strong engineering fundamentals and the ability to independently own complex technical problems. - Comfortable working in a fast-moving environment where systems and processes are being built from the ground up. PREFERRED QUALIFICATIONS - Experience with modern inference-serving frameworks such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or similar technologies. - Experience optimizing LLM inference workloads or large-scale AI serving platforms. - Background operating API-ba…

Salary estimate

$187,000 – $253,000/yr
Provided by the employer.

Skills for this role

Machine LearningLLMAutomation

Resume tips for Inference Engineer applicants

Interview preparation

Prepare concrete STAR-format stories that show Inference Engineer outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Inference Engineer problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About designworkstalent

designworkstalent is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles