Jobedly Post a Job

Sr. Software Engineer - Inference Engine (Platform Software)

furiosa-ai · Seoul HQ
Full-timeSoftwareTechnology$136,000–$184,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Senior Software Engineer Inference Engine role

Senior Software Engineer Inference Engine positions focus on delivering results in their domain. This page aggregates open Senior Software Engineer Inference Engine roles and what employers typically expect.

ABOUT THE JOB Software Engineer (Inference Engine) is responsible for developing and optimizing a high-performance inference engine for Large Language Models (LLMs) and multimodal LLMs running on FuriosaAI NPUs. In this role, you will proactively research and apply the state-of-the-art inference optimization techniques to our inference engine. You will work in close collaboration with the compiler and hardware teams to enhance the engine's performance to its full potential. RESPONSIBILITIES - Design and implement FuriosaAI’s next-generation inference engine for large and multimodal language models—comparable in capability to frameworks such as vLLM and SGLang—optimized for throughput, latency, and memory efficiency. - Design and implement advanced inference optimizations—such as speculative decoding, KV-cache management, tensor/model parallelism, memory-efficient execution, and scheduling—in our production inference engine. - Design and develop capabilities for distributed and scalable inference, including prefill–decode (PD) and encode–prefill–decode (EPD) disaggregation, disaggregated speculative decoding, and hierarchical and external KV-cache storage such as HiCache and Mooncake. - Collaborate closely with the Compiler team to co-design and optimize execution for FuriosaAI NPUs, improving system-level throughput, latency, and memory utilization. - Proactively research, evaluate, and integrate state-of-the-art inference optimization techniques and key features of LLM serving frameworks into our production inference engine. MINIMUM QUALIFICATIONS - BS degree in Computer Science, Engineering, or a related field, with at least 3 years of relevant industry experience, or equivalent practical experience - Proficiency in Rust or C++ programming skill - Knowledge and passion of deep learning, LLM, and/or generative AI models - Excellent problem-solving and data analysis skills. - Strong communication and collaboration skills. PREFERRED QUALIFICATIONS - Experience in building inference serving systems for large models, encompassing batching, scheduling, caching, and load balancing. - A deep understanding of performance optimization systems. - Proficiency in C++/CUDA or Triton kernel development - Contributions to open-source inference frameworks such as vLLM, SGLang, or TensorRT-LLM. CONTACT - recruit@furiosa.ai

Salary estimate

$136,000 – $184,000/yr
Provided by the employer.

Skills for this role

C++RUSTLLMCommunicationData Analysis

Resume tips for Senior Software Engineer Inference Engine applicants

Interview preparation

Prepare concrete STAR-format stories that show Senior Software Engineer Inference Engine outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Senior Software Engineer Inference Engine problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About furiosa-ai

furiosa-ai is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles