Jobedly Post a Job

Senior Staff LLM Inference Engineer

d-matrix · Remote
RemoteFull-timeArchitectureTechnology$179,000–$241,000/yr
Apply on Jobedly ⚡ One-click AI Apply

About the Senior Staff Llm Inference Engineer role

Senior Staff Llm Inference Engineer positions focus on delivering results in their domain. This page aggregates open Senior Staff Llm Inference Engineer roles and what employers typically expect.

At d-Matrix, we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible. Our culture is one of respect and collaboration. We value humility and believe in direct communication. Our team is inclusive, and our differing perspectives allow for better solutions. We are seeking individuals passionate about tackling challenges and are driven by execution. Ready to come find your playground? Together, we can help shape the endless possibilities of AI. D-Matrix Frontier Group sits at the leading edge of what’s possible with LLM inference on heterogeneous hardware. Our charter spans the full stack: from pathfinding emerging use cases and novel deployment patterns to deep optimization of inference kernels, to building proof-of-concept systems that showcase D-Matrix’s unique computational fabric. We are an applied research and engineering team that moves fast, ships real systems, and works directly with product and hardware teams to shape the roadmap. We build the tools, runtimes, and frameworks that let frontier AI models run efficiently and cost-effectively across heterogeneous deployments — combining D-Matrix silicon with CPUs, GPUs, and custom accelerators. Our work powers everything from benchmarking and evaluation pipelines to production-grade inference serving. This Role We are hiring end-to-end inference engineers who are comfortable going from a novel research idea to a deployed, optimized system. You will work at every layer of the inference stack — from kernel-level optimization to distributed orchestration to high-level serving APIs. This role could be a great match for you if you: • Have deep intuition for modern generative AI architectures and how to squeeze performance out of them at inference time. • Are familiar with the internals of open-source inference frameworks (vLLM, SGLang, TensorRT-LLM, etc.) and can extend or replace them when needed. • Enjoy pathfinding new use cases — exploring heterogeneous deployment topologies and building early-stage POCs that prove out new ideas. • Are results-oriented with a strong bias toward action; you own problems end-to-end from prototype to optimization to handoff. • Are energized by working at the intersection of novel hardware and frontier models, and want your work to directly influence how next-generation AI silicon is used. • Value clear communication and thrive in a small, high-ownership team environment. Responsibilities • Identify and prototype emerging LLM inference use cases suited to heterogeneous hardware deployments. • Build compelling proof-of-concept systems that demonstrate D-Matrix capabilities to customers, partners, and internal stakeholders. • Develop and tune custom kernels and operator-level optimizations to maximize throughput and minimize latency. • Drive quantization, sparsity, and batching strategies tailored to D-Matrix computational model. • Build and maintain inference runtimes, serving frameworks, and evaluation tooling. • Contribute to distributed inference systems: tensor/pipeline parallelism, disaggregated prefill/decode, KV-cache management. • Work closely with hardware architects to provide firmware and compiler teams with actionable inference workload insights. • Partner with product and business development to translate POCs into customer-facing demonstrations. • Contribute to technical publications, whitepapers, and open-source projects that advance D-Matrix visibility. Required Qualifications • Bachelor’s degree in Computer Science, Electrical Engineering, or a related field, and 10+ years of relevant engineering experience; or equivalent demonstrated experience. • Master’s or PhD in Computer Science, Electrical Engineering, or a related field preferred, with 6+ years of relevant industry experience. • Strong proficiency in Python and C/C++. • Hands-on experience optimizing LLM inference — attenti…

Salary estimate

$179,000 – $241,000/yr
Provided by the employer.

Skills for this role

PythonC++LLMCommunication

Resume tips for Senior Staff Llm Inference Engineer applicants

Interview preparation

Prepare concrete STAR-format stories that show Senior Staff Llm Inference Engineer outcomes you drove.

Research the employer's product and recent news before the interview.

Be ready to explain how you'd approach a typical Senior Staff Llm Inference Engineer problem end to end.

Have thoughtful questions ready about the team, tools and success metrics.

About d-matrix

d-matrix is actively hiring on Jobedly. Explore their open roles and what it's like to work there.

Apply on Jobedly ⚡ One-click AI Apply

Similar jobs

Companies hiring for similar roles