All open roles

Senior ML Researcher

Bring machine learning and mathematical optimization into the decisions that shape real inference systems.

San Francisco, In-person

About Tandemn

Tandemn is an open-source, self-hosted planning and execution system for AI inference. Koi plans workload configuration and placement across heterogeneous GPU fleets, and the Tandemn execution engine connects those plans to existing serving and orchestration systems. Our team brings together researchers, systems engineers, and mathematicians working on the practical challenges of running AI workloads across cloud and on-premises infrastructure. You will work directly with the team building this system and own substantial problems from definition through implementation and evaluation. Learn more about Tandemn.

The role

Own research at the intersection of statistical learning, constrained optimization, and inference systems. You will advance Koi, Tandemn’s cluster-wide planner, developing methods that connect workload requirements and hardware behavior to deployment decisions. This is a senior, hands-on research role: you should be comfortable defining a problem, identifying its assumptions, implementing a method, and evaluating it against real serving behavior.

Responsibilities

Develop models of inference performance and use them to make workload configuration and placement decisions across heterogeneous GPUs. Work with real serving traces and systems such as vLLM or SGLang to understand batching, parallelism, memory limits, and hardware topology. Design reproducible experiments and ablations, implement methods in Python and PyTorch, and collaborate with systems engineers to put research into the planning and execution loop. Own the problem formulation, experimental methodology, and technical communication, with attention to how a method behaves when workloads or infrastructure change.

Qualifications

You bring 4+ years of ML research or applied research experience, including academic or industrial work, and have independently taken substantial projects from an initial question through implementation and evaluation. A PhD or research-focused master’s degree in computer science, statistics, applied mathematics, or a related discipline, or an equivalent research record, is expected. You are fluent in Python and PyTorch, have run experiments on GPU infrastructure, and understand probability, optimization, and experimental design. Experience with inference serving, performance modeling, causal methods, or scheduling is especially relevant. We look for evidence of original work in publications, open-source implementations, or deployed research systems, with a clear account of your individual contribution.

Apply for this role

Send your resume, relevant work, and a short introduction to hello@tandemn.com. Tell us about the work you have owned and what you would bring to Tandemn. The button below drafts an email with the role in the subject; attach your resume before sending.

Apply by email