About Tandemn
Tandemn is an open-source, self-hosted planning and execution system for AI inference. Koi plans workload configuration and placement across heterogeneous GPU fleets, and the Tandemn execution engine connects those plans to existing serving and orchestration systems. Our team brings together researchers, systems engineers, and mathematicians working on the practical challenges of running AI workloads across cloud and on-premises infrastructure. You will work directly with the team building this system and own substantial problems from definition through implementation and evaluation. Learn more about Tandemn.
The role
Own research at the intersection of statistical learning, constrained optimization, and inference systems. You will advance Koi, Tandemn’s cluster-wide planner, developing methods that connect workload requirements and hardware behavior to deployment decisions. This is a senior, hands-on research role: you should be comfortable defining a problem, identifying its assumptions, implementing a method, and evaluating it against real serving behavior.
Responsibilities
Develop models of inference performance and use them to make workload configuration and placement decisions across heterogeneous GPUs. Work with real serving traces and systems such as vLLM or SGLang to understand batching, parallelism, memory limits, and hardware topology. Design reproducible experiments and ablations, implement methods in Python and PyTorch, and collaborate with systems engineers to put research into the planning and execution loop. Own the problem formulation, experimental methodology, and technical communication, with attention to how a method behaves when workloads or infrastructure change.
Qualifications
You bring 4+ years of ML research or applied research experience, including academic or industrial work, and have independently taken substantial projects from an initial question through implementation and evaluation. A PhD or research-focused master’s degree in computer science, statistics, applied mathematics, or a related discipline, or an equivalent research record, is expected. You are fluent in Python and PyTorch, have run experiments on GPU infrastructure, and understand probability, optimization, and experimental design. Experience with inference serving, performance modeling, causal methods, or scheduling is especially relevant. We look for evidence of original work in publications, open-source implementations, or deployed research systems, with a clear account of your individual contribution.
