All open roles

Senior Kubernetes Systems Engineer

Build the reliable systems that turn inference plans into production deployments across GPU fleets.

San Francisco, In-person

About Tandemn

Tandemn is an open-source, self-hosted planning and execution system for AI inference. Koi plans workload configuration and placement across heterogeneous GPU fleets, and the Tandemn execution engine connects those plans to existing serving and orchestration systems. Our team brings together researchers, systems engineers, and mathematicians working on the practical challenges of running AI workloads across cloud and on-premises infrastructure. You will work directly with the team building this system and own substantial problems from definition through implementation and evaluation. Learn more about Tandemn.

The role

Own the systems that translate inference plans into reliable production operations. You will work on Tandemn’s execution engine and Kubernetes integrations across heterogeneous GPU infrastructure, connecting scheduling decisions with workload lifecycle management and observed serving performance. We are looking for an experienced engineer who has built and operated distributed systems, understands Kubernetes beyond deployment manifests, and can reason carefully about failure, concurrency, and operational tradeoffs.

Responsibilities

Design and build the controllers, custom resources, and reconciliation paths that turn workload plans into dependable deployments. Own how workloads start, reconfigure, recover, and report their state when nodes fail or infrastructure changes. Integrate with Kubernetes scheduling, GPU allocation, and inference engines, and build the metrics and diagnostics needed to understand real serving behavior. Carry changes through integration testing, rollout, and production operation, working directly with researchers and users to resolve failures across the control plane, networking, containers, and GPU stack. Help set clear interfaces and invariants that keep the system understandable and reliable.

Qualifications

You bring 4+ years of systems or infrastructure engineering experience and have designed, built, and operated distributed systems in production. You have hands-on experience with Kubernetes controllers or operators, custom resources, reconciliation, scheduling, and workload lifecycle management, rather than only deploying applications to a cluster. You are fluent in Go or Python and can reason about concurrency, state consistency, retries, and partial failure. Practical experience with Linux, container runtimes, networking, observability, and at least one of AWS, GCP, or Azure is expected, along with ownership of incident response and safe releases. GPU workloads, NVIDIA GPU Operator, KubeRay, KServe, vLLM, SGLang, or multi-cluster infrastructure are especially relevant. Be prepared to discuss a system you owned and the engineering decisions that made it reliable.

Apply for this role

Send your resume, relevant work, and a short introduction to hello@tandemn.com. Tell us about the work you have owned and what you would bring to Tandemn. The button below drafts an email with the role in the subject; attach your resume before sending.

Apply by email