Tandemn

Tandemn is an end-to-end cluster orchestration system for running LLM inference across heterogeneous, multi-cloud GPU fleets. At its core is Koi, a centralized planning algorithm that decides how workloads should be configured and placed across the entire cluster.

Koi observes the available hardware and every workload running on it, jointly considering each job’s requirements, priorities, and resource demands to produce a single cluster-wide deployment plan. Every five minutes, it incorporates new serving telemetry and updates that plan as traffic, capacity, and workload conditions change.

Koi is paired with Orca, Tandemn’s execution engine, which translates these plans into live infrastructure state. Built on Kubernetes, Orca works with lower-level frameworks such as NVIDIA Dynamo and Ray to execute placement, configuration, scaling, and reconfiguration decisions.

Together, Koi and Orca enable Tandemn to continuously configure, launch, observe, and update every model deployment across the cluster—autonomously.

Everything is open source and available on GitHub here.

Koi: Fleet-wide Planning.

The best way to run your GPU fleet isn't static. Workloads change. Traffic changes. Available capacity changes. And every placement decision changes the options available for the workloads that come next.

That's why Koi doesn't just optimize the job in front of it.

Koi applied a self-evolving causal relationship to learn from what's happening across your fleet, predict what's coming next, and continually determine better deployment and placement decisions. Every few minutes, Koi reevaluates the fleet and adapts configurations and workload distribution as conditions change.

That ability to learn, predict, and look beyond the next placement decision is critical. A choice that looks optimal right now can strand capacity or constrain a more important workload minutes later.

Koi optimizes not just for now, but for what comes next.

Five workloads sharing one GPU clusterStatic, illustrative demand curves for customer chat, coding assistance, finance summarization, RAG evaluations, and overnight embeddings backfill.

Mathematical Foundation

Koi is a self-calibrating algorithmic planner for inference fleets. It creates a cluster-wide deployment plan by evaluating performance, SLO risk, and cost of swapping deployment configurations.

Koi models how deployment choices affect serving outcomes:

X→V→YX\rightarrow V\rightarrow Y

Across this model, Koi tracks 146 inference parameters spanning hardware, models, runtime, and workload behavior—covering billions of possible configurations.

XX represents deployment choices that are directly configurable, VV system behavior from the GPU such as KV cache pressure, and YY outcomes such as cost, latency, throughput, and SLO performance. Because the configuration space of these variables is massive (146 params gives billions of configurations), Koi is uniquely able to run inference combinations that a human could never manually account for.

Koi updates confidence in each relationship as production evidence arrives:

p∼Beta(α,β)p\sim\mathrm{Beta}(\alpha,\beta)

Supporting evidence raises α\alpha; contradicting evidence raises β\beta. Koi uses the posterior mean and variance as its confidence and uncertainty estimates:

c=E[p]=αα+β,Var[p]=αβ(α+β)2(α+β+1)c=\mathbb E[p]=\frac{\alpha}{\alpha+\beta},\qquad \mathrm{Var}[p]=\frac{\alpha\beta}{(\alpha+\beta)^2(\alpha+\beta+1)}

As evidence accumulates, confidence improves and uncertainty falls.

Koi predicts using a calibrated serving simulator trained on profiled data, as well as roofline algorithms:

yi=(cost/token,p99 TTFT,p99 TPOT,throughput,SLO margin)\mathbf y_i=(\text{cost/token},\text{p99 TTFT},\text{p99 TPOT},\text{throughput},\text{SLO margin})

Koi scores the tradeoffs with an augmented Tchebycheff objective:

J=−[max⁡jgj+ρ∑jgj]J=-\left[\max_jg_j+\rho\sum_jg_j\right]

Here, gjg_j is the normalized gap from ideal performance. The max term protects the weakest metric; the small sum term helps break ties.

Koi also values candidates that reduce uncertainty. For a candidate ladder L′L':

EIG(L′)=∑e∈edges(L′)aeVar[pe]+wM∑M∈mechs(L′)aMVar[pM]\mathrm{EIG}(L')=\sum_{e\in\mathrm{edges}(L')}a_e\mathrm{Var}[p_e]+w_M\sum_{M\in\mathrm{mechs}(L')}a_M\mathrm{Var}[p_M]

The terms count causal relationships and mechanisms exercised by the candidate. Across the cluster, repeated tests of the same relationship are counted once:

Ae(P)=1−∏i(1−ae(Li′))A_e(P)=1-\prod_i\left(1-a_e(L_i')\right)

This keeps a plan from overvaluing duplicate tests.

Koi accounts for prediction error when assessing SLO risk:

upperj=y^j+Qq(rj)+2ϵσ^j,lowerj=y^j+Q1−q(rj)−2ϵσ^j\mathrm{upper}_j=\hat y_j+Q_q(r_j)+2\epsilon\hat\sigma_j,\qquad \mathrm{lower}_j=\hat y_j+Q_{1-q}(r_j)-2\epsilon\hat\sigma_j

The robust probability of violating an objective is bounded by:

Pr⁡DRO[gj>0]≤P^emp,j+ϵ Lipjσ^j\Pr_{\mathrm{DRO}}[g_j>0]\le\hat P_{\mathrm{emp},j}+\frac{\epsilon\,\mathrm{Lip}_j}{\hat\sigma_j}

Risk across objectives combines as:

Pr⁡DROany=1−∏j(1−Pr⁡DRO[gj>0])\Pr_{\mathrm{DRO}}^{\mathrm{any}}=1-\prod_j\left(1-\Pr_{\mathrm{DRO}}[g_j>0]\right)

Koi combines performance, learning value, SLO risk, and switching cost into one candidate score:

σi(Li′)=Ji+βt EIG(Li′)−γ Pr⁡DRO,i−λt SwitchCosti\sigma_i(L_i')=J_i+\beta_t\,\mathrm{EIG}(L_i')-\gamma\,\Pr_{\mathrm{DRO},i}-\lambda_t\,\mathrm{SwitchCost}_i

Switching cost accounts for cold starts, overlap, teardown, and transition risk.

Koi scores the full fleet plan, accounting for how each placement affects the others:

Σ(P)=∑i∈ladder actionsσi(Li′)\Sigma(P)=\sum_{i\in\mathrm{ladder\ actions}}\sigma_i(L_i')

Equivalently:

Σ(P)=∑i[Ji+βtEIG(Li′)−γPr⁡DRO,i]−λt∑iSwitchCosti\Sigma(P)=\sum_i\left[J_i+\beta_t\mathrm{EIG}(L_i')-\gamma\Pr_{\mathrm{DRO},i}\right]-\lambda_t\sum_i\mathrm{SwitchCost}_i

Koi selects the highest-value feasible plan:

Pt∈arg⁡max⁡P∈FtΣ(P)P_t\in\arg\max_{P\in\mathcal F_t}\Sigma(P)

Ft\mathcal F_t contains plans that meet capacity, deployment, SLO, and churn constraints. Active changes are capped at:

∣{i∈At:Li≠Li(t−1)}∣≤Bt\left|\{i\in A_t:L_i\ne L_i^{(t-1)}\}\right|\le B_t

Switching is both priced and limited.

After deployment, Koi compares actual results with predictions. Persistent deviations trigger a drift signal, and Koi flags a prediction when the signal crosses its threshold:

rt=observedt−predictedtr_t=\mathrm{observed}_t-\mathrm{predicted}_t
St+=max⁡(0,St−1++rt−δ),St−=min⁡(0,St−1−+rt+δ)S_t^+=\max(0,S_{t-1}^++r_t-\delta),\qquad S_t^- = \min(0,S_{t-1}^-+r_t+\delta)
St+>horSt−<−h,δ=0.5σ,h=4σS_t^+>h\qquad\text{or}\qquad S_t^-<-h,\qquad \delta=0.5\sigma,\qquad h=4\sigma

Koi checks both system behavior V and serving outcomes Y. Matching both is a repeatable success (Q1); a mismatch shows whether the cause or outcome differed (Q2–Q4).

Koi checks whether performance relationships hold across clouds, regions, and GPU types (ICP; at least three environments and 15 samples). Validation guides exploration, risk limits, and reconfiguration:

αICP=0.05,p>1−αICP⇒accept,p<αICP⇒reject\alpha_{\mathrm{ICP}}=0.05,\qquad p>1-\alpha_{\mathrm{ICP}}\Rightarrow\text{accept},\qquad p<\alpha_{\mathrm{ICP}}\Rightarrow\text{reject}

Koi reaches steady operation as results become repeatable, unnecessary changes stay low, and risk estimates match production:

qt≥1−ϵQ,ρtswap≤ϵswap,∣κt−κ∗∣≤dq_t\ge1-\epsilon_Q,\quad \rho_t^{\mathrm{swap}}\le\epsilon_{\mathrm{swap}},\quad |\kappa_t-\kappa^*|\le d

Together, this loop helps Koi continuously improve cluster-wide deployment decisions from production results.

Works With Your Existing Stack.

Tandemn is not a rip-and-replace solution. It sits above the existing inference stack as a global orchestrator, with a cluster-wide view of workloads, models, hardware types, and available capacity.

Koi uses that global context to decide where each workload should run and how it should be configured, then produces a deployment plan for downstream systems to execute. Whether your stack uses vLLM, SGLang, NVIDIA Dynamo, Ray Serve, or Kubernetes-native infrastructure, Tandemn adapts its plan to the execution environment already in place. Orca is designed to run cleanly within your existing infrastructure. You can also modify Koi’s plans however you need to work with custom systems you’ve built.

The entire open-source system runs inside your infrastructure, so you keep the tools, hardware, and serving stack you already use—while adding a global planning layer on top.

Kubernetes

NVIDIA Dynamo

Ray Serve

Your Diverse GPU Fleet Is Suddenly Your Friend.

Koi uses its knowledge of the workload DAG to plan across GPU types and families in any combination of clouds and regions. This lets you use whatever capacity is available—whether it is one GPU type in one cloud or a mix of hardware across providers and regions—instead of waiting for a specific block of capacity in one region.

Each workload has distinct compute and memory requirements. Koi matches those needs with available capacity while accounting for the serving engine, software stack, hardware configuration, and performance characteristics of each pool.

With all these combinations and dependencies, placing workloads manually is impractical. Koi evaluates how each choice affects the wider DAG and adjusts plans as capacity changes, helping each workload make the best use of the resources you can access.

AWS

Google Cloud

Azure

NVIDIA

AMD

Orca: From Plan to Running Model.

Orca turns Koi's validated plan into deployments on your existing Kubernetes infrastructure. It applies actions such as placing a workload, changing its configuration, keeping it running, or deferring it until capacity is available. Orca coordinates multi-worker launches as one rollout and carries out the plan through the infrastructure you already use.

Orca connects deployments to serving frameworks such as NVIDIA Dynamo and Ray Serve, and inference engines such as vLLM and SGLang. These systems handle model serving, while Orca manages their deployment lifecycle, including coordinated launches and configuration changes.

As workloads run, Orca records the actions taken, deployment outcomes, and performance telemetry. Koi compares those results with its predictions to see how each configuration performs in practice. That evidence informs the next planning cycle, connecting each cluster-wide plan to its execution and results.