Tandemn is an end-to-end cluster orchestration system for running LLM inference across heterogeneous, multi-cloud GPU fleets. At its core is Koi, a centralized planning algorithm that decides how workloads should be configured and placed across the entire cluster.
Koi observes the available hardware and every workload running on it, jointly considering each job’s requirements, priorities, and resource demands to produce a single cluster-wide deployment plan. Every five minutes, it incorporates new serving telemetry and updates that plan as traffic, capacity, and workload conditions change.
Koi is paired with Orca, Tandemn’s execution engine, which translates these plans into live infrastructure state. Built on Kubernetes, Orca works with lower-level frameworks such as NVIDIA Dynamo and Ray to execute placement, configuration, scaling, and reconfiguration decisions.
Together, Koi and Orca enable Tandemn to continuously configure, launch, observe, and update every model deployment across the cluster—autonomously.
Everything is open source and available on GitHub here.
Koi: Fleet-wide Planning.
The best way to run your GPU fleet isn't static. Workloads change. Traffic changes. Available capacity changes. And every placement decision changes the options available for the workloads that come next.
That's why Koi doesn't just optimize the job in front of it.
Koi applied a self-evolving causal relationship to learn from what's happening across your fleet, predict what's coming next, and continually determine better deployment and placement decisions. Every few minutes, Koi reevaluates the fleet and adapts configurations and workload distribution as conditions change.
That ability to learn, predict, and look beyond the next placement decision is critical. A choice that looks optimal right now can strand capacity or constrain a more important workload minutes later.
Koi optimizes not just for now, but for what comes next.
Mathematical Foundation
Koi is a self-calibrating algorithmic planner for inference fleets. It creates a cluster-wide deployment plan by evaluating performance, SLO risk, and cost of swapping deployment configurations.
Koi models how deployment choices affect serving outcomes:
X→V→Y
Across this model, Koi tracks 146 inference parameters spanning hardware, models, runtime, and workload behavior—covering billions of possible configurations.
X represents deployment choices that are directly configurable, V system behavior from the GPU such as KV cache pressure, and Y outcomes such as cost, latency, throughput, and SLO performance. Because the configuration space of these variables is massive (146 params gives billions of configurations), Koi is uniquely able to run inference combinations that a human could never manually account for.
Koi updates confidence in each relationship as production evidence arrives:
p∼Beta(α,β)
Supporting evidence raises α; contradicting evidence raises β. Koi uses the posterior mean and variance as its confidence and uncertainty estimates:
c=E[p]=α+βα,Var[p]=(α+β)2(α+β+1)αβ
As evidence accumulates, confidence improves and uncertainty falls.
Koi predicts using a calibrated serving simulator trained on profiled data, as well as roofline algorithms:
The terms count causal relationships and mechanisms exercised by the candidate. Across the cluster, repeated tests of the same relationship are counted once:
Ae(P)=1−i∏(1−ae(Li′))
This keeps a plan from overvaluing duplicate tests.
Koi accounts for prediction error when assessing SLO risk:
Ft contains plans that meet capacity, deployment, SLO, and churn constraints. Active changes are capped at:
{i∈At:Li=Li(t−1)}≤Bt
Switching is both priced and limited.
After deployment, Koi compares actual results with predictions. Persistent deviations trigger a drift signal, and Koi flags a prediction when the signal crosses its threshold:
rt=observedt−predictedt
St+=max(0,St−1++rt−δ),St−=min(0,St−1−+rt+δ)
St+>horSt−<−h,δ=0.5σ,h=4σ
Koi checks both system behavior V and serving outcomes Y. Matching both is a repeatable success (Q1); a mismatch shows whether the cause or outcome differed (Q2–Q4).
Koi checks whether performance relationships hold across clouds, regions, and GPU types (ICP; at least three environments and 15 samples). Validation guides exploration, risk limits, and reconfiguration:
αICP=0.05,p>1−αICP⇒accept,p<αICP⇒reject
Koi reaches steady operation as results become repeatable, unnecessary changes stay low, and risk estimates match production:
qt≥1−ϵQ,ρtswap≤ϵswap,∣κt−κ∗∣≤d
Together, this loop helps Koi continuously improve cluster-wide deployment decisions from production results.
Works With Your Existing Stack.
Tandemn is not a rip-and-replace solution. It sits above the existing inference stack as a global orchestrator, with a cluster-wide view of workloads, models, hardware types, and available capacity.
Koi uses that global context to decide where each workload should run and how it should be configured, then produces a deployment plan for downstream systems to execute. Whether your stack uses vLLM, SGLang, NVIDIA Dynamo, Ray Serve, or Kubernetes-native infrastructure, Tandemn adapts its plan to the execution environment already in place. Orca is designed to run cleanly within your existing infrastructure. You can also modify Koi’s plans however you need to work with custom systems you’ve built.
The entire open-source system runs inside your infrastructure, so you keep the tools, hardware, and serving stack you already use—while adding a global planning layer on top.
Kubernetes
NVIDIA Dynamo
Ray Serve
Your Diverse GPU Fleet Is Suddenly Your Friend.
Koi uses its knowledge of the workload DAG to plan across GPU types and families in any combination of clouds and regions. This lets you use whatever capacity is available—whether it is one GPU type in one cloud or a mix of hardware across providers and regions—instead of waiting for a specific block of capacity in one region.
Each workload has distinct compute and memory requirements. Koi matches those needs with available capacity while accounting for the serving engine, software stack, hardware configuration, and performance characteristics of each pool.
With all these combinations and dependencies, placing workloads manually is impractical. Koi evaluates how each choice affects the wider DAG and adjusts plans as capacity changes, helping each workload make the best use of the resources you can access.
AWS
Google Cloud
Azure
NVIDIA
AMD
Orca: From Plan to Running Model.
Orca turns Koi's validated plan into deployments on your existing Kubernetes infrastructure. It applies actions such as placing a workload, changing its configuration, keeping it running, or deferring it until capacity is available. Orca coordinates multi-worker launches as one rollout and carries out the plan through the infrastructure you already use.
Orca connects deployments to serving frameworks such as NVIDIA Dynamo and Ray Serve, and inference engines such as vLLM and SGLang. These systems handle model serving, while Orca manages their deployment lifecycle, including coordinated launches and configuration changes.
As workloads run, Orca records the actions taken, deployment outcomes, and performance telemetry. Koi compares those results with its predictions to see how each configuration performs in practice. That evidence informs the next planning cycle, connecting each cluster-wide plan to its execution and results.