You deploy Tandemn’s planning algorithm inside your own infrastructure, on a CPU machine or designated host in your VPC or on-premises environment. It runs there as a global orchestrator above your inference stack, with a unified view of workloads, models, hardware, and available capacity across your clusters, clouds, and regions. From that view, it determines how deployments should be configured and placed together so your fleet can serve more work within each workload’s performance requirements.
The planning algorithm turns this global context into a cluster-wide deployment plan that specifies what should run, where it should run, and how it should be configured. Those plans can be adapted for any lower-level system or framework to execute through its own deployment interfaces. Whether your environment uses Kubernetes, Nvidia Dynamo, SkyPilot, Ray, or custom orchestration software, Tandemn supplies the planning layer while your existing tools continue to manage infrastructure and serve models.
Tandemn’s Kubernetes-based execution engine also runs inside your infrastructure. It translates those plans into workload launches, scaling, and configuration updates through the frameworks already in place. Teams with custom execution systems can instead connect the deployment plans to their own workflows. This lets the same global planning approach work across different execution environments without requiring every cluster to use the same tools, GPU types, or cloud provider.
As deployments run, serving telemetry feeds back into the planner so it can compare predicted and observed behavior and update the next fleet-wide plan. Models, workload data, serving telemetry, and planning decisions stay within your infrastructure. Tandemn is fully self-hosted, with no Tandemn-managed control plane and no data egress required for planning or execution. You operate the system in your environment, keeping your existing stack and adding the global orchestration needed to unlock more capacity from it.

