Tandemn is an open-source, algorithmic planner for self-hosted AI inference that unlocks more usable capacity from your existing GPU fleet. It places and configures workloads across heterogeneous hardware, clouds, and regions so you can run more jobs and serve more demand without expanding your infrastructure.
It brings fragmented GPU capacity across clusters, cloud providers, and serving regions into one cohesive pool for planning. Different GPU types and hardware configurations retain their individual capabilities, while Tandemn matches each workload to suitable capacity wherever it is available across the fleet.
The algorithm models inference end to end, connecting model configurations and hardware characteristics to runtime behavior and serving performance. It uses this understanding to navigate a space of billions of possible deployment configurations and continuously refines its planning from observed results. As workloads and demand change, it jointly selects configurations and placements that make better use of available resources while meeting each job’s serving requirements. The result is higher goodput and greater effective fleet capacity: more useful inference work from the GPUs you already have.

