# Tandemn > Tandemn is an open-source, self-hosted planning and execution system for AI inference. It helps operators run more online and batch workloads on their existing GPU fleets by configuring and placing workloads across heterogeneous hardware, clouds, regions, and on-premises infrastructure. Operators provide the model, workload type, performance objectives, priorities, and infrastructure or policy constraints. Koi, Tandemn's planning algorithm, determines a cluster-wide deployment plan; the Tandemn Manager executes plans through existing infrastructure and returns observed performance to improve future plans. Tandemn runs in the operator's environment and works alongside existing inference engines and orchestration systems. Use the current Koi research paper for technical methods and measured results. The original Koi whitepaper is historical and is listed separately below. Claims vary by benchmark and baseline; consult the paper's evaluation tables for the exact comparison. ## Company and product - [Homepage](https://www.tandemn.com/): Company overview and value proposition. - [Product overview](https://www.tandemn.com/product/overview): How Tandemn combines workload requirements with available GPU resources. - [Planning algorithm](https://www.tandemn.com/product/algorithm): Koi's causal models, objectives, and cluster-wide planning. - [Deployment](https://www.tandemn.com/product/deploy): Self-hosted operation and integration with existing infrastructure. - [Open source](https://www.tandemn.com/product/open-source): Open-source components and enterprise offering. - [About Tandemn](https://www.tandemn.com/about): Team and company background. - [Contact](https://www.tandemn.com/contact): Request a demo or ask about deployment. ## Research and writing - [Current Koi research paper](https://www.tandemn.com/blog/koi): Technical explanation, evaluation setup, benchmark tables, and limitations. - [Research and writing index](https://www.tandemn.com/blog): All published articles and case studies. - [AWS networking for fast inference](https://www.tandemn.com/blog/aws-fast-networking-is-deep-tribal-knowledge): GPU networking details that affect inference performance. - [Six LLM cost surprises](https://www.tandemn.com/blog/gangmuk-ep-3): How workload shape, GPU choice, parallelism, and topology affect serving cost. ## Optional - [Original Koi whitepaper, V1](https://www.tandemn.com/blog/koi-whitepaper-v1): Historical precursor to the current research paper; links to the original PDF. - [Deployment documentation](https://docs.tandemn.com/introduction): Technical documentation. - [Tandemn on GitHub](https://github.com/Tandemn-Labs): Open-source projects. - [XML sitemap](https://www.tandemn.com/sitemap.xml): Canonical, indexable pages for search crawlers.