Research paperGet more from the GPUs you already have.
Tandemn is an open-source, algorithmic planner for self-hosted AI inference that unlocks more usable capacity from your existing GPU fleet. It places and configures workloads across heterogeneous hardware, clouds, and regions so you can run more jobs and serve more demand without expanding your infrastructure.
Improve All Your Key Metrics - Simultaneously
GPU optimization usually forces tradeoffs between latency, throughput, utilization, SLOs, and cost. Tandemn continuously adapts GPU selection, parallelism strategies, replica counts, and workload placement to actual demand. Across heterogeneous fleet simulations, it admitted more online and batch jobs while improving SLO-honoring goodput, TTFT attainment, and cost per token.
One fleet, allocated in real time.
Tandemn treats fragmented, heterogeneous GPUs as one adaptive capacity pool. As inference demand shifts, it reallocates workloads so more jobs fit on the fleet you already run.
Proven by Research. Validated on Real Hardware.
Tandemn does more than improve GPU utilization.
We've tested it across diverse workloads, GPU configurations, time horizons, and performance objectives, comparing Tandemn with alternative scheduling approaches.
The result: more usable GPU capacity while simultaneously improving latency, goodput, SLO attainment, and cost efficiency.
Research & Writing
Ideas, research, and notes from the Tandemn team.



