Skip to content

Get more from the GPUs you already have.

Tandemn is an open-source, algorithmic planner for self-hosted AI inference that unlocks more usable capacity from your existing GPU fleet. It places and configures workloads across heterogeneous hardware, clouds, and regions so you can run more jobs and serve more demand without expanding your infrastructure.

Improve All Your Key Metrics - Simultaneously

GPU optimization usually forces tradeoffs between latency, throughput, utilization, SLOs, and cost. Tandemn continuously adapts GPU selection, parallelism strategies, replica counts, and workload placement to actual demand. Across heterogeneous fleet simulations, it admitted more online and batch jobs while improving SLO-honoring goodput, TTFT attainment, and cost per token.

+30%more jobs admitted on average
+45%SLO-honoring goodput
+19%TTFT targets met
−31%cost per token
See Results

One fleet, allocated in real time.

Tandemn treats fragmented, heterogeneous GPUs as one adaptive capacity pool. As inference demand shifts, it reallocates workloads so more jobs fit on the fleet you already run.

ChatbotEmbeddingsRerankingVision
Four workloads sharing GPU capacityCompare GPU demand with Tandemn against the same fleet without Tandemn. The dashed line marks total cluster capacity; red shading shows demand above that limit. Hover to inspect a point, or use the arrow keys when focused.
Shared capacity map99/128 GPU in use

Proven by Research. Validated on Real Hardware.

Tandemn does more than improve GPU utilization.

We've tested it across diverse workloads, GPU configurations, time horizons, and performance objectives, comparing Tandemn with alternative scheduling approaches.

The result: more usable GPU capacity while simultaneously improving latency, goodput, SLO attainment, and cost efficiency.

Up to 80%More capacity
Up to 45%Lower cost

Research & Writing

Ideas, research, and notes from the Tandemn team.