Introducing Koi: Get More from the GPUs You Already Have
LatestTandemn’s planning algorithm helps teams get more from their existing GPUs. In a three-hour heterogeneous-cluster simulation, it delivered 45% more throughput at 31% lower cost per token, while improving job admission and SLO attainment across evaluations.