HomeBlogAboutContact

Research Blog

Sep 14, 2026

Introducing Koi: Get More from the GPUs You Already Have

Latest

Tandemn’s planning algorithm helps teams get more from their existing GPUs. In a three-hour heterogeneous-cluster simulation, it delivered 45% more throughput at 31% lower cost per token, while improving job admission and SLO attainment across evaluations.

Jun 5, 2026

Six LLM Serving Cost Surprises

Cost-optimal LLM serving is a joint optimization across six dimensions — model × workload shape × GPU type × parallelism (TP/PP) × engine settings × cloud topology.

May 5, 2026

AWS Fast Networking Is Deep Tribal Knowledge

A practical walk-through of the hidden AWS networking stack behind fast multinode inference on A100 and L40S clusters.

Apr 29, 2026

Complexity of LLM Inference Optimization

A benchmark sweep showing why cost-optimal LLM inference is a live optimization problem across TP, PP, GPU count, SLO, and workload shape.

Apr 20, 2026

Your LLM Inference Cluster Is Probably Operating Suboptimally

A deep dive into what actually drives cost and throughput in self-hosted LLM inference — and why most clusters silently overpay 2–5×.

Apr 6, 2026

How Retailers Run Batch Inference on Product Catalogs

A case study in catalog enrichment, attribute extraction, and the infrastructure problem underneath.

May 15, 2025

Tandemn Raises $1.7M to Build the Missing Layer for AI Infrastructure

Tandemn has raised $1.7 million to build open infrastructure for running AI workloads across heterogeneous compute.

Cluster-level inference orchestration for teams running AI workloads across heterogeneous infrastructure.

Explore

HomeBlogContact

Resources

DocsGitHubLinkedInYouTube