You focus on the models.

We handle the infrastructure.

Train foundation models without thinking about infrastructure.

A CLI for GPUs

CLI COMMANDS

tp cluster create 8xB200 --num-nodes 4

Multi-node GPU clusters on demand. Always 3.2Tbs Infiniband.

tp storage create 100000

Shared storage volumes for clusters. 10GB/s reads & 5GB/s writes. Faster than local NVMe.

tp job push train.toml c-abcde12345

Kick off training jobs with a Git-style interface.

Why TensorPool?

FEATURES

High Performance Storage

Storage infrastructure engineered for distributed training. Optimized for high throughput & low latency workloads with sustained aggregate performance: up to 300 GB/s read, 150 GB/s writes, 1.5M read IOPS, 750k write IOPS.

Git style CLI

Training with no orchestration burden. Manage experiments with a familiar push/pull paradigm that eliminates idle GPU time. Normal SSH access also offered.

Capacity

TensorPool provides on-demand multi-node clusters. We do this by scheduling & binpacking users onto TensorPool operated clusters in partnership with compute providers (Nebius, Lightning AI, Verda, Lambda Labs, Google Cloud, Microsoft Azure, and others).

Frequently Asked Questions

What instance types are available?

How can I use TensorPool at my company?

When should I use TensorPool?