Your cart 0 items
Your cart is empty.
Searching the site...
No matches. Try another word.
Servers & compute

AI training cluster

A tightly coupled multi-GPU cluster you own and run in your own hall. Dense GPU nodes joined by an NVLink-class internal mesh and a 400 Gb/s RDMA fabric, fed by all-flash storage and one scheduler, cooled direct-to-chip so a rack can pull past 60 kW without a wind tunnel.

The problem

Waiting

A training run is only as fast as its slowest link, and on borrowed infrastructure that link is almost never the GPU. Rented instances land on whatever rack the provider has free, so eight machines that should share one switch end up three hops apart and an all-reduce that should take milliseconds stalls on the network. Costs meter by the hour whether the GPUs are saturated or blocked on a straggler, and the moment you pause to inspect a checkpoint the meter keeps running. Data has to be copied out of your own building and back, which the dataset owners and compliance rarely love. And once a run scales past a few nodes, utilisation quietly collapses because nothing was placed to keep the GPUs fed. A cluster you own, wired as one machine and scheduled as one machine, fixes the part the hourly price hides: keeping every GPU busy.

Design targets
400 Gb/s RDMA per GPU, rail to the spine
> 90% Scaling efficiency to 64 GPUs
< 45 C Coolant to the cold plate
1.06 Design PUE at the rack
The system

One machine

Compute, fabric and storage, each sized so nothing downstream of the GPU becomes the bottleneck. Open a layer for what it does and how far it goes.

Eight per node
Imagery in production

GPU nodes

Dense, all-to-all inside the box.

8 GPUs, NVLink-class mesh HBM3 memory, PCIe Gen5 host 700 W class per accelerator, liquid cooled
Eight per node
Imagery in production

GPU nodes

Dense, all-to-all inside the box.

8 GPUs, NVLink-class mesh HBM3 memory, PCIe Gen5 host 700 W class per accelerator, liquid cooled

Each node carries eight accelerators wired to each other over an NVLink-class switch mesh, so every GPU reaches every other GPU inside the chassis at full bandwidth without touching the network. Host side is dual server CPUs on PCIe Gen5 with GPUDirect, so data moves from NVMe and the fabric straight into GPU memory without a bounce through system RAM. This is the unit you scale. One node trains a mid-size model on its own; the fabric is what turns thirty-two of them into one.

Rail-optimised
Imagery in production

The fabric

Where scaling is won or lost.

400 Gb/s RDMA per GPU Non-blocking rail-optimised fat-tree GPUDirect RDMA, NCCL tuned
Rail-optimised
Imagery in production

The fabric

Where scaling is won or lost.

400 Gb/s RDMA per GPU Non-blocking rail-optimised fat-tree GPUDirect RDMA, NCCL tuned

Between nodes, every GPU gets its own 400 Gb/s RDMA rail into a non-blocking fat-tree, so a gradient all-reduce across the whole cluster runs GPU to GPU with no CPU in the path. Rail-optimised means GPU 0 on every node shares a leaf, GPU 1 shares the next, and so on, which keeps collective traffic off the spine and holds bandwidth flat as you add nodes. This is the layer rented infrastructure gets wrong. Placed right, scaling efficiency stays above ninety percent to sixty-four GPUs instead of falling off at eight.

Flash + scheduler
Imagery in production

Storage and scheduler

Keeping the GPUs fed and shared.

All-NVMe parallel filesystem Hundreds of GB/s aggregate read Slurm-class scheduler, gang + fair-share
Flash + scheduler
Imagery in production

Storage and scheduler

Keeping the GPUs fed and shared.

All-NVMe parallel filesystem Hundreds of GB/s aggregate read Slurm-class scheduler, gang + fair-share

Datasets and checkpoints live on an all-flash parallel filesystem that delivers hundreds of gigabytes per second aggregate, enough that data loading never starves a full cluster of GPUs mid-epoch. A Slurm-class scheduler gang-schedules a job so all its GPUs start together, enforces fair-share between teams, and checkpoints to fast storage so a long run survives a node fault and resumes rather than restarts. One cluster, many teams, no one team able to strand the GPUs.

Solution optimized products

The build

Every line item selected and integrated so the cluster behaves as one machine, not a shelf of servers.

Imagery in production
Compute

GPU training node

8 GPU, liquid

Eight accelerators on an NVLink-class mesh, dual Gen5 host CPUs, HBM3 memory, direct-to-chip cold plates. The unit you add to grow the cluster.

Imagery in production
Fabric

400 Gb/s RDMA switch

Non-blocking

Rail-optimised leaf and spine giving every GPU its own 400 Gb/s RDMA path, GPUDirect and NCCL tuned at commissioning.

Imagery in production
Storage

All-flash parallel store

Hundreds of GB/s

NVMe parallel filesystem for datasets and checkpoints, sized so loading never starves the GPUs during an epoch.

Imagery in production
Facility

Coolant distribution unit

Per rack, N+1

Liquid-to-liquid CDU feeding the cold plates at under 45 C, with redundant pumps and leak detection to the rack controller.

Imagery in production
Software

CalyOS Server + scheduler

On cluster

Provisioning, the Slurm-class scheduler, fair-share and gang scheduling, telemetry and checkpoint management, all on hardware you own.

Why own it

Yours

On-prem

Owned

The cluster sits in your hall, on your network, under your scheduler. Datasets never leave the building, the GPUs are busy for you and not metered by the hour, and when a run pauses so does the only cost that matters. You buy the machine once and run it flat out.

Imagery in production

GPUs busy, not metered

You size the cluster to your workload and run it around the clock. There is no hourly meter to race and no idle-time bill.

Data stays in the building

Training data and checkpoints live on your storage inside your network, which keeps dataset owners and compliance on side.

Placed to scale

Nodes, rails and storage are wired as one machine, so scaling efficiency holds instead of collapsing past a few nodes.

One capital cost

You buy the configured cluster once. There is no per-GPU-hour licence and no forced subscription behind it.

“We were burning a fortnight of cloud budget on runs that sat at forty percent GPU utilisation. On our own cluster the same model trains at over ninety, and the bill stopped moving when we stopped training.”
Head of ML infrastructure, a genomics company
Under the hood

Three layers

Imagery in production
Compute

Eight GPUs that behave as one

Inside a node, eight accelerators share an NVLink-class switch mesh, so any GPU reaches any other at full bandwidth without a network hop. HBM3 keeps model shards and activations close to the cores, and a dual Gen5 host with GPUDirect lets data land in GPU memory straight from storage and the fabric. A single node is enough to train a mid-size model; the point of the mesh is that model-parallel shards inside the box pay almost no communication tax.

Imagery in production
Fabric

A rail for every GPU

Between nodes, each GPU owns a 400 Gb/s RDMA rail into a non-blocking rail-optimised fat-tree. Collective operations like all-reduce run GPU to GPU over GPUDirect RDMA with no CPU in the path, and rail optimisation keeps that traffic off the spine so bandwidth stays flat as the cluster grows. NCCL is tuned to the topology at commissioning, which is the difference between ninety percent scaling and thirty.

Imagery in production
Storage and schedule

Fed, shared and restartable

An all-flash parallel filesystem serves datasets and checkpoints at hundreds of gigabytes per second so data loading never starves the GPUs. A Slurm-class scheduler gang-schedules each job, enforces fair-share between teams, and writes periodic checkpoints to fast storage, so a fault costs minutes of rollback rather than a whole run.

Sized to fit

Scale

The same building block from a single rack to a hall.

Starter

One rack, up to 32 GPUs

Four training nodes, one leaf switch, a flash store and a CDU. Trains real models today and seeds the fabric to grow.

Pod

A row, 128 to 256 GPUs

Leaf and spine across several racks, one parallel filesystem, one scheduler domain. The point at which rail optimisation earns its keep.

Hall

Multiple pods, one schedule

Pods federated under one fair-share scheduler and one namespace, so teams share capacity without stranding GPUs.

The topology at a glance

8 GPUs per node
400 Gb/s RDMA rail per GPU
Non-blocking Rail-optimised fat-tree
1 Scheduler namespace
Engineering questions

Fabric

InfiniBand or Ethernet?

Either. The default is a 400 Gb/s InfiniBand-class rail-optimised fabric; where you standardise on Ethernet we build the same non-blocking topology on 400 GbE with RoCEv2 and lossless configuration. Both give GPUDirect RDMA and are tuned to NCCL at commissioning.

Can we mix in nodes we already own?

Usually yes. If your existing GPU nodes expose the right RDMA NICs and PCIe generation we can bring them into the fabric and the scheduler as a separate partition, though a scaling number is only guaranteed on the nodes we integrate.

How does model-parallel traffic stay off the network?

Tensor and pipeline parallel groups are placed inside a node so their heavy traffic rides the NVLink-class mesh, while data-parallel all-reduce crosses the RDMA fabric. The scheduler and NCCL topology file enforce that placement rather than leaving it to chance.

Performance

Utilisation

The number that decides the bill is not peak FLOPS on a spec sheet, it is how much of it you actually use during a run. Everything here is built to hold model FLOPs utilisation high as you scale.

Measured on the floor
> 90% Scaling efficiency, 8 to 64 GPUs
40 to 60% Model FLOPs utilisation, typical
< 5% Time lost to comms at 32 GPUs
Minutes Rollback after a node fault
The difference

Coupled

Owned, coupled clusterRented instances
GPU placement Nodes on one fabric, rails planned Wherever the provider has capacity
All-reduce path GPU to GPU over RDMA, no CPU Through host and shared network
Scaling efficiency Over 90% to 64 GPUs Falls off past a few nodes
Cost while paused Nothing, you own it Metered by the hour, idle or not
Where the data lives Your storage, your network Copied out and back
“The all-reduce time barely moves between eight GPUs and sixty-four. That is the whole game for us, and it is why a run that took nine days now takes under four.”
Principal research engineer, an autonomous systems lab
In the hall

Thermals

Imagery in production
Cooling

Heat leaves at the chip

Each accelerator sits under a direct-to-chip cold plate fed with coolant at under 45 C from a per-rack coolant distribution unit. Taking heat away as liquid at the source is why a rack can dissipate past 60 kW without the deafening airflow an air-cooled hall needs, and why the design PUE sits near 1.06 rather than 1.5. Leak detection reports to the rack controller and the CDU carries N+1 pumps.

Imagery in production
Power

Sized for a full rack

A training rack draws tens of kilowatts continuously, so power is delivered over busbar with redundant feeds and per-node metering. We specify the PDU, breaker and feed for your actual node count and give facilities the real steady-state and peak draw, not a nameplate figure, so nothing trips when the whole cluster steps into a run at once.

Imagery in production
Telemetry

Every watt and degree logged

GPU temperature, coolant flow, power draw and fabric health stream into one dashboard, with alerts before a threshold rather than after. The same telemetry drives scheduler decisions, so a node running hot or a failing link is drained of jobs automatically instead of taking a run down with it.

What we run

Kept

Managed

Uptime

A cluster is only useful when it is up and the GPUs are busy. We commission it to a measured scaling number, monitor thermals and fabric health, and hold spares so a failed node is drained and swapped without standing the whole run down. You get the utilisation, not the pager.

Imagery in production

Cooled at the source

Direct-to-chip liquid at under 45 C, per-rack CDU with redundant pumps and leak detection.

Power specified for real load

Busbar delivery, redundant feeds and the true steady-state and peak draw handed to facilities.

Spares on the shelf

A failed node is drained by the scheduler and swapped from local spares, its jobs resumed from the last checkpoint.

Commissioned to a number

We hand over a measured scaling-efficiency figure on your models, not a spec-sheet claim.

Commissioning

Standing it up

  1. Site and facility survey

    Floor loading, power feed, coolant supply and return, and network hand-off checked against the build before anything ships.

  2. Install and plumb

    Racks landed, nodes seated, cold plates and CDU plumbed, busbar and fabric cabled to the rail plan and leak-tested.

  3. Tune and prove

    NCCL tuned to the topology, then a scaling test on your own model to a measured efficiency figure, not a synthetic benchmark.

  4. Handover

    Scheduler, fair-share policy and telemetry configured to your teams, with a run-book and the commissioning results for your file.

What the hall needs to provide

40 to 120 kW Per training rack
< 45 C Coolant supply to plates
1.06 Design PUE at rack
N+1 Pumps and power feeds
Specify it with us

Sizing

Send us the models you train and the hall you have. We will size the nodes, fabric and cooling to a scaling number and a power budget you can take to facilities.

Resources

Papers

The briefs, the sizing method and worked deployments. Enough to scope the cluster before a single node ships. Open a card to read what is inside.

The library

Reading

Briefs for the people who sign, white papers for the people who integrate, and worked cases for the people who run it.

Brief
Imagery in production

Solution brief

The one-page case for owning it.

2 pages For decision makers
Brief
Imagery in production

Solution brief

The one-page case for owning it.

2 pages For decision makers

A two-page summary: the waiting problem on rented infrastructure, the design targets, the build from a single rack to a hall, and the ownership case against hourly GPU rental. Written for the person who signs the capital request, not the person who racks it. Ends with a headline specification and a total-cost comparison over three years.

White paper
Imagery in production

Cluster sizing white paper

How we size nodes, fabric and storage.

18 pages For technical evaluators
White paper
Imagery in production

Cluster sizing white paper

How we size nodes, fabric and storage.

18 pages For technical evaluators

The method we use to size a cluster: reading your model and batch size to a per-GPU memory and communication profile, then choosing node count, rail bandwidth and storage throughput so nothing downstream of the GPU is the bottleneck. Includes the scaling-efficiency model, the storage-throughput rule of thumb per GPU, and worked examples at 32, 128 and 256 GPUs with the expected model FLOPs utilisation.

White paper
Imagery in production

Fabric and NCCL tuning guide

For your own engineers.

22 pages For integrators
White paper
Imagery in production

Fabric and NCCL tuning guide

For your own engineers.

22 pages For integrators

For your engineers: the fabric topology, the rail-optimised placement rule, GPUDirect RDMA setup, and the NCCL topology and tuning file we ship, plus how the scheduler enforces intra-node parallel placement. Enough to reproduce the commissioning tuning and to bring your own nodes into the fabric as a partition.

Use case
Imagery in production

Genomics model training

One rack, 32 GPUs.

ConfigStarter Rundata-parallel
Use case
Imagery in production

Genomics model training

One rack, 32 GPUs.

ConfigStarter Rundata-parallel

A genomics company running variant-calling model training on one Starter rack of 32 GPUs, data-parallel, dataset held on the local flash store so nothing leaves their network. Scaling held above 90 percent across the rack and a nine-day run came in under four, with the bill flat while they paused to inspect checkpoints.

Use case
Imagery in production

Autonomy perception stack

A pod, 256 GPUs.

ConfigPod Runtensor + data parallel
Use case
Imagery in production

Autonomy perception stack

A pod, 256 GPUs.

ConfigPod Runtensor + data parallel

An autonomy lab training a perception stack across a 256-GPU pod, tensor-parallel inside each node over the NVLink-class mesh and data-parallel across the RDMA fabric. Rail optimisation kept all-reduce time nearly flat from 8 to 64 GPUs, and the scheduler resumed the run from checkpoint after a node fault in minutes.

Use case
Imagery in production

Shared research cluster

Many teams, one schedule.

ConfigHall Runfair-share
Use case
Imagery in production

Shared research cluster

Many teams, one schedule.

ConfigHall Runfair-share

A research group sharing one hall between a dozen teams under a fair-share scheduler. Gang scheduling starts each job's GPUs together, fair-share stops any one team monopolising capacity, and checkpointing lets long runs yield to short ones without losing work.

Request the files

Files

Sent as PDF on request, tuned to the models and hall you send us.

Specify it with us

Build

Tell us the workload and we will come back with a configured cluster, a scaling number and a power and cooling budget. You own what we build.