Skip to main content

IBM AI Appliance

Enterprise AI in one rack, on your own floor

IBM Fusion HCI with the IBM Storage Scale System 6000: GPU compute, high performance storage, Red Hat OpenShift and the AI services that run on top. Built and tested in the IBM factory, delivered and supported by Li9, and serving models the day the rack is powered up.

IBM Storage Fusion HCI rack, the 42U cabinet the IBM AI Appliance is built in
Up to
32
NVIDIA H200 NVL GPUs
4.5 TB of GPU memory in one pool
Up to
3.74PB
usable storage
Performance and capacity tiers, one namespace
Up to
170GB/s
delivered read
13 million operations a second
42U
one rack
Built and tested before it ships

Configurations

One design, three sizes, growing in place

Every level shares the same control and worker nodes, the same switches and the same storage policy. What changes is how many GPUs are in the rack and how much flash sits behind them.

Entry Level

A first production AI platform, with room to grow in place

NVIDIA H200 NVL GPUs
8
GPU nodes
2 nodes, 4 GPUs each
GPU memory
1.1 TB
Performance tier (flash)
147 TB usable
Capacity tier
Optional
Delivered read
approx 86 GB/s
Control and worker nodes
3 and 3

Mid-Level

A third GPU node, more flash, and a capacity tier for the lakehouse

NVIDIA H200 NVL GPUs
12
GPU nodes
3 nodes, 4 GPUs each
GPU memory
1.7 TB
Performance tier (flash)
590 TB usable
Capacity tier
1.18 PB usable
Delivered read
approx 130 GB/s
Control and worker nodes
3 and 3

Maximum

Every slot filled, with both storage tiers at full width

NVIDIA H200 NVL GPUs
32
GPU nodes
4 nodes, 8 GPUs each
GPU memory
4.5 TB
Performance tier (flash)
1.18 PB usable
Capacity tier
2.56 PB usable
Delivered read
approx 170 GB/s
Control and worker nodes
3 and 3
  • Models up to 564 GB on four GPUs bridged by NVLink
  • One Storage Scale namespace, with policy moving warm data to the capacity tier
  • Growth past the maximum is a second rack of the same design

The hardware in the rack

IBM Fusion HCI GPU node
2 to 4

GPU nodes

128 cores (2 x AMD EPYC 9555), 1152 GB memory, up to 8 NVIDIA H200 NVL GPUs

IBM Fusion HCI control and worker node
3 control, 3 worker

Control and worker nodes

64 cores (2 x Intel Xeon Gold 6438), 1 TB memory and NVMe storage, each

200GbE high speed Ethernet switch
2

High speed switches

32 ports of 200GbE each, lossless RoCEv2, paired for redundancy

1GbE management switch
2

Management switches

48 ports of 1GbE each, for server and switch management

What runs on it

The AI services arrive with the platform

Each of these is deployed from the Fusion catalogue onto the same GPUs, the same storage and the same security model, rather than assembled from parts after the hardware lands.

Models as a Service

A governed, shared inference service. Teams get an OpenAI compatible endpoint and a self-service API key rather than a GPU of their own, and the platform keeps track of who is using what.

Developer Hub

One front door for AI, container and virtual machine projects, with the approved models, data services and project templates already in it.

Content-Aware Storage

Your documents made searchable in place. Results carry each user's own file permissions, so retrieval cannot hand someone a passage they could not open themselves.

Data lakehouse

watsonx.data deployed from a service tile, with zero copy access to the data already sitting on the array. The capacity tier is the lakehouse data set.

NVIDIA AI Blueprints

Retrieval, agents, and video search and summarisation, deployed from the catalogue onto the same GPU pool rather than bought and integrated one at a time.

Why the storage decides it

A GPU that waits on disk is a GPU you paid for twice

As a model reads a long conversation it builds up working state. Hold that state on flash and the next question picks up where the last one left off. Throw it away and the GPU rebuilds it from the beginning, every time.

The array is sized so that model loads, document retrieval and that saved state never become the bottleneck, which is also why the appliance clears the NVIDIA Enterprise Reference Architecture storage guideline several times over.

Returning to a long conversation, per session

Rebuilt on the GPUapprox 25 sec
Read back from the applianceapprox 1 sec

About 40 GB of saved state for a long context request, read back at the speed of the array instead of recomputed on the accelerator.

3.4x

the NVIDIA Enterprise Reference Architecture storage guideline, at 32 GPUs

13M

operations a second for retrieval, IBM published for the array

Before it ships

Built, integrated and tested in the IBM factory

The rack is assembled and pre-configured to best practice or to your stated requirements, then tested across five layers and documented for your specific build. The same tests run again on site after installation, so what was proved in the factory is proved again on your floor.

  1. 1

    Network

    What the integrated 200GbE path delivers, in aggregate and per port

  2. 2

    Storage

    Array delivery behind the network, measured against IBM's published figures

  3. 3

    GPU data path

    Storage read straight into GPU memory, with no copy through the host

  4. 4

    AI service

    A timed model load and a timed restore of saved conversation state, per node

  5. 5

    Platform

    The cluster, the operators and the services, as handed over

Working with Li9

What Li9 delivers around the appliance

Li9 architects the appliance with you, works it through IBM, and stays with it once it is on the floor.

01

Written requirements

Models, users, data sources, integrations, security and site constraints, captured and agreed before anything is sized.

02

Validated configuration and architecture

GPUs, storage and nodes sized against those requirements, with the path from pilot to production written down.

03

Factory integration and test

Built and pre-configured in the IBM factory, then tested across five layers and documented for your specific build.

04

Installation and handover

Racked, connected and re-tested on site with the same tests that ran in the factory, handed over with the results.

05

Day 2 support

IBM supports the whole appliance, with one number to call for any part of it. Li9 stays on for the architecture, the upgrades and the next use case.

Questions

What customers ask first

IBM Fusion HCI with NVIDIA H200 NVL GPU nodes and an IBM Storage Scale System 6000, integrated in one 42U rack and running Red Hat OpenShift. It arrives built, configured and tested, and it runs AI workloads beside ordinary containers and virtual machines.

Bring enterprise AI onto your own floor

Tell us the models, the users and the data you want to put to work. Li9 will size the IBM AI Appliance against them and walk you through the architecture first.