CASE STUDIES

PROOF INPRODUCTION.

Real infrastructure. Real workloads. Real outcomes. These are the systems we've built and the problems we've solved.

CASE-01 / AI INFRASTRUCTURE

Series B AI Research Lab

512-GPU Training Cluster for Large Language Model Development

Challenge

A well-funded AI research lab needed to scale from a 64-GPU pilot to a 512-GPU production training cluster in under 90 days — without disrupting ongoing research runs.

Solution

We designed a 4-rack, 512x H100 SXM5 cluster with 400G InfiniBand HDR fabric, NVMe-oF shared storage at 40TB/s aggregate bandwidth, and a custom job scheduler integration. Delivered on day 87.

Key Specs

512x NVIDIA H100 SXM5 GPUs
400G InfiniBand HDR fabric
40TB/s aggregate NVMe-oF storage
87-day delivery timeline

Outcome

3.2x throughput improvement over prior infrastructure. First production training run completed on day 91.

CASE-02 / CLOUD INFRASTRUCTURE

Cloud Infrastructure Provider

Custom Inference Accelerator Card for Hyperscale Deployment

Challenge

A cloud provider needed a custom inference accelerator that could be deployed at hyperscale — with specific power envelope, PCIe form factor, and latency targets that no off-the-shelf card could meet.

Solution

Full-stack engagement: FPGA-based accelerator design, custom firmware, PCIe driver development, and integration with the provider's existing orchestration layer. 18-month program from spec to production.

Key Specs

Custom FPGA-based inference accelerator
75W TDP, PCIe Gen5 x16 form factor
Custom firmware and Linux kernel driver
18-month spec-to-production timeline

Outcome

40% lower inference latency vs. GPU baseline at 60% of the power cost. Deployed across 3 data center regions.

CASE-03 / ML PLATFORM

Enterprise ML Platform Company

HPC Storage Architecture for Petabyte-Scale Training Data

Challenge

An ML platform company was bottlenecked on data ingestion — their training jobs were GPU-idle 35% of the time waiting on storage I/O. They needed a storage architecture that could keep up with 256 GPUs.

Solution

Redesigned storage fabric using NVMe-oF over RoCEv2, deployed a Lustre parallel filesystem with custom tuning for ML workload access patterns, and implemented a tiered caching layer for hot dataset reuse.

Key Specs

NVMe-oF over RoCEv2 fabric
Lustre parallel filesystem
Tiered caching for hot dataset reuse
256-GPU storage saturation target

Outcome

GPU idle time from storage I/O dropped from 35% to under 4%. Training throughput increased 1.9x with no additional GPU spend.

Working on something similar?

We've seen most of these problems before. Let's talk about yours.

Scope a Project
BlackSalt Technology Group

HPC & GPU engineering for AI/ML labs and cloud infrastructure at scale.

© 2026 BlackSalt Technology Group LLC. All rights reserved.

HPC · GPU · FIRMWARE · INFRASTRUCTURE