Check out our 2026 USA Salary Survey
Take me there

Storage Engineer - Hosting

1723830
  • $200,000 base salary
  • United States - Remote
  • Permanent
  • 200000
  • Artificial Intelligence
  • AI Network


Join a Founders Fund-backed NVIDIA cloud partner building the high-performance infrastructure that powers the world’s most ambitious AI research. In the world of GPUaaS, the bottleneck is rarely the compute; it’s the data. 

You will be a Storage Engineer who understands that AI at scale requires more than just capacity; it requires massive throughput, ultra-low latency, and the ability to feed thousands of GPUs without a hiccup. You will architect and build the data layer that supports foundation model training and enterprise-grade production inference.

If you are interested in this exciting opportunity, get in touch and apply today! 

Responsibilities:

  • Design & Deploy AI Storage: Architect and implement high-performance parallel file systems (Weka, Lustre, or similar) optimised specifically for GPU-heavy workloads and multi-node training.

  • Optimise Data Pipelines: Fine-tune storage performance to ensure maximum GPUDirect Storage (GDS) efficiency, minimising latency between the storage fabric and the GPU memory.
  • Manage Scale & Reliability: Build and maintain petabyte-scale storage clusters across multiple global data centers, ensuring 99.99% uptime for mission-critical AI research labs.
  • Infrastructure Integration: Partner with Network and Data Center engineers to configure high-speed storage networking (InfiniBand/400G Ethernet) and ensure seamless backend connectivity.
  • Automate Storage Ops: Develop Terraform providers, Ansible playbooks, or Python scripts to automate the provisioning, monitoring, and scaling of storage resources.
  • Troubleshoot Complex I/O: Act as the Tier-3 lead for storage-related performance degradation, identifying root causes in the filesystem, network, or Linux kernel.

Skills/Must have:

  • Specialised Storage Expertise: 5+ years of experience with high-performance storage solutions (WekaIO, VAST Data, BeeGFS, or DDN) in a Linux-heavy environment.

  • AI Infrastructure Knowledge: Deep understanding of how storage interacts with NVIDIA GPU stacks (HGX/DGX) and the specific I/O patterns of ML training (checkpoints, small file reads, etc.).

  • Networking Proficiency: Hands-on experience with InfiniBand, RoCEv2, and NVMe-over-Fabrics (NVMe-oF).

  • Systems Automation: Strong scripting skills in Python, Go, or Bash, and experience with IaC tools like Terraform or Pulumi.

  • Linux Internals: Deep knowledge of the Linux storage stack, including XFS/ZFS, LVM, and kernel tuning for high-throughput networking.

Benefits:

  • 10% bonus
  • Stock options 

Salary:

  • $200,000 base salary


Ben Davies Director Global AI Infrastructure

Apply for this role