Storage Engineer - Hosting
1723830
Posted: 23/07/2026
- $200,000 base salary
- United States - Remote
- Permanent
- 200000
- Artificial Intelligence
- AI Network
Join a Founders Fund-backed NVIDIA cloud partner building the high-performance infrastructure that powers the world’s most ambitious AI research. In the world of GPUaaS, the bottleneck is rarely the compute; it’s the data.
You will be a Storage Engineer who understands that AI at scale requires more than just capacity; it requires massive throughput, ultra-low latency, and the ability to feed thousands of GPUs without a hiccup. You will architect and build the data layer that supports foundation model training and enterprise-grade production inference.
If you are interested in this exciting opportunity, get in touch and apply today!
Responsibilities:
- Design & Deploy AI Storage: Architect and implement high-performance parallel file systems (Weka, Lustre, or similar) optimised specifically for GPU-heavy workloads and multi-node training.
- Optimise Data Pipelines: Fine-tune storage performance to ensure maximum GPUDirect Storage (GDS) efficiency, minimising latency between the storage fabric and the GPU memory.
- Manage Scale & Reliability: Build and maintain petabyte-scale storage clusters across multiple global data centers, ensuring 99.99% uptime for mission-critical AI research labs.
- Infrastructure Integration: Partner with Network and Data Center engineers to configure high-speed storage networking (InfiniBand/400G Ethernet) and ensure seamless backend connectivity.
- Automate Storage Ops: Develop Terraform providers, Ansible playbooks, or Python scripts to automate the provisioning, monitoring, and scaling of storage resources.
- Troubleshoot Complex I/O: Act as the Tier-3 lead for storage-related performance degradation, identifying root causes in the filesystem, network, or Linux kernel.
Skills/Must have:
- Specialised Storage Expertise: 5+ years of experience with high-performance storage solutions (WekaIO, VAST Data, BeeGFS, or DDN) in a Linux-heavy environment.
- AI Infrastructure Knowledge: Deep understanding of how storage interacts with NVIDIA GPU stacks (HGX/DGX) and the specific I/O patterns of ML training (checkpoints, small file reads, etc.).
- Networking Proficiency: Hands-on experience with InfiniBand, RoCEv2, and NVMe-over-Fabrics (NVMe-oF).
- Systems Automation: Strong scripting skills in Python, Go, or Bash, and experience with IaC tools like Terraform or Pulumi.
- Linux Internals: Deep knowledge of the Linux storage stack, including XFS/ZFS, LVM, and kernel tuning for high-throughput networking.
Benefits:
- 10% bonus
- Stock options
Salary:
- $200,000 base salary
Ben Davies
Director Global AI Infrastructure