Senior AI Network Engineer - AI Infrastructure
- $220,000 – $350,000 Base Salary
- United States
- Permanent
- 300000
- Artificial Intelligence
- AI Network
Looking for a role with plenty of growth opportunities?
Join one of North America's fastest-growing AI infrastructure providers, delivering large-scale GPU cloud platforms for AI training, fine-tuning, and inference workloads. Backed by significant investment and advanced infrastructure expertise, the organization builds high-performance AI environments powered by NVIDIA GPUs, high-speed networking, storage, and Kubernetes.
This company is seeking a Senior AI Network Engineer to design, deploy, and optimize networking infrastructure for large-scale GPU clusters. The role offers the opportunity to solve complex HPC and AI networking challenges while supporting low-latency, high-bandwidth platforms built for next-generation AI workloads.
Ready to make a move? Get in touch and apply today!
Responsibilities:
- Design, deploy, and operate high-performance AI networking infrastructure supporting large-scale GPU clusters.
- Build and optimise low-latency Ethernet and InfiniBand fabrics for distributed AI training and inference workloads.
- Configure and support NVIDIA Spectrum switches, Cumulus Linux, and modern spine-leaf network architectures.
- Optimise network performance for RDMA, RoCE, GPUDirect RDMA, NCCL, and large-scale distributed training.
- Collaborate with platform, storage, Kubernetes, and infrastructure teams to maximise cluster performance and reliability.
- Develop network automation using Infrastructure-as-Code, CI/CD, and configuration management tools.
- Implement monitoring, telemetry, and observability across AI networking environments.
- Troubleshoot complex networking, hardware, and distributed systems issues across production GPU infrastructure.
- Support capacity planning, network scaling, and future infrastructure expansion.
Skills / Must Have:
- 5+ years of experience in Network Engineering supporting large-scale data centre, cloud, HPC, or AI infrastructure environments.
- Strong knowledge of spine-leaf networking and large-scale data centre architectures.
- Hands-on experience with NVIDIA Spectrum switches and Cumulus Linux.
- Deep understanding of Ethernet fabrics supporting AI and HPC workloads.
- Experience with RoCE, RDMA, InfiniBand, BGP, EVPN-VXLAN, MLAG, and modern routing protocols.
- Experience supporting GPU infrastructure and distributed AI training environments.
- Strong Linux systems knowledge and scripting experience using Python, Bash, or Ansible.
- Experience with automation, Infrastructure-as-Code, and network telemetry platforms.
Desirable Skills:
- Experience supporting NVIDIA H100, H200, B200, or Blackwell GPU deployments.
- Knowledge of NCCL, CUDA networking optimisation, GPUDirect RDMA, and distributed AI workloads.
- Experience with Kubernetes networking and cloud-native infrastructure.
- Familiarity with storage networking technologies including VAST, Weka, or BeeGFS.
- Background working within hyperscalers, GPU cloud providers, AI infrastructure companies, or HPC environments.
- Experience deploying multi-thousand GPU clusters.
Benefits:
- Competitive salary with annual bonus and equity opportunities.
- Opportunity to help build one of North America's fastest-growing AI infrastructure platforms.
- Work with cutting-edge NVIDIA GPU technology and hyperscale networking environments.
- High-impact engineering role with significant technical ownership.
- Collaborative engineering culture with minimal bureaucracy.
- Flexible working arrangements and excellent career progression.
Salary:
- $220,000 – $350,000 Base Salary