Launch your career with Hamilton Barnes’ Graduate Hub
Take me there

Network Engineer (AI Infrastructure) - Hosting

1738922
  • Circa $225,000 base salary
  • San Francisco, California, United States
  • Permanent
  • 200000
  • Artificial Intelligence
  • AI Network


Ready to take the next step in your career?

Join a rapidly scaling AI cloud infrastructure provider building next-generation GPU platforms for large-scale AI training, experimentation, and inference, significantly expanding operations across the United States alongside continued international growth.

An exciting opportunity has arisen for a Senior Network Engineer to design, deploy, and operate ultra-low-latency, high-throughput network fabrics supporting large GPU clusters. The ideal candidate will work with technologies such as NVIDIA Spectrum and Cumulus Linux, helping build scalable 400G network environments optimized for AI and HPC traffic patterns.

Ready to make a move? Get in touch and apply today!


Responsibilities:

  • Design, deploy, and operate large-scale 400G network fabrics supporting AI and HPC workloads
  • Build and maintain high-performance data center networking environments using NVIDIA Spectrum switches and Cumulus Linux
  • Optimize network performance, latency, throughput, and resilience across GPU clusters
  • Troubleshoot complex Layer 2/Layer 3 networking issues in distributed HPC environments
  • Automate network provisioning, configuration management, and operational workflows
  • Collaborate with infrastructure, platform, and ML engineering teams to ensure efficient GPU cluster communication
  • Support network observability, telemetry, and capacity planning initiatives
  • Contribute to network architecture strategy and scalability planning for future deployments


Skills/Must Have:

  • Deep experience in network engineering within data center, HPC, cloud, or large-scale infrastructure environments
  • Strong hands-on experience with NVIDIA Spectrum switching platforms and Cumulus Linux
  • Experience designing and operating 100G/400G Ethernet fabrics
  • Deep understanding of modern data center networking protocols (BGP, EVPN, VXLAN, MLAG, ECMP)
  • Strong Linux systems knowledge and network automation skills
  • Experience with automation tools and scripting (Python, Ansible, Bash, Terraform preferred)
  • Familiarity with AI/HPC traffic patterns and GPU cluster networking requirements
  • Strong troubleshooting and performance optimization capabilities in high-throughput environments


Benefits:

  • Stock options 
  • Remote working options and allowance 


Salary:

  • Circa $225,000 base salary 
Ben Davies Director Global AI Infrastructure

Apply for this role