Check out our 2026 USA Salary Survey
Take me there

HPC Solution Architect - AI Infrastructure

1728697
  • $300,000 base salary
  • San Francisco, California, United States
  • Permanent
  • 300000
  • Artificial Intelligence
  • AI Network
  • AI Software


Keen to join a company that champions growth and development?

Join a fast-scaling AI compute platform operating at the forefront of global AI infrastructure. The organisation is building the technology that connects AI training and inference workloads with a global network of compute providers, supporting some of the world's leading AI labs and cloud operators.

This company is seeking the first engineer in this function to build and define the role from the ground up. This greenfield opportunity is ideal for someone with deep HPC experience who wants to shape technical standards, establish best practices, and create the foundations of a rapidly growing AI infrastructure platform.

If you would like to learn more about this opportunity, feel free to reach out and apply today!


Responsibilities:

  • You will be vetting prospective compute providers across GPU hardware, network fabric, storage, and orchestration — and making the call on whether a cluster qualifies
  • You will be building the acceptance test suite, benchmark methodology, and quality thresholds from scratch, replacing tribal knowledge with rigorous written standards
  • You will be executing hands-on validation where it counts: burn-in testing, fabric validation (InfiniBand/RoCE), NCCL benchmarks, and storage performance testing
  • You will be guiding providers through technical onboarding, working directly with their engineers to close gaps on a predictable timeline
  • You will be partnering with compute procurement on technical due diligence - giving a clear read on remediation cost before contracts are signed
  • You will be the standing technical relationship with existing providers, catching architectural drift before it becomes a customer issue
  • You will be feeding learnings back into provider-facing standards so each qualification cycle is faster than the last


Skills / Must Have:

  • Deep HPC experience, you will have designed, built, or operated GPU clusters at meaningful scale
  • Strong network fabric knowledge across both InfiniBand and RoCE; you will be able to evaluate topology and diagnose underperformance
  • Hands-on with distributed orchestration and the HPC software stack: Slurm, Kubernetes, OpenMPI or equivalent
  • Data-centre literacy, power, cooling, cabling, physical-layer realities; you will know what questions to ask when you walk a facility
  • Benchmarking judgment  - you will know which numbers matter for large-scale training and how to design tests that can't be gamed
  • You will be writing technical standards that providers can build against without you in the room
  • You will be comfortable delivering a failing grade to a provider who wants your business  - and keeping the relationship intact


Benefits:

  • Large equity
  • Comprehensive healthcare, dental, and vision (you and dependents)
  • 401(k)
  • Unlimited PTO


Salary:

  • $300,000 base salary 
Ben Davies Director Global AI Infrastructure

Apply for this role