Check out our 2026 USA Salary Survey
Take me there

HPC Engineer - AI Infrastructure

1727256
  • $275,000 Base
  • San Francisco, California, United States
  • Permanent
  • 250000
  • Artificial Intelligence


Ready to take the next step in your career?

Join a VC-backed GPU cloud company at the founding stage, building the next generation of managed GPU compute and inference infrastructure. Backed by leading venture investors, the organization delivers fully managed AI infrastructure and offers the opportunity to help shape the technical foundation of a rapidly growing GPU cloud platform.

This company is seeking a Founding Engineer to take ownership of core GPU cloud infrastructure across Kubernetes, Slurm, GPU orchestration, distributed storage, high-bandwidth networking, and inference platforms. This hands-on role offers founding-level ownership, significant influence over platform architecture, and the opportunity to work closely with the founders to build and scale production AI infrastructure from the ground up.

Don’t miss out on this exciting opportunity and apply today!


Responsibilities:

  • Design, build, and operate large-scale Kubernetes and Slurm based GPU clusters in production
  • Build and manage GPU orchestration layers on top of core scheduling infrastructure
  • Architect and operate distributed object storage, NVMe storage clusters, and storage systems for large-scale AI workloads
  • Design and manage high-bandwidth networking supporting distributed training and inference at scale
  • Design, deploy, and optimise a scalable inference stack for serving large AI models with low-latency, high-throughput GPU acceleration
  • Build telemetry, observability, and automated remediation across the GPU fleet
  • Implement power-aware infrastructure mechanisms including GPU power capping, power-aware scheduling, and telemetry loops
  • Own operational health, reliability, and performance of the platform end to end
  • Work directly with founders on architecture, roadmap, and technical strategy
  • Help define engineering culture, standards, and hiring as one of the first technical team members


Skills/Must Have:

  • 3 to 6 years of hands-on experience operating large-scale AI or HPC infrastructure in production
  • Proven experience operating large-scale Kubernetes and Slurm clusters
  • Experience building and managing GPU orchestration layers on top of core schedulers
  • Strong understanding of distributed object storage, NVMe storage clusters, and high-bandwidth networking
  • Deep knowledge of storage architectures for large-scale AI infrastructure
  • Comfort operating with founding-level ownership across the full infrastructure stack
  • Based in or willing to relocate to San Francisco


Benefits:

  • Founding engineer equity
  • Full benefits package


Salary:

  • $275,000 Base
Sam Hammersley Senior AI Infrastructure Consultant

Apply for this role