Launch your career with Hamilton Barnes’ Graduate Hub
Take me there

Senior SRE - AI Infrastructure

1732890
  • $250,000 gross per year +
  • San Francisco, California, United States
  • Permanent
  • 250000
  • Artificial Intelligence


Keen to join a company that champions growth and development?

Join a seed-stage AI infrastructure company building large-scale training and inference platforms that were previously accessible only to hyperscalers. The business began with a single managed GPU cluster that quickly reached capacity and has since expanded into a global platform spanning infrastructure, networking, and orchestration.

Don’t miss out on this exciting opportunity and apply today!


Responsibilities:

  • Design, deploy, and maintain large-scale GPU clusters (H100/H200/B200) for training and inference workloads.
  • Build automation pipelines for provisioning, scaling, and monitoring compute resources across Slurm and Kubernetes environments.
  • Develop observability, alerting, and auto-healing systems for high-availability GPU workloads.
  • Collaborate with ML, networking, and platform teams to optimise resource scheduling, GPU utilisation, and data flow.
  • Implement infrastructure-as-code, CI/CD pipelines, and reliability standards across thousands of nodes.
  • Diagnose performance bottlenecks and drive continuous improvements in reliability, latency, and throughput.

Skills / Must Have:

  • 7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting large-scale compute environments.
  • Strong hands-on experience with Kubernetes and Slurm for cluster orchestration and workload management.
  • Deep knowledge of Linux systems, networking, and GPU infrastructure (NVIDIA H100/H200/B200 preferred).
  • Proficiency in Python, Go, or Bash for automation, tooling, and performance tuning.
  • Experience with observability stacks (Prometheus, Grafana, Loki) and incident response frameworks.
  • Familiarity with high-performance computing (HPC) or AI/ML training infrastructure at scale.
  • Background in reliability engineering, distributed systems, or hardware acceleration environments is a strong plus.


Benefits:

  • IPO Equity
  • 10% comapny bonus
  • 401K 4% match


Salary:

  • $250,000 gross per year +
Ben Davies Director Global AI Infrastructure

Apply for this role