Launch your career with Hamilton Barnes’ Graduate Hub
Take me there

Site Reliability Engineer - Hosting

1738900
  • Circa $250,000 base salary
  • San Francisco, California, United States
  • Permanent
  • 250000
  • Artificial Intelligence
  • AI Software


Keen to join a company that champions growth and development?

Join a rapidly scaling AI cloud infrastructure provider building next-generation GPU platforms for large-scale AI training, experimentation, and inference, significantly expanding operations across the United States alongside continued international growth.

The organization is inviting applications from experienced professionals for the role of Senior / Staff Site Reliability Engineer to scale and stabilize large-scale HPC and cloud environments powering GPU-intensive workloads. The ideal candidate will work alongside platform, ML, and infrastructure teams to architect reliable systems, automate operations, and improve observability across distributed compute clusters.

Ready to make a move? Get in touch and apply today!


Responsibilities:

  • Ensure the reliability, scalability, and performance of HPC and cloud infrastructure environments
  • Design, build, and maintain automation, observability, and monitoring frameworks for GPU compute clusters
  • Collaborate with ML, data, and platform engineering teams to deliver highly available infrastructure systems
  • Improve CI/CD pipelines, deployment workflows, and operational tooling
  • Contribute to infrastructure architecture discussions and long-term platform strategy
  • Diagnose performance bottlenecks across distributed systems and HPC workloads
  • Support and optimize Slurm-based GPU cluster environments
  • Participate in an on-call rotation supporting mission-critical infrastructure operations


Skills/Must Have:

  • Deep experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or related fields
  • Strong experience supporting HPC or large-scale distributed compute environments
  • Deep Linux expertise (Ubuntu/Debian preferred)
  • Strong scripting and automation skills using Python, Go, or Bash
  • Hands-on experience with public cloud platforms or modern GPU cloud providers
  • Strong understanding of networking fundamentals (DNS, TCP/IP, routing, performance optimization)
  • Experience with Infrastructure-as-Code tooling such as Terraform and Ansible
  • Proven experience operating Slurm-based GPU/HPC clusters
  • Ability to troubleshoot distributed systems and optimize workload scheduling/performance


Benefits:

  • Stock options
  • Remote working option and allowance 


Salary:

  • Circa $250,000 base salary
Ben Davies Director Global AI Infrastructure

Apply for this role