Launch your career with Hamilton Barnes’ Graduate Hub
Take me there

Member of Technical Staff (AI Platform) - AI Infrastructure

1729961
  • $250,000 to $300,000 base salary
  • San Francisco, California, United States
  • Permanent
  • 300000
  • Artificial Intelligence


Keen to join a company that champions growth and development?

Join one of the most exciting AI infrastructure companies in the market, building a platform that deploys and operates large-scale GPU clusters for some of the world's leading AI labs. Working with tens of thousands of GPUs across a global network of data centres, the organisation delivers the infrastructure powering next-generation AI training and inference at scale.

This role offers the opportunity to work on complex, production-scale AI infrastructure without the layers of bureaucracy often found in larger organisations. You'll have the freedom to shape how the platform evolves while solving engineering challenges that have a direct impact on customers.

Ready to make a move? Get in touch and apply today!


Responsibilities:

  • You will be building and operating the automated systems that take a cluster from bare machines to customer-ready, end to end
  • You will be managing machine lifecycle between tenants: join, wipe, verify, rejoin, reliably and at scale
  • You will be operating Kubernetes and Postgres across the fleet, keeping critical infrastructure healthy under real production load
  • You will be contributing to custom Kubernetes operators that power the control plane
  • You will be scaling clusters from tens of nodes to thousands, solving the problems that only show up at the edges of scale
  • You will be participating in on-call rotations and responding to production incidents with the depth of knowledge to resolve them quickly


Skills / Must Have:

  • A track record of impressive technical work you can speak to in depth, the years matter less than the quality of the work and its real-world impact
  • 2+ years of on-call experience for critical production services
  • Deep Kubernetes experience, not just consuming it but understanding what is happening underneath
  • Strong Linux fundamentals across kernel, cgroups, containers, networking, and storage
  • Experience operating databases, monitoring, CI/CD, and cloud infrastructure at scale
  • Familiarity with fleet management and capacity planning


Desirable Skills:

  • Experience writing Kubernetes operators
  • Background in GPU infrastructure, HPC scheduling, or bare-metal hardware


Benefits:

  • Meaningful equity
  • Fully remote across North America
  • Full insurance coverage for you and your dependants


Salary:

  • $250,000 to $300,000 base salary
Ben Davies Director Global AI Infrastructure

Apply for this role