Member of Technical Staff (AI Platform) - AI Infrastructure
- $250,000 to $300,000 base salary
- San Francisco, California, United States
- Permanent
- 300000
- Artificial Intelligence
Keen to join a company that champions growth and development?
Join one of the most exciting AI infrastructure companies in the market, building a platform that deploys and operates large-scale GPU clusters for some of the world's leading AI labs. Working with tens of thousands of GPUs across a global network of data centres, the organisation delivers the infrastructure powering next-generation AI training and inference at scale.
This role offers the opportunity to work on complex, production-scale AI infrastructure without the layers of bureaucracy often found in larger organisations. You'll have the freedom to shape how the platform evolves while solving engineering challenges that have a direct impact on customers.
Ready to make a move? Get in touch and apply today!
Responsibilities:
- You will be building and operating the automated systems that take a cluster from bare machines to customer-ready, end to end
- You will be managing machine lifecycle between tenants: join, wipe, verify, rejoin, reliably and at scale
- You will be operating Kubernetes and Postgres across the fleet, keeping critical infrastructure healthy under real production load
- You will be contributing to custom Kubernetes operators that power the control plane
- You will be scaling clusters from tens of nodes to thousands, solving the problems that only show up at the edges of scale
- You will be participating in on-call rotations and responding to production incidents with the depth of knowledge to resolve them quickly
Skills / Must Have:
- A track record of impressive technical work you can speak to in depth, the years matter less than the quality of the work and its real-world impact
- 2+ years of on-call experience for critical production services
- Deep Kubernetes experience, not just consuming it but understanding what is happening underneath
- Strong Linux fundamentals across kernel, cgroups, containers, networking, and storage
- Experience operating databases, monitoring, CI/CD, and cloud infrastructure at scale
- Familiarity with fleet management and capacity planning
Desirable Skills:
- Experience writing Kubernetes operators
- Background in GPU infrastructure, HPC scheduling, or bare-metal hardware
Benefits:
- Meaningful equity
- Fully remote across North America
- Full insurance coverage for you and your dependants
Salary:
- $250,000 to $300,000 base salary