Member of Technical Staff - AI Infrastructure
1729962
Posted: 06/08/2026
- $250,000 Base salary
- San Francisco, California, United States
- Permanent
- 250000
- Artificial Intelligence
- AI Software
Responsibilities:
- You will be building and evolving the core AI infrastructure software stack, orchestration, scheduling, cluster management, and the agentic operations layer that automates placement, healing, and recovery
- You will be contributing to the inference platform, including the token gateway, microVM pooling, and model serving infrastructure built for high utilisation and low cold-start latency
- You will be working across bare metal, Slurm, Kubernetes, and InfiniBand environments, writing software that abstracts the complexity for customers without hiding it from the engineers who need to debug it
- You will be building Infrastructure as Code tooling that automates deployment at scale and tracks changes across a fleet of 10,000+ servers
- You will be integrating with NVIDIA inference microservices, custom model pipelines, and customer-specific workloads — writing software that makes complex GPU infrastructure feel simple to consume
- You will be contributing to the SRE tooling layer; monitoring, continuous optimisation, GPU utilisation balancing, and automated remediation
Skills / Must Have:
- Strong systems software background; you will be comfortable at the intersection of distributed systems, infrastructure automation, and low-level performance
- Experience building orchestration, scheduling, or cluster management software at scale; Kubernetes, Slurm, or equivalent
- Solid understanding of GPU infrastructure; you will not need to crimp cables, but you will need to understand what InfiniBand, NVLink, and NCCL mean for the software you write
- Production engineering mindset; you will have shipped software that runs in anger on real infrastructure, not just in staging
- Strong Go, Python, or Rust; comfort with Linux systems programming underneath it all
Desirable Skills:
- Experience building agentic or autonomous systems that operate infrastructure without human-in-the-loop
- Background in inference serving, vLLM, TensorRT-LLM, NVIDIA NIM, or equivalent
- Exposure to bare metal provisioning, IPMI/BMC management, or large-scale server fleet tooling
- Familiarity with hypervisor or microVM technology (Firecracker, Cloud Hypervisor, or similar)
Benefits:
- Early-stage equity
- Direct access to leadership and genuine technical ownership
Salary:
- $250,000 Base salary
Ben Davies
Director Global AI Infrastructure