Check out our 2026 USA Salary Survey
Take me there

Software Engineer - AI Infrastructure

1722654
  • $300,000 - $500,000+
  • San Francisco, California, United States
  • Permanent
  • 300000
  • Artificial Intelligence
  • AI Software


Are you looking for an exciting new opportunity? 

Join a stealth-mode hyperscale infrastructure startup building a 300MW+ AI compute platform designed to power the next generation of large-scale training, inference, and sovereign cloud environments. Backed by significant capital and industry-leading talent, the company is creating one of the most ambitious AI infrastructure projects currently under development.

The Principal Software Engineer will take ownership of the software architecture that connects thousands of GPUs, high-performance networking fabrics, storage systems, schedulers, and cloud orchestration platforms. This hands-on technical leadership role will be responsible for designing and building the infrastructure software that powers one of the world's largest AI compute environments. The successful candidate will solve complex challenges across distributed systems, cluster orchestration, workload scheduling, GPU resource management, infrastructure automation, and platform scalability at exceptional scale. 

This is a unique opportunity to design software that enables tens of thousands of GPUs to operate as a unified platform while helping shape the future of hyperscale AI infrastructure and supporting next-generation AI workloads. If you would like to learn more about this opportunity, feel free to reach out and apply today!


Responsibilities:

  • Design and develop core platform software powering large-scale HPC and AI infrastructure environments.
  • Build distributed systems responsible for GPU scheduling, resource allocation, workload orchestration, and cluster management.
  • Develop software integrations across Slurm, Kubernetes, cloud orchestration platforms, and custom scheduling frameworks.
  • Create infrastructure services for provisioning, telemetry, observability, health monitoring, and automated remediation.
  • Optimize software performance across high-bandwidth GPU fabrics, storage platforms, and distributed compute environments.
  • Collaborate with networking, platform, storage, and hardware teams to maximize cluster efficiency and utilization.
  • Design APIs, automation frameworks, and developer tooling that improve infrastructure operability and customer experience.
  • Contribute to long-term platform architecture supporting hyperscale AI and HPC deployments.
  • Drive software reliability, scalability, and operational excellence across thousands of nodes and GPUs.
  • Mentor engineers and establish software engineering standards across the infrastructure organization.


Skills / Must Have:

  • 8+ years of experience developing software for distributed systems, infrastructure platforms, HPC environments, or large-scale cloud services.
  • Strong software engineering expertise in Go, Rust, C++, Python, or similar systems-level languages.
  • Experience building infrastructure software rather than traditional business applications.
  • Deep understanding of Linux systems, operating systems internals, networking, and distributed computing concepts.
  • Experience with HPC schedulers such as Slurm or large-scale Kubernetes environments.
  • Strong knowledge of infrastructure automation, APIs, service-oriented architectures, and cloud-native platforms.
  • Experience designing highly scalable, fault-tolerant distributed systems.
  • Understanding of observability, telemetry, monitoring, and operational tooling at scale.
  • Ability to operate as a technical leader while remaining highly hands-on.


Highly Desirable:

  • Experience supporting NVIDIA GPU environments including H100, H200, B200, GB200, DGX, or HGX platforms.
  • Knowledge of CUDA, NCCL, NVLink, NVSwitch, GPUDirect RDMA, or GPU resource management.
  • Experience with InfiniBand, RoCE, Spectrum-X, Cumulus Linux, or high-performance networking fabrics.
  • Familiarity with AI infrastructure platforms, distributed training systems, and ML workload orchestration.
  • Experience working within hyperscalers, GPU cloud providers, HPC vendors, AI labs, or large-scale infrastructure startups.
  • Contributions to open-source infrastructure, Kubernetes, Slurm, or distributed systems projects.


Benefits:

  • Significant early-stage equity participation.
  • Founding engineer opportunity within a next-generation hyperscale AI infrastructure platform.
  • Direct influence over architecture decisions across software, infrastructure, and AI platform design.
  • Opportunity to build systems operating at unprecedented GPU scale.
  • Flexible remote or hybrid working arrangements.
  • Work alongside industry leaders in AI infrastructure, cloud computing, and hyperscale operations.
  • Significant Equity Package
  • Executive-Level Incentives Depending on Experience


Salary:

  • $300,000 - $500,000+
Ben Davies Director Global AI Infrastructure

Apply for this role