Check out our 2026 USA Salary Survey
Take me there

Technical Program Manager (Provider Management) - AI Infrastructure

1728861
  • $300,000 base salary
  • San Francisco, California, United States
  • Permanent
  • 300000
  • Artificial Intelligence
  • AI Data Center


Are you looking for an exciting new opportunity? 

Join a fast-scaling AI compute platform supporting leading AI labs and data centers by routing training and inference workloads across a global provider network. Operating in one of the fastest-growing technology markets, the organization is building the infrastructure that powers the next generation of AI applications.

This opportunity is ideal for a Provider Program Manager looking to build and own a function from the ground up. The role involves leading provider relationships, managing incident response, developing operational frameworks, and creating the processes and standards that will support the company’s continued growth. Working closely with a senior team, this role offers the opportunity to make a direct impact within a rapidly expanding AI infrastructure business.

If you would like to learn more about this opportunity, feel free to reach out and apply today!


Responsibilities:

  • You will be managing new provider and site onboarding, capacity expansions, hardware and network remediation, and ongoing provider relationships
  • You will be acting as incident commander during major provider-side incidents; assembling the right people across SRE, the provider, and the affected customer, owning communication throughout, and driving post-incident remediation
  • You will be keeping a real plan for each programme; milestones, owners, dependencies, risks — with visibility maintained across all parties including the provider
  • You will be holding providers to their contractual commitments through structured check-ins, tracked action items, and escalation to provider leadership when things slip
  • You will be coordinating internally to ensure provider issues don't stall; pulling in SRE, Engineering, Product, and Sales/CS as needed
  • You will be catching capacity and quality risk early, before it becomes a customer-facing problem
  • You will be turning what you learn into playbooks and provider-facing standards so each onboarding is faster and cleaner than the last


Skills / Must Have:

  • Several years running technical programmes in infrastructure, TPM or technical project management with genuine execution ownership, ideally involving external vendors or partners you didn't control
  • Incident management experience, you will have commanded or run point on production incidents involving multiple organisations, and you are comfortable with the off-hours reality that comes with that
  • Enough technical depth to hold your own with data-centre engineers and SREs on GPUs, networking, and storage, you will not need to debug an InfiniBand fabric yourself, but you will need to follow the conversation and know when someone is hand-waving
  • A track record of getting teams you don't manage, including external partners, aligned and moving, including through rough patches where the relationship is strained
  • Calm, direct communication, you will be the person who delivers bad news early rather than good news late


Benefits:

  • Sock options 
  • Comprehensive healthcare, dental, and vision (you and dependents)
  • 401(k)
  • Unlimited PTO


Salary:

  • $300,000 base salary 
Ben Davies Director Global AI Infrastructure

Apply for this role