Senior Solution Architect - AI Infrastructure
- $275,000 Base Salary
- San Francisco, California, United States
- Permanent
- 250000
- Artificial Intelligence
- AI Network
Looking for a role with plenty of growth opportunities?
Join an early-stage company building GPU compute infrastructure from the ground up. As one of the first technical hires, you'll help shape the company's core infrastructure while working directly with the founders in a high-ownership, hands-on environment.
This company is seeking a Senior Solutions Architect / Network Infrastructure Engineer to lead the design, deployment, and optimisation of GPU compute clusters across networking, software, and infrastructure. This customer-facing role combines hands-on engineering with technical leadership, including cluster architecture, fabric engineering, deployment, troubleshooting, and customer onboarding, while influencing the future direction of the platform.
Don’t miss out on this exciting opportunity and apply today!
Responsibilities:
- Design compute, storage, and networking topology for new cluster deployments
- Specify node configurations, redundancy, and scaling plans
- Make real-time build decisions on-site during deployment, not just on paper
- Bridge data center specifications to cluster deployment, including power whips and rack fit-out
- Design and personally implement GPU interconnect fabric using InfiniBand and/or RoCEv2
- Plan and validate bandwidth, topology, and east-west throughput at scale
- Diagnose and fix network bottlenecks under load, hands-on rather than in theory
- Get workloads running well across both CUDA and ROCm stacks
- Own driver and firmware compatibility, NCCL/RCCL tuning, and performance benchmarking
- Troubleshoot low-level issues directly across drivers, firmware, and fabric managers
- Lead technical onboarding and workload validation for new customers
- Act as the direct escalation point when a customer's cluster underperforms, diagnosing it yourself
- Translate customer workload requirements into concrete infrastructure and configuration decisions
- Document runbooks and SOPs that reflect how the infrastructure actually gets built and fixed
Skills/Must Have:
- Direct, hands-on experience standing up GPU clusters at production scale, with specific clusters you built rather than systems you oversaw
- Real InfiniBand or RoCEv2 design and implementation experience
- Comfortable working across both CUDA and ROCm ecosystems, or clearly demonstrated ability to ramp fast on a new GPU vendor stack
- Direct customer-facing technical experience, comfortable being the person a customer talks to when something's wrong
- Genuine preference for staying hands-on over moving into pure management
- Based in or willing to relocate to San Francisco (onsite)
Benefits:
- Founding-level ownership and visibility
- Direct access to and collaboration with the founders
- Onsite role in San Francisco
Salary:
- $275,000 Base Salary