Senior Solutions Architect (Post Sales) - AI Infrastructure
- £150k Base
- United Kingdom
- Permanent
- 150000
- Artificial Intelligence
Looking for a role with plenty of growth opportunities?
Join a high-growth GPU cloud and AI infrastructure platform provider delivering Kubernetes-native, GPU-accelerated solutions for enterprise AI and machine learning workloads. The organisation is experiencing rapid growth as demand for AI compute continues to accelerate, offering the opportunity to work with cutting-edge infrastructure and large-scale distributed training environments.
This opportunity is ideal for a Senior Solutions Architect looking to play a key customer-facing role within a senior post-sales team across EMEA. The role involves advising enterprise customers on designing, deploying, and scaling production AI/ML platforms while working closely with product and engineering teams. It offers significant technical ownership, direct influence, and the opportunity to help shape the platform as both its capabilities and customer base continue to grow.
Ready to take your expertise to the next level? Get in touch today!
Responsibilities:
- Design end-to-end AI/ML platform architectures spanning inference, training, and data pipelines for enterprise customers
- Develop reference architectures for GPU cluster deployment, LLM serving, and multi-tenant ML infrastructure
- Advise on GPU fabric topology, including NVLink, InfiniBand, and RoCEv2, for distributed training environments
- Act as the primary technical advisor and escalation point for assigned customers, leading root cause analysis on complex production issues
- Deliver technical workshops, proof-of-concept engagements, and executive-level presentations on AI infrastructure strategy
- Design observability strategies across DCGM, OpenTelemetry, eBPF, and GPU metrics pipelines
- Partner with customer platform, MLOps, and data science stakeholders to translate workload requirements into scalable architecture
- Feed customer insights back into the product and engineering roadmap, and mentor junior members of the Solutions Architecture team
Skills/Must Have:
- Experience: 8+ years in infrastructure, platform, or solutions engineering, including 3+ years focused specifically on AI/ML infrastructure or MLOps
- Core Tech/Domain: Deep, hands-on Kubernetes expertise (cluster lifecycle, workloads, operators, RBAC) plus direct experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred)
- Methodology/Protocols: Working knowledge of distributed training concepts (NCCL, tensor and pipeline parallelism), LLM inference serving and optimisation (vLLM, NIM, TGI), and GPU networking (GPU Operator, MIG, SR-IOV, IB/RoCEv2)
- Soft Skills: Confident, credible communicator able to run technical discussions with engineers through to executive stakeholders; strong troubleshooting mindset and customer advisory presence
Desirable Skills:
- Experience with Run:AI or Slurm for GPU scheduling and workload optimisation
- Familiarity with PyTorch or TensorFlow
- Public cloud experience (AWS, Azure, or GCP) across networking, IAM, and managed Kubernetes
- Certifications such as CKA, CKAD, AWS Solutions Architect, Azure Solutions Architect, or GCP Professional Cloud Architect
- Understanding of multi-tenant GPU isolation (SR-IOV VFs, DPU offload)
Benefits:
- Competitive base salary with performance-related bonus
- Remote-first/flexible working across EMEA
- Private healthcare and pension contribution
- Training and certification allowance
- Genuine exposure to cutting-edge GPU and AI infrastructure at enterprise scale
Salary:
- £150k Base