Launch your career with Hamilton Barnes’ Graduate Hub
Take me there

Principal Solutions Architect - AI Infrastructure

1737494
  • £150,000 to £180,000 base plus performance-related bonus
  • United Kingdom
  • Permanent
  • 150000
  • Artificial Intelligence


Looking for a role with plenty of growth opportunities?

Join a rapidly growing cloud and AI infrastructure company helping enterprises modernise Kubernetes operations, infrastructure orchestration, and AI/ML workloads across hybrid and multi-cloud environments, working with customers building next-generation cloud-native and AI-powered platforms.

This pioneering Kubernetes-native platform provider is hiring a Principal Solutions Architect to anchor its EMEA post-sales function. The ideal candidate will act as the primary technical authority for strategic accounts across EMEA, working at the intersection of distributed systems, GPU infrastructure, and MLOps, with genuine autonomy and direct exposure to large-scale GPU fabric and distributed training environments. The role also carries a mentorship remit over junior team members within a small, senior Solutions Architecture function.

If you would like to learn more about this opportunity, feel free to reach out and apply today!


Responsibilities:

  • Design end-to-end AI/ML platform architectures spanning inference, training and data pipelines
  • Develop reference architectures for GPU cluster deployment, LLM serving and multi-tenant ML infrastructure
  • Advise on GPU fabric topology (NVLink, InfiniBand, RoCEv2) for distributed training workloads
  • Act as the primary technical advisor and escalation point, leading root cause analysis on complex production issues
  • Deliver technical workshops, proofs of concept and executive-level presentations to senior stakeholders
  • Design observability strategies using DCGM, OpenTelemetry, eBPF and GPU metrics pipelines
  • Partner with customer platform, MLOps and data science stakeholders to translate requirements into architecture
  • Feed customer insights back into the product and engineering roadmap, and mentor junior Solutions Architects


Skills/Must Have:

  • Experience: 8+ years in infrastructure, platform or solutions engineering, including 3+ years specifically in AI/ML infrastructure or MLOps
  • Core Tech/Domain: Deep hands-on Kubernetes expertise (cluster lifecycle, workloads, operators, RBAC) and direct NVIDIA GPU infrastructure experience (H100/H200/B200 preferred)
  • Methodology/Protocols: Distributed training knowledge (NCCL, tensor/pipeline parallelism), LLM inference serving and optimisation (vLLM, NIM, TGI), and GPU networking (GPU Operator, MIG, SR-IOV, IB/RoCEv2)
  • Soft Skills: Strong customer-facing and advisory communication skills, credible at both engineer and executive level


Desirable Skills:

  • Run:AI or Slurm scheduling experience
  • PyTorch/TensorFlow familiarity
  • Public cloud experience (AWS/Azure/GCP) including networking, IAM and managed Kubernetes
  • CKA, CKAD, AWS Solutions Architect, Azure Solutions Architect or GCP Professional Cloud Architect certification
  • Multi-tenant GPU isolation knowledge (SR-IOV VFs, DPU offload)


Benefits:

  • Performance-related bonus
  • Private healthcare
  • Pension contribution
  • Training and certification allowance
  • Remote-first, flexible working across EMEA


Salary:

  • £150,000 to £180,000 base plus performance-related bonus

Jamie Maher Head of AI Infrastructure

Apply for this role