Launch your career with Hamilton Barnes’ Graduate Hub
Take me there

Member of Technical Staff (Inference) - AI Infrastructure

1738119
  • $250,000
  • United States
  • Permanent
  • 250000
  • Artificial Intelligence


Ready to take the next step in your career?

Join a high-growth infrastructure company that makes Apple hardware available and performant at data centre scale for AI workloads, building cloud platforms that expose high-performance hardware as elastic compute through developer-friendly interfaces.

The organization is currently on the lookout for an Inference Engineer to establish the technical foundation for inference, defining the architecture, writing performance-critical code, and taking the system through deployment and production operation. The ideal candidate will work across kernels, compilers, inference engines, memory management, model parallelism, networking, scheduling, and serving APIs, collaborating closely with Platform, Fleet, and Network Infrastructure engineering to turn individual machines into a reliable inference service, owning the tradeoffs among latency, throughput, model quality, reliability, and cost.

Ready to make a move? Get in touch and apply today!


Responsibilities:

  • Build and optimize the inference engine. Implement model execution, continuous batching, prefill and decode scheduling, KV-cache management, prefix caching, and memory allocation. Bring new model architectures into production and implement optimizations such as quantization, speculative decoding, and chunked prefill.
  • Write performance-critical kernels and runtime code. Implement and optimize operations such as matrix multiplication, attention, and mixture-of-experts execution. Work across Metal, MLX, and compiler and runtime internals to improve memory access, operator fusion, synchronization, and CPU/GPU execution. Validate numerical correctness alongside performance.
  • Design for Apple Silicon. Build around unified memory, memory bandwidth, compute capabilities, and operating-system behavior. Profile execution on real hardware, diagnose bottlenecks, and choose model representations and execution strategies based on measured results.
  • Build distributed inference and its communication layer. Implement model sharding and parallel execution across machines. Develop and optimize collective communication, data movement, and the overlap between computation and communication. Work with network engineers on transport performance, topology, and failure handling.
  • Own inference scheduling and routing. Build request queues, admission control, cache-aware routing, model placement, and autoscaling. Manage capacity and tenant fairness while meeting latency and throughput targets across changing traffic patterns and model sizes.
  • Ship a production inference service. Build model loading and deployment workflows, serving APIs, streaming responses, cancellation, and backpressure. Integrate the runtime with fleet and platform systems. Own observability, safe releases, incident diagnosis, and recovery for the inference stack.
  • Make performance reproducible. Build benchmarks and profiling tools that measure time to first token, inter-token latency, tail latency, throughput, memory use, power, and cost per token. Test realistic workloads, concurrency, and context lengths. Catch performance and model-quality regressions before release.
  • Set the engineering direction. Turn customer workloads into technical priorities, choose and contribute to open-source projects, and establish clear interfaces across the stack. Use coding agents to investigate, implement, and validate changes, backed by correctness checks and reproducible benchmarks.


Skills/Must Have:

  • A record of building and optimizing production inference engines, GPU compute software, or distributed ML systems, with substantial depth in at least one and hands-on work across multiple layers.
  • Strong systems programming skills in C++ or Rust, proficiency in Python, and experience working inside performance-critical libraries and runtimes.
  • A practical understanding of transformer inference, including attention, prefill and decode, batching, KV caches, quantization, and their compute and memory costs.
  • Experience with GPU programming and performance analysis, including memory hierarchies, parallel execution, synchronization, and numerical precision.
  • Strong distributed-systems fundamentals, including scheduling, concurrency, networking, failure handling, and resource management.
  • Experience taking software from architecture through production deployment, debugging, and ongoing operation.
  • The ability to turn an ambiguous performance problem into a measured bottleneck, an implementation, and a verified improvement.
  • The judgment and ownership to establish a new technical area, prioritize the work, and explain tradeoffs clearly to engineers and customers.


Benefits:

  • Full Benefits 


Salary:

  • $250,000
Sam Merchant Principal AI Infrastructure Consultant

Apply for this role