Senior ML Systems Engineer
We're partnering with a well-funded, research-driven organisation at the frontier of large-scale ML infrastructure. This is a hands-on technical role for someone who enjoys going deep on performance modelling, distributed systems, and real hardware behaviour — with direct influence over architecture decisions at scale.
What you'll do
Build simulation models for compute, memory, interconnect, and communication behaviour across large-scale ML systems
Develop tools to simulate training and inference workloads across distributed accelerator clusters
Model distributed execution patterns including collectives, synchronisation, and communication bottlenecks
Run experiments and benchmarks on real ML systems to calibrate and validate simulation models
Analyse end-to-end performance: throughput, latency, scaling efficiency, and cost/performance tradeoffs
Collaborate with hardware, software, networking, and ML teams to communicate findings through design recommendations
What we're looking for
Master's or PhD in CS, Electrical or Computer Engineering, or related field
Strong background in ML systems, distributed systems, performance engineering, or simulation
Experience analysing compute, communication, and memory behaviour in large-scale ML systems
Hands-on benchmarking, profiling, and measurement of ML systems
Familiarity with distributed training concepts: data/tensor/pipeline parallelism, collectives, synchronisation
Proficiency in Python, C++, or Rust
To find out more please reach out to Charles Duran.