Machine Learning Performance Engineer
- Location
- Greater London, England, United Kingdom
about efficient large-scale training, low-latency inference in real-time systems and high-throughput inference in research. Part of this is improving straightforward CUDA, but the interesting part needs a whole-systems approach, including storage systems, networking and host- and GPU-level considerations. Zooming in, we also want … end. Low-level GPU knowledge of PTX, SASS, warps, cooperative groups, Tensor Cores and the memory hierarchy. Debugging and optimisation experience using tools like CUDA GDB, NSight Systems, NSight Compute-sight-systems and nsight-compute. Library knowledge of Triton, CUTLASS, CUB, Thrust, cuDNN and cuBLAS. Intuition about the latency ...