Software Development Engineer - AI/ML, AWS Neuron, Multimodal Inference
Cambridge, Cambridgeshire, United Kingdom
Amazon
engineers and runtime engineers to create, build and tune distributed inference solutions with Trn1. Experience optimizing inference performance for both latency and throughput on these large models using Python, Pytorch or JAX is a must. Deepspeed and other distributed inference libraries are central to this and extending all of this for the Neuron based system is key. Key job responsibilities … This role will help lead the efforts building distributed inference support into Pytorch, Tensorflow using XLA and the Neuron compiler and runtime stacks. This role will help tune these models to ensure highest performance and maximize the efficiency of them running on the customer AWS Trainium and Inferentia silicon and the TRn1 , Inf1 servers. Strong software development using C Python More ❯
Employment Type: Permanent
Salary: GBP Annual
Posted: