systems stack.Build robust systems that ensure reproducible, debuggable, large-scale runs.Qualifications:Strong engineering experience in large-scale distributed training or HPC systems.Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops.Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar).Comfort debugging performance issues across … following would also be good to have for this role:Experience with training LLMs or other large transformer architectures.Contributions to ML frameworks (PyTorch, JAX, DeepSpeed, Megatron, xFormers, etc.).Familiarity with evaluation and serving frameworks (vLLM, TensorRT-LLM, custom KV caches).Experience with data pipeline optimization, sharded datasets, or caching strategies.Background ...