moving parts: rollouts, replay buffers, reward signals, data filtering, policy updates, evaluation loops, and failure analysis Are familiar with deep learning frameworks such as PyTorch or JAX Are proficient in Python, including concurrency, asynchronous programming, multiprocessing, and performance optimization Can debug distributed GPU workloads across CUDA runtime, container runtime, driver … versions, NCCL or equivalent communication layers, networking, storage, scheduling, and checkpointing Have experience with profiling tools across the stack, for example py‐spy, PyTorch profiler, Nsight, perf, tracing, metrics, logs, or custom instrumentation Have experience with inference stacks such as vLLM, SGLang, TensorRT‐LLM, Dynamo, or custom serving infrastructure ...