work is externally credible. Partner directly with MLEs to ensure research prototypes become usable production components. Define and execute research programs in efficient LLM and VLM inference with measurable production impact. Invent, evaluate, and productionize methods for quantization, QAT, distillation, speculative decoding, KV-cache reuse, KV-cache compression, long-context … EMNLP, ASPLOS, OSDI, SOSP, ISCA, HPCA, or comparable venues. Experience deploying ML models or inference optimizations in production. Experience with vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, FlashAttention, FlashInfer, Triton, CUDA, or PyTorch internals. Experience with post-training, SFT, DPO, RLHF, RLAIF, preference optimization, or synthetic data generation when connected ...