different range of applications. This involves a blend of technical expertise and collaborative problem-solving to ensure both efficiency and quality throughout the entire LLM deployment lifecycle. The role includes opportunities for both IC and TL opportunities, and is open to both Software Engineering and Research Engineering backgrounds. There … bottlenecks. Experience with distributed ML systems optimization and parallelism (e.g., data, model, or pipeline parallelism). Familiarity with writing performance-optimized kernels. Understanding of LLM architecture and inference performance dynamics (e.g., Transformer models, memory bandwidth and compute bounds, KV cache scaling). Responsibilities: Collaborate closely with Research teams to understand ...