graphs, test time compute, CoT pipelines, tool use & API calling, etc.)Data-centric Machine Learning (Synthetic data, curriculum learning, learned data mixtures, etc.)Evaluation (Benchmarking best practices, humans/LLMs as a judge, red teaming/adversarial testing, hallucination detection, etc.)We work collaboratively with TR Labs (TR's applied … other RLHF methodsHands-on experience implementing and scaling supervised fine-tuning, preference learning, and reinforcement learning pipelines for LLMsExperience building LLM evaluation frameworks, benchmarking systems, or automated testing pipelinesHands-on experience with agentic workflows, tool-using AI systems, or multi-agent coordination (examples include: langgraph, AutoGPT, LLamaIndex)Experience with data ...