join a specialist Public Sector Technology team. This hands-on role focuses on building evaluation frameworks, tooling and harnesses for AI systems, particularly LLM and agentic AI solutions. The role requires ownership, experimentation and rapid delivery. Key Responsibilities Design evaluation frameworks and tooling Develop evaluation harnesses Evaluate model and agentic … agentic AI Experience with AI evaluation frameworks Ability to code independently Strong problem-solving and communication skills Comfortable working autonomously. Technical Experience AI/LLM Evaluation, RAG Evaluation, Ragas, Agentic AI Evaluation, Evaluation Harnesses, LLM Testing and Benchmarking, Prompt Engineering, Python Development, Model and Agent Integration. Ideal Background AI Evaluation ...