AI Quality Analyst - Outside IR35
We're looking for an AI Quality Analyst to join a cutting-edge AI/ML team and play a pivotal role in testing, challenging and improving the next generation of intelligent systems.
This is not a traditional QA role.
You'll sit at the intersection of AI, quality engineering, data and product, investigating how AI systems behave in the real world
AI Quality Analyst - Contract
£600 per day | Outside IR35 | 4-Day Week | Hybrid | Initial 6 Months
The OpportunityWe're looking for an AI Quality Analyst to join a cutting-edge AI/ML team and play a pivotal role in testing, challenging and improving the next generation of intelligent systems.
This is not a traditional QA role.
You'll sit at the intersection of AI, quality engineering, data and product, investigating how AI systems behave in the real world - including where they fail, where they produce unexpected results, where they can be manipulated, and where their outputs fall short of expectations.
You'll have the freedom to think beyond predefined test cases, challenge assumptions and uncover the edge cases that others miss.
What You'll Be DoingTest & Evaluate AI
- Design and develop evaluation strategies, benchmarks and test suites for AI/ML models.
- Assess model accuracy, consistency, reliability and real-world performance.
- Identify edge cases, failure modes, hallucinations, bias and unexpected behaviour.
- Evaluate AI and potentially agentic systems across a range of scenarios and use cases.
Break Things - Intentionally
- Conduct exploratory and adversarial testing to uncover weaknesses.
- Participate in AI red-teaming, deliberately challenging models with unusual, ambiguous and adversarial inputs.
- Investigate why models fail rather than simply recording that they have failed.
- Identify patterns that can inform model, data and product improvements.
Turn Testing Into Insight
- Perform detailed error and failure analysis.
- Categorise and prioritise defects according to severity, impact and risk.
- Analyse evaluation results and translate findings into clear recommendations.
- Communicate complex AI behaviours to both technical and non-technical stakeholders.
Build Better Evaluation Data
- Source, curate and validate datasets used for AI evaluation.
- Support annotation and ground-truth creation.
- Ensure evaluation data is accurate, representative and fit for purpose.
- Identify data quality issues that could influence model performance.
Build the Quality Framework
- Develop and maintain repeatable AI evaluation processes.
- Create automated tests and tooling where appropriate.
- Set up and maintain test and demonstration environments.
- Help evolve the team's AI quality, testing and evaluation methodology.
We're looking for someone who is naturally curious about how AI behaves and enjoys finding things that other people don't.
You could come from a traditional QA/testing background, AI/ML, data analysis, automation or another technical discipline. What matters most is your ability to combine technical testing expertise with analytical thinking and curiosity.
You'll ideally have:
- Experience in QA, software testing, quality engineering or data analysis.
- A good understanding of the machine learning/AI lifecycle and common AI failure modes.
- Experience with data validation, annotation or large datasets.
- Strong analytical and investigative skills.
- Experience with test automation tools such as Playwright, Selenium, Cypress or REST-assured.
- Experience testing web and/or mobile applications.
- Python or another scripting language for automation and data manipulation.
- Experience using tools such as Jira and test case/defect management platforms.
- Excellent written and verbal communication skills.
We'd love to hear from candidates with experience in:
- Testing or evaluating AI/ML systems
- LLM evaluation and model benchmarking
- Agentic AI testing
- AI red teaming / adversarial testing
- Prompt evaluation and testing
- Model performance metrics
- Data quality and integrity
- Computer vision
- AI safety and responsible AI
- Automated evaluation frameworks
Robert Walters Operations Limited is an employment business and employment agency and welcomes applications from all candidates