ML Ops Engineer
- Location
- Greater London, England, United Kingdom
model validation, monitoring for model drift, data drift, and latency bottlenecks. Monitor and optimise cloud spend across high-cost GPU/CPU clusters across AWS, GCP, or Azure. Implement auto-scaling strategies, spot instance policies, and dynamic resource allocation to eliminate infrastructure waste. Establish benchmarking and telemetry to track unit … with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins. Experience with vLLM, Ray, MLflow, LangChain/LangSmith, DeepSpeed, or Hugging Face TGI. Solid background in AWS/GCP/Azure, Kubecost, and GPU cost optimisation techniques. Strong skills in Python, Bash, or Go; deep knowledge of Linux kernel tuning and performance ...