Machine Learning Operations Engineer
- Location
- Greater London, England, United Kingdom
batch and online inference services in containerised cloud environments Define and meet availability, latency, throughput and recovery objectives for ML services Monitor service health, infrastructure, data-quality signals, data drift, prediction drift and model performance decay Establish dashboards, alerting and operational runbooks so failures are detected and resolved quickly … Support automated or controlled retraining, model promotion, rollback and model retirement Debug production issues across model, application, infrastructure and critical data-dependency layers Reliability, security and engineering quality Improve system robustness, scalability and cost efficiency through automation, observability and infrastructure as code Write production ...