1 to 25 of 136 Permanent AIOps Jobs in London

Architect & Delivery Lead (68018)

Location
Greater London, England, United Kingdom
model transformations for global financial services clients Published thought leadership in SRE, platform engineering, or IT transformation Experience with AI/ML‐driven operations (AIOps) and emerging automation technologies ITIL 4 Managing Professional or Strategic Leader certification Benefits We help take care of your today and tomorrow with industry‐leading ...

Director of Platform Engineering

Location
Greater London, England, United Kingdom
particularly where operational teams are required to provide evidence for access, change, incident, backup, resilience and monitoring controls. Practical experience applying AI, Agentic AI, AIOps or intelligent automation to operational use cases such as incident triage, anomaly detection, runbook automation, knowledge search, root cause support or service desk workflows. Knowledge ...

AI Platform & Site Reliability Engineering Consultant/Senior Consultant - Cloud Operating Model

Location
Greater London, England, United Kingdom
cloud operating model transformations, using operational metrics, serviced data and insights to recommend improvements. Support the introduction of intelligent monitoring, service analytics and AIOps-aligned practices where appropriate.· Deliver Business Outcomes: Use skills in strategy development, service management, agile delivery, DevOps, business process mapping and organisational design to deliver digitally ...

Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President

Location
Greater London, England, United Kingdom
building, and deploying Large Language Model (LLM) orchestration frameworks (e.g., LangChain, Temporal, or custom agentic loops) to coordinate multi-step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root-cause analysis, anomaly detection ...

Director of Platform Engineering

Location
Greater London, England, United Kingdom
Google GKE Preferred: GitOps tooling such as Argo CD or Flux Preferred: ISO 27001, SOC 2 or similar assurance frameworks Preferred: AI, Agentic AI, AIOps or intelligent automation experience Preferred: Observability and reliability practices, including SRE principles, service-level objectives, alert tuning, capacity planning and production readiness reviews Preferred: SaaS ...

Performance & Observability Engineer

Location
Greater London, England, United Kingdom
Increase the percentage of incidents identified before user impact. Error Budget Utilization (%) – Ensure system reliability is balanced with innovation velocity. Automation & AI-Driven Observability (AIOps) % of Issues Resolved via Automated Remediation – Reduce manual intervention in incident response. Reduction in On-Call Burden (%) – Minimize alerts requiring human intervention. Anomaly Detection Accuracy ...

Lead AI Engineer

Hiring Organisation
Kainos
Location
London, United Kingdom
Salary
£ 100 K
advanced AI solutions leveraging state-of-the-art machine learning, generative and agentic AI technologies. You will drive the adoption of modern AI frameworks, AIOps best practices and scalable cloud-native architectures. Your role will involve hands-on technical leadership, collaborating with customers to translate business challenges into trustworthy ...

AI & SRE Consultant

Location
City Of London, England, United Kingdom
transformation. The Opportunity You will help clients assess, redesign and optimise cloud operating models, ensuring they are scalable, resilient and ready for Generative AI, AIOps and intelligent automation. Key Responsibilities Advise clients on cloud transformation and cloud operating models. Assess existing service management and operational environments. Design future-state operating ...

site reliability engineer

Location
Greater London, England, United Kingdom
culture Nice to have: Experience in financial services or other highly regulated, mission‐critical environments, Certifications in cloud technologies such as AWS, Exposure to AIOps platforms or advanced observability tooling #J-18808-Ljbffr ...

AI & SRE Consultant

Hiring Organisation
Akkodis
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£88000 - £96000/annum
transformation. The Opportunity You will help clients assess, redesign and optimise cloud operating models, ensuring they are scalable, resilient and ready for Generative AI, AIOps and intelligent automation. Key Responsibilities Advise clients on cloud transformation and cloud operating models. Assess existing service management and operational environments. Design future-state operating ...

Director of Site Reliability Engineering

Location
Greater London, England, United Kingdom
culture Nice to have Experience in financial services or other highly regulated, mission‐critical environments Certifications in cloud technologies, such as AWS Exposure to AIOps platforms or advanced observability tooling We offer EPAM Employee Stock Purchase Plan (ESPP) Protection benefits including life assurance, income protection and critical illness cover Private ...

Senior Manager, Network & Security Solutions Architect – Cloud & Edge, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
London, United Kingdom
Salary
£ 120 K
Ansible, Python scripting, AWS CloudFormation, Azure Resource Manager (ARM), and CI/CD pipelines to support automated deployment and management of network infrastructure.Experience with AIOps platforms, network analytics, machine learning for anomaly detection, or advanced security analytics in a networking context.Experience integrating network solutions with IT Service Management (ITSM) platforms ...

Lead Product Manager AIOPs

Hiring Organisation
S&P Global
Location
London, United Kingdom
Salary
£ 80 K
About the Role:Grade Level (for internal use):11The Team:The AIOps team is responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights.DTS Platform & Tools – Service Enablement … serve as thought leaders in AIOps, partnering across IT Operations, SRE, engineering, infrastructure, service management, and application teams to solve enterprise operational challenges. Our mission is to improve reliability, reduce operational complexity, optimize technology investments, and enable more proactive and resilient technology operations by applying AI.Responsibilities and Impact ...

Lead Product Manager AIOPs

Location
Greater London, England, United Kingdom
About the Role: Grade Level (for internal use): 11 The Team: The AIOps team is responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools … Service Enablement: We serve as thought leaders in AIOps, partnering across IT Operations, SRE, engineering, infrastructure, service management, and application teams to solve enterprise operational challenges. Our mission is to improve reliability, reduce operational complexity, optimize technology investments, and enable more proactive and resilient technology operations by applying AI. Responsibilities ...

Global IT Senior Director and Platform Team Lead - Gen AI Platforms and Agentic Development Kit

Hiring Organisation
The Boston Consulting Group
Location
London, United Kingdom
Salary
£ 100 K
platform teams.Establish and run enterprise‐grade AI observability frameworks to monitor performance, safety, and interpretability.Continuously enhance LLM Ops, Evaluation and Quality pipelines.Ensure robust AIOps/DevSecOps, CI/CD, Infrastructure‐as‐Code, and automated delivery pipelines.What You'll BringLeadership & DeliveryOver 15 years of experience in engineering leadership, platform or product ...

Senior Python Engineer

Location
Greater London, England, United Kingdom
deployments of AI systems. Provide guidance and mentoring to other team members on best practices in AI engineering. Use best practices (e.g., MLOps, AIOps) to improve products/services and processes related to AI. Optimise existing model serving and data pipelines to meet changing performance and security requirements. Hold requirements ...

AWS Data Platform Architect, Technology Consulting- London, Leeds, Manchester or Newcastle

Hiring Organisation
Momentum Worldwide
Location
London, United Kingdom
Salary
£ 70 K
innovation and market-leading practices across multiple clients and industries- Strong understanding of cloud-centric solution design, the full data lifecycle, and MLOps/AIOps- Experience mentoring engineers and architects, building capability in junior team members and client counterparts- You will be required to hold active SC clearance for this ...

AI Technical Platform Leader

Hiring Organisation
Willis Towers Watson
Location
London, United Kingdom
Salary
£ 80 K
responsible AI and security controls into the AI platform by design, including identity, data protection, content safety and auditability.• Define and implement AI operations (AIOps) practices for AI Platforms: Monitoring and observability, forensics on security anomalies, token/quota capacity planning, and workload optimization.• Serve as a senior technical escalation ...

Product Manager - Data

Location
Greater London, England, United Kingdom
Ubuntu Pro, Compliance, Standards, Security Engineering, and Managed Services on cloud and on prem AI/ML & MLOps - Open source AI/ML solutions, AIOps automation, model lifecycle management, Kubeflow, MLFlow, KServe, and AI infrastructure on cloud and edge IoT - Ubuntu on embedded devices and/or edge servers, device ...

Vice President, Site Reliability Engineering

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
aligned to operational and business priorities. Build and optimize monitoring, observability, and alerting capabilities using tools such as Prometheus, Grafana, AppDynamics, and Splunk. Apply AIOps capabilities to improve event correlation, anomaly detection, root cause analysis, predictive insights, and proactive issue prevention. Partner with engineering, infrastructure, production support, security, and risk … across technical and non-technical stakeholders. Preferred Qualifications Experience building centralized internal platforms or shared engineering services for operational or enterprise users. Experience applying AIOps, machine learning, or intelligent automation within production support or reliability engineering environments. Exposure to CI/CD pipelines, infrastructure as code, API-driven automation ...

Vice President, Site Reliability Engineering

Hiring Organisation
The Bank of New York Mellon
Location
London, United Kingdom
Salary
£ 80 K
health measures aligned to operational and business priorities.Build and optimize monitoring, observability, and alerting capabilities using tools such as Prometheus, Grafana, AppDynamics, and Splunk.Apply AIOps capabilities to improve event correlation, anomaly detection, root cause analysis, predictive insights, and proactive issue prevention.Partner with engineering, infrastructure, production support, security, and risk teams … collaborate effectively across technical and non-technical stakeholders.Preferred QualificationsExperience building centralized internal platforms or shared engineering services for operational or enterprise users.Experience applying AIOps, machine learning, or intelligent automation within production support or reliability engineering environments.Exposure to CI/CD pipelines, infrastructure as code, API-driven automation, and modern software ...

Head of EMEA Infrastructure

Location
Greater London, England, United Kingdom
teams are equipped and engaged throughout the change journey. AI-Driven Operations: Champion the adoption of AI-driven operational tooling across EMEA infrastructure, including AIOps platforms for predictive monitoring, automated incident response, and intelligent capacity management. Evaluate and drive use cases where AI enhances operational efficiency, reduces manual toil ...

Infrastructure Architect

Location
Greater London, England, United Kingdom
data integration, secure access and responsible AI controls. Evaluate emerging Microsoft capabilities, including Copilot and AI services, and define group‐wide adoption patterns. Promote AIOps to improve incident detection, automated remediation and operational insight. Stay current with technology trends, Microsoft roadmaps and emerging capabilities to inform future architecture decisions. Role ...

Director of Software Engineering (AIOps) - Executive Director

Location
City Of London, England, United Kingdom
Director of Software Engineering (AIOps) - Executive Director LONDON, LONDON, United Kingdom and 1 more Job Information Job Identification 210707625 Job Category Software Engineering Business Unit Corporate Sector Posting Date 08/05/2026, 05:42 PM Locations 4 John Carpenter St, London, Greater London, EC4Y 0JP, GB 315 Argyle … Engineers across the whole firm Innovates, designs and delivers technical solutions that can be leveraged across multiple businesses and domains, this will include AIOps strategy for the future Sets direction and governance for agentic AI-enabled engineering and SDLC/TLM automation within a technical area to drive measurable improvements ...

Director of Software Engineering (AIOps) - Executive Director

Hiring Organisation
JP Morgan Chase
Location
London, United Kingdom
Salary
£ 120 K
Operations and Engineers across the whole firmInnovates, designs and delivers technical solutions that can be leveraged across multiple businesses and domains, this will include AIOps strategy for the futureSets direction and governance for agentic AI-enabled engineering and SDLC/TLM automation within a technical area to drive measurable improvements … infrastructure or other cloud design and implementationPreferred qualifications, capabilities, and skillsUsage of large scale Observability Platforms (example products IBM Netcool, (Watson CloudPak AiOps), Tivoli, SCOM, SMARTS, Dynatrace, Splunk, Elastic, Prometheus, Grafana or Messaging systems for telemetry transport like Kafka, OTEL) Formal training or certification on software engineering concepts and expert ...