701 to 725 of 2,062 Observability Jobs in the UK

Staff Platform Engineer, AI Enablement

Hiring Organisation
Jobleads-UK
Location
United Kingdom
years, not quarters. The Team Our Foundation team owns the platform every engineer at Gigs builds on: cloud infrastructure, CI/CD, observability, golden paths, and the shared services that keep our product running. That charter is expanding. We believe the next big improvement in how Gigs ships software comes ...

Founding Backend Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
messy, heterogeneous external data, at national scale. In your first 90 days you'd ship multi-tenant authentication, a hardened ingestion layer with the observability we can put an SLA behind, and the benchmarking harness for those accuracy claims. Real production milestones, not onboarding theatre. On AI tooling and ownership ...

Senior Software Engineer - London

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
e.g., E2E/Cypress). Backend Excellence: Engineers sophisticated backend solutions involving API versioning, caching strategies, and complex data migration plans. Operational Maturity: Leads observability and SRE practices; defines SLOs, manages incident responses, and conducts blameless post-mortems. Security & Risk: Oversees operational security, including secrets hygiene and dependency risk management ...

Head of AI Solutions, COO Technology - MD (C16)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
agentic safety at enterprise scale. AI/ML Engineering & MLOps: Full AI/ML lifecycle ownership: training pipelines, model deployment, versioning, monitoring, drift detection, observability (e.g., Weights & Biases, Arize), and lifecycle management using platforms such as MLflow, Vertex AI, or SageMaker. Cloud AI Platforms: Demonstrated deployment of AI workloads ...

Head of Production Management - JP Morgan Personal Investing

Hiring Organisation
17918
Location
London, United Kingdom
ability to influence across technical and non-technical audiences in complex, regulated environments. Strong technical background in modern production environments including cloud-native infrastructure, observability, CI/CD pipelines, and automation frameworks. Proven track record defining and governing production management standards - change, incident, capacity, and automation - across multiple engineering teams ...

Head of Production Management - JP Morgan Personal Investing

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
ability to influence across technical and non-technical audiences in complex, regulated environments. Strong technical background in modern production environments including cloud-native infrastructure, observability, CI/CD pipelines, and automation frameworks. Proven track record defining and governing production management standards - change, incident, capacity, and automation - across multiple engineering teams ...

Head of AI Solutions, COO Technology - MD (C16)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
agentic safety at enterprise scale* AI/ML Engineering & MLOps: Full AI/ML lifecycle ownership: training pipelines, model deployment, versioning, monitoring, drift detection, observability (e.g., Weights & Biases, Arize), and lifecycle management using platforms such as MLflow, Vertex AI, or SageMaker* Cloud AI Platforms: Demonstrated deployment of AI workloads ...

AI Solution Architect Senior Manager

Hiring Organisation
Jobleads-UK
Location
United Kingdom
leveraging Genesys Cloud AI capabilities, conversational AI platforms, generative AI technologies, and third-party AI ecosystems.* Establish architectural standards for AI solution scalability, security, observability, reliability, compliance, and responsible AI practices.* Lead architectural reviews and provide guidance on AI deployment models, integration patterns, data architecture, and operationalization strategies.* Serve ...

Senior Software Engineer - London

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
e.g., E2E/Cypress). Backend Excellence: Engineers sophisticated backend solutions involving API versioning, caching strategies, and complex data migration plans. Operational Maturity: Leads observability and SRE practices; defines SLOs, manages incident responses, and conducts blameless post-mortems. Security \& Risk: Oversees operational security, including secrets hygiene and dependency risk management ...

Senior Software Engineer - London

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
e.g., E2E/Cypress). Backend Excellence: Engineers sophisticated backend solutions involving API versioning, caching strategies, and complex data migration plans. Operational Maturity: Leads observability and SRE practices; defines SLOs, manages incident responses, and conducts blameless post-mortems. Security & Risk: Oversees operational security, including secrets hygiene and dependency risk management ...

Director of Client Engineering

Hiring Organisation
Tide Platform
Location
Moffat, Dumfries & Galloway, UK
Employment Type
Full-time
evaluate our current stack and workforce model to ensure we are maximising engineering leverage without sacrificing the right tool for the job. Product Virtualisation & Observability: Advocate for systems that provide stakeholders globally with full visibility into the product experience across all configurations. Delivery Excellence: Be responsible for the delivery ...

AI Harness Engineers

Hiring Organisation
Capgemini
Location
Greater London, United Kingdom
Employment Type
Full Time
Enterprise as the semantic knowledge graph, an agent memory plane serving episodic and precedent memory over MCP, MCP-native connectors, OpenTelemetry and Grafana for observability, all on CNCF-conformant Kubernetes with Helm and Argo CD, deployable to any hyperscaler or on-prem. A tool-for-tool match is not expected ...

Senior Engineering Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
market, and eligibility — closer to the customer than most engineering roles. Craft that's respected, not traded off . Strict typing, layered testing, ADRs, observability — and AI-native ways of working as the baseline, not an experiment. A company mid-transformation, told straight . Multiverse is rebuilding itself into ...

Site Reliability Engineer

Hiring Organisation
SR2 | Socially Responsible Recruitment | Certified B CorporationTM
Location
London, UK
Employment Type
Full-time
Description Site Reliability Engineer (SRE) DevSecOps | Cloud Engineering | Observability | Production Environments | London SR2 is supporting a major 3-year programme and looking for an experienced Site Reliability Engineer (SRE) to join the Production Engineering team. This function underpins the reliability, security, and performance of all live environments, from production systems … native infrastructure. Beyond supporting live systems, this team also acts as a centre of excellence, guiding project teams in adopting best practices across DevSecOps, observability, and cost optimisation. Key Responsibilities: Build, maintain, and support production and demo environments Automate infrastructure provisioning and deployment workflows (Terraform, GitHub Actions, GitOps) Package ...

Director of DevOps & SRE

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
high-severity incidents and major client environment changes, ensuring appropriate change management and CAB governance, with occasional off-hours support for the teams. Observability and service levels. SLO/SLA and monitoring strategy across Grafana and our wider observability tooling, building alert coverage that is meaningful rather than noisy. … shared-responsibility operating model between DevOps/SRE and product engineering. Vendor and cost management. Relationships across cloud, security, CI/CD, and observability platforms, including spend and renewal negotiation. You Have: 8+ years in DevOps, SRE, or infrastructure engineering, including 3+ years leading and developing engineering teams. Deep hands ...

Remote UK SRE: Observability & Reliability Lead

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Orex Nova, Inc. is seeking an experienced SRE/Platform Engineer to help keep our systems fast, reliable, and observable. This fully remote role covers the UK and requires strong incident response experience. You will ...

SRE Architect (68019) (DEAI DS) Cloud & Data Engineering United Kingdom

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us. Job Description Mandatory Skills: Observability, Resiliency, Service Management, Reliability, Performance engineering, Scalability, release management, Cloud cost management. Role Description Skills: ROLE PURPOSE Lead the Site Reliability Engineering practice, driving … reactive operations to proactive, engineering-led reliability. Own the definition and enforcement of non-functional requirements (NFRs) using FMEA-based resiliency frameworks, and champion observability, self-healing automation, automated incident management, and database operations automation. Ensure systems are resilient, performant, cost-optimised, and continuously improving. KEY RESPONSIBILITIES Define and enforce ...

Principal Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
governance, automation, security controls, and developer experience established with our GCP platform. Develop and enhance platform services, self‐service capabilities, CI/CD tooling, observability solutions, and automation frameworks that improve developer productivity and operational effectiveness. Contribute to our AI transformation by building the platform capabilities, tooling, operational patterns … teams to define technical direction, establish engineering standards, and ensure platform capabilities meet current and future business requirements. Drive continuous improvement across platform reliability, observability, automation, security, cost efficiency, and developer experience through data‐driven engineering practices. Mentor and support engineers across the organisation, promoting platform engineering best practices, technical ...

Software Engineer/ SRE (Linux)

Hiring Organisation
Visa
Location
Basingstoke, Hampshire, UK
Employment Type
Full-time
platform strategy. In this role, you'll ensure our development platform and tools let engineers focus on innovation instead of infrastructure. You'll promote observability best practices and automate resolution of recurring issues, working closely with software engineering teams to support security, availability, and performance. Responsibilities include triaging issues, collaborating … implement, and maintain systems for high availability, scalability, and performance. Monitor and improve application reliability through proactive measures and incident response. Develop and maintain observability solutions (metrics, logging, tracing). Participate in on-call rotations and drive root cause analysis for incidents. Collaboration & Continuous Improvement Partner with engineering teams ...

Cloud Platform Architect (68020)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
important to us. Job Description Mandatory Skills Plat Engineering for on‐premises & cloud, IaC, CI/CD pipeline & tools, security, network, landing zones, observability, self‐servicing, DB automation Role Purpose Lead the design, build, and operational excellence of the organisation’s cloud platform across private and public cloud environments. … coach 2 Cloud Platform Engineers, establish coding standards, conduct architecture reviews, and drive knowledge sharing Collaborate with SRE and Security teams to embed reliability, observability, and compliance into the platform from the ground up Produce architectural decision records (ADRs), runbooks, and platform documentation for operational excellence Technical Skills & Expertise Expert ...

Cloud Platform Architect (68020)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
important to us. Job Description Mandatory Skills: Plat Engineering for on‐premises & cloud, IaC, CI/CD pipeline & tools, security, network, landing zones, observability, self‐servicing, DB automation ROLE PURPOSE Lead the design, build, and operational excellence of the organisation’s cloud platform across private and public cloud environments. … coach 2 Cloud Platform Engineers, establish coding standards, conduct architecture reviews, and drive knowledge sharing Collaborate with SRE and Security teams to embed reliability, observability, and compliance into the platform from the ground up Produce architectural decision records (ADRs), runbooks, and platform documentation for operational excellence TECHNICAL SKILLS & EXPERTISE Expert ...

Site Reliability Engineer - Service Assurance Systems

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
range of systems across the Viasat estate. This data is shared between internal platforms and distributed through Google Cloud Platform (GCP), where it underpins observability and monitoring capabilities that give operations teams real-time insight into service health and performance. To ensure the data is accurate and fit for purpose … SLAs. Manage and oversee application deployments across development, staging and production environments, ensuring smooth and reliable release processes. Monitor application and infrastructure health using observability tools such as Prometheus and AWS CloudWatch; proactively identify and respond to anomalies and performance degradation. Own and maintain CI/CD pipelines, working ...

Vice President, Site Reliability Engineering

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
implement, and continuously improve Service Level Indicators, Service Level Objectives, and service health measures aligned to operational and business priorities. Build and optimize monitoring, observability, and alerting capabilities using tools such as Prometheus, Grafana, AppDynamics, and Splunk. Apply AIOps capabilities to improve event correlation, anomaly detection, root cause analysis, predictive … enterprise or production environments. Demonstrated ability to define and operationalize SLIs, SLOs, dashboards, alerts, and health indicators. Hands-on experience with enterprise monitoring and observability platforms including Prometheus, Grafana, AppDynamics, and Splunk. Strong troubleshooting, analytical, and problem-solving skills in complex distributed or production environments. Strong verbal and written communication ...

Vice President, Site Reliability Engineering

Hiring Organisation
17918
Location
Westminster, West End, United Kingdom
implement, and continuously improve Service Level Indicators, Service Level Objectives, and service health measures aligned to operational and business priorities. Build and optimize monitoring, observability, and alerting capabilities using tools such as Prometheus, Grafana, AppDynamics, and Splunk. Apply AIOps capabilities to improve event correlation, anomaly detection, root cause analysis, predictive … enterprise or production environments. Demonstrated ability to define and operationalize SLIs, SLOs, dashboards, alerts, and health indicators. Hands-on experience with enterprise monitoring and observability platforms including Prometheus, Grafana, AppDynamics, and Splunk. Strong troubleshooting, analytical, and problem-solving skills in complex distributed or production environments. Strong verbal and written communication ...

Software Engineer, General

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
services for platform control planes, resource management, and inference serving Design and implement database schemas for storing platform state, resource metadata, billing data, and observability metrics Work with distributed storage systems, message queues (Kafka, RabbitMQ), and databases (PostgreSQL, Redis) to build reliable platform components Build event‐driven architectures for asynchronous … workflows Exposure to cloud platforms (AWS/GCP/Azure) and their core services is a plus Experience with infrastructure‐as‐code (Terraform) or observability tools (Prometheus, Grafana) is beneficial Bonus/Good to Have HPC & Cluster Management: Experience handling large-scale HPC clusters using Kubernetes and Slurm ...