601 to 625 of 2,093 Observability Jobs in the UK

Senior Software Engineer - Pay Sustainable Engineering

Hiring Organisation
Hackajob Ltd
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent
e.g., E2E/Cypress). Backend Excellence: Engineers sophisticated backend solutions involving API versioning, caching strategies, and complex data migration plans. Operational Maturity: Leads observability and SRE practices; defines SLOs, manages incident responses, and conducts blameless post-mortems. Security & Risk: Oversees operational security, including secrets hygiene and dependency risk management ...

Senior Applied AI Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
evaluation datasets and automated eval pipelines that give the team confidence in AI feature quality before and after changes. Instrument AI features for production observability: logging, drift detection, quality monitoring, and alerting. Collaborate with Product and Design to scope AI features from first principles; you are a co-author ...

Senior Engineering Manager

Hiring Organisation
Multiverse Group
Location
London, UK
Employment Type
Full-time
market, and eligibility — closer to the customer than most engineering roles.Craft that's respected, not traded off. Strict typing, layered testing, ADRs, observability — and AI-native ways of working as the baseline, not an experiment.A company mid-transformation, told straight. Multiverse is rebuilding itself into a product-powered, AI-native ...

Lead Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
United Kingdom
enforce SLOs, SLAs, and error budgets Develop and configure monitoring dashboards and alerts in tools like Grafana and Azure Monitor. Installation and configuration of Observability Platform including tools like Grafana, Prometheus, Azure Monitor, Open telemetry etc. Developing bicep modules for monitoring infrastructure and deploy it. Optimize system performance, cost … reproducing issues in a local environment. Multi-tasking and time-management to prioritise and switch between varied tasks. Significant experience in platform engineering, observability, and provisioning. Proven ability to develop and implement a strategic vision for platform services, observability, and provisioning. Strong understanding of cyber security principles, governance, and compliance ...

Senior Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Knutsford, England, United Kingdom
drive reliability, scalability and performance across critical banking systems. This role combines hands‐on SRE engineering with technical leadership, with a strong focus on observability, automation, continuous improvement and optimisation. Responsibilities: * Build and maintain reliable, scalable and secure infrastructure platforms and solutions. * Apply SRE and software engineering practices to improve … lead complex troubleshooting and root cause analysis. * Develop automation using programming and scripting to reduce manual intervention and improve efficiency. * Develop and improve observability, monitoring, instrumentation and performance capabilities. * Use data and reliability metrics to drive continuous improvement and optimisation. * Lead technical discussions, blameless retrospectives and problem‐solving activities. * Work ...

Senior Site Reliability Engineer

Hiring Organisation
GCS
Location
Glasgow, City of Glasgow, United Kingdom
Employment Type
Permanent
Salary
£75000 - £95000/annum Bonus
drive reliability, scalability and performance across critical banking systems. This role combines hands-on SRE engineering with technical leadership, with a strong focus on observability, automation, continuous improvement and optimisation. Responsibilities: * Build and maintain reliable, scalable and secure infrastructure platforms and solutions. * Apply SRE and software engineering practices to improve … lead complex troubleshooting and root cause analysis. * Develop automation using programming and scripting to reduce manual intervention and improve efficiency. * Develop and improve observability, monitoring, instrumentation and performance capabilities. * Use data and reliability metrics to drive continuous improvement and optimisation. * Lead technical discussions, blameless retrospectives and problem-solving activities. * Work ...

Senior DevOps Engineer - CI/CD, SRE & Observability

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Jobtailor est à la recherche d'un expert DevOps/SRE à Londres pour concevoir et optimiser des pipelines CI/CD, automatiser le provisioning et renforcer l observabilité dans des environnements distribués à haute ...

Remote SRE: Platform Reliability & Observability

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Orexnova is seeking an experienced SRE/Platform Engineer to join a fully remote UK team. You’ll own incident response, blameless post-mortems and drive reliability improvements across services while partnering with the SRE ...

Principal Cloud SRE: Multi-Cloud Reliability & Observability

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Veson Nautical is hiring a Principal Site Reliability Engineer in London, hybrid, to design and operate scalable cloud infrastructure across GCP and AWS. You will lead greenfield builds, drive automation with Terraform, and mentor peers ...

Senior SRE: Cloud, Automation & Observability

Hiring Organisation
Jobleads-UK
Location
Newbury, England, United Kingdom
Vodafone Group Plc in Newbury is seeking an experienced Site Reliability Engineer to support our cloud platforms and developer experience. You will own design and operation of scalable AWS-based services, collaborating with cross-functional ...

Lead DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
partnership with engineering, infrastructure, security, and operations teams.It is an opportunity to improve reliability, scalability, security, and delivery while advancing DevOps, platform engineering, observability, and AI-enabled infrastructure tooling.**Responsibilities*** Lead the design and evolution of scalable, secure, and highly available trading infrastructure.* Provide technical direction, mentorship, and guidance … assisted engineering tools to accelerate development, automation, troubleshooting, documentation, and operational workflows.* Identify, design, and help implement AI-enabled operational capabilities for infrastructure automation, observability, incident response, and platform engineering.* Contribute to the strategy for safe, practical, and secure adoption of AI across infrastructure and DevOps practices.* Lead and participate ...

Lead DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
partnership with engineering, infrastructure, security, and operations teams. It is an opportunity to improve reliability, scalability, security, and delivery while advancing DevOps, platform engineering, observability, and AI-enabled infrastructure tooling. Responsibilities Lead the design and evolution of scalable, secure, and highly available trading infrastructure. Provide technical direction, mentorship, and guidance … assisted engineering tools to accelerate development, automation, troubleshooting, documentation, and operational workflows. Identify, design, and help implement AI-enabled operational capabilities for infrastructure automation, observability, incident response, and platform engineering. Contribute to the strategy for safe, practical, and secure adoption of AI across infrastructure and DevOps practices. Lead and participate ...

Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
data and AI workflows. It’s an excellent opportunity for an experienced Platform/DevOps Engineer to work with cloud, Kubernetes, CI/CD, observability, and emerging AI infrastructure while helping establish scalable, secure, and reliable engineering practices. This is an opportunity to join an innovative, progressive, and collaborative team. … agent orchestration AI Evaluation & Quality: Eval harnesses and golden datasets, LLM-as-judge and human-in-the-loop review, regression suites, and red-teaming Observability & Monitoring: Prometheus, Grafana, Datadog, Splunk, Elastic/ELK, OpenTelemetry, including GenAI tracing and token, latency, and cost telemetry Platform Security & Policy-as-Code: HashiCorp Vault ...

Senior Cloud Engineer, AI Platform SRE

Hiring Organisation
Jobleads-UK
Location
Leeds, England, United Kingdom
tools and prompt changes. It means the SLOs and on-call practice that make reliability an asset commitment rather than a hope, and the observability that makes AI-specific failure modes visible, including drift, silent quality regression, cost blowouts and agent loops. You'll be the operational conscience … Take part in the programme's on-call rotation, and build runbooks clear enough that someone else can use them at 3am. AI-specific observability: Instrument latency, token usage, cost per request, model and agent error rates, retrieval quality and drift, with dashboards and alerting that surface problems before users ...

Senior Cloud Engineer, AI Platform SRE

Hiring Organisation
Jobleads-UK
Location
Manchester, England, United Kingdom
tools and prompt changes. It means the SLOs and on-call practice that make reliability an asset commitment rather than a hope, and the observability that makes AI-specific failure modes visible, including drift, silent quality regression, cost blowouts and agent loops. You'll be the operational conscience … Take part in the programme's on-call rotation, and build runbooks clear enough that someone else can use them at 3am. AI-specific observability: Instrument latency, token usage, cost per request, model and agent error rates, retrieval quality and drift, with dashboards and alerting that surface problems before users ...

Senior Cloud Engineer, AI Platform SRE

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
tools and prompt changes. It means the SLOs and on-call practice that make reliability an asset commitment rather than a hope, and the observability that makes AI-specific failure modes visible, including drift, silent quality regression, cost blowouts and agent loops. You'll be the operational conscience … Take part in the programme's on-call rotation, and build runbooks clear enough that someone else can use them at 3am. AI-specific observability: Instrument latency, token usage, cost per request, model and agent error rates, retrieval quality and drift, with dashboards and alerting that surface problems before users ...

Senior Cloud Engineer, AI Platform SRE

Hiring Organisation
Jobleads-UK
Location
City of Edinburgh, Scotland, United Kingdom
tools and prompt changes. It means the SLOs and on-call practice that make reliability an asset commitment rather than a hope, and the observability that makes AI-specific failure modes visible, including drift, silent quality regression, cost blowouts and agent loops. You'll be the operational conscience … Take part in the programme's on-call rotation, and build runbooks clear enough that someone else can use them at 3am. AI-specific observability: Instrument latency, token usage, cost per request, model and agent error rates, retrieval quality and drift, with dashboards and alerting that surface problems before users ...

Lead DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
cloud infrastructure, delivery platforms, and operational capabilities. You will remain hands-on across the engineering lifecycle, from architecture and infrastructure design through deployment, observability, incident response, and continuous improvement.We expect you to operate with a high degree of autonomy, make strategic and architectural decisions within your area, and resolve complex … teams to productionize AI solutions and ensure services are ready to operate reliably at scale.* Establish engineering standards and reusable patterns for infrastructure, security, observability, resilience, documentation, and operational readiness.* Lead architectural decisions and evaluate trade-offs across reliability, security, scalability, performance, cost, and maintainability.* Take ownership of operational risks ...

Lead Cloud Platform Engineer (Kubernetes) - Remote

Hiring Organisation
Jobleads-UK
Location
United Kingdom
managed platform services, including capacity planning, performance tuning, cost optimisation, patching, and lifecycle management Partner with software engineering teams to support application deployment, troubleshooting, observability, and platform adoption Monitor platform health and respond to incidents, conducting root cause analysis and implementing preventative improvements Identify technical debt and contribute to platform …/CD principles and experience building and maintaining automated delivery pipelines Experience working with container technologies and cloud‐native architectures Knowledge of observability, monitoring, logging, and incident management practices Strong troubleshooting and problem‐solving skills across infrastructure, platform, and application layers Experience supporting software development teams in deploying and operating ...

Senior Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
Engineer containerized environments with advanced orchestration, networking, and security. Deliver internal platforms and self-service capabilities that improve developer experience and reduce friction. Implement observability stacks—metrics, logs, traces, and proactive alerting for reliability. Champion security and compliance across infrastructure and delivery pipelines. Architect secure, scalable networking solutions for hybrid … management (e.g., Ansible, Chef). Strong AWS architecture skills and cost optimisation strategies Advanced containerization and orchestration experience (Docker, Kubernetes, etc.). Proficiency in observability tools (Prometheus, Grafana, ELK, OpenTelemetry). Security-first mindset with hands‐on experience in access control, encryption, and incident response. Solid networking knowledge - protocols, routing ...

Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
deploying Large Language Model (LLM) orchestration frameworks (e.g., LangChain, Temporal, or custom agentic loops) to coordinate multi‐step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root‐cause analysis, anomaly detection, and semantic … plus. Preferred Qualifications: AWS certifications (Solutions Architect Professional, DevOps Engineer, etc.). Experience building self‐service platforms for development teams. Familiarity with observability and monitoring tools (Prometheus, Grafana, Datadog, CloudWatch). Background in financial services or other highly regulated environments. About the company At the company, we commit our people ...

Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
deploying Large Language Model (LLM) orchestration frameworks (e.g., LangChain, Temporal, or custom agentic loops) to coordinate multi-step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root-cause analysis, anomaly detection, and semantic … technical decisions across teams. Experience working in regulated industries is a plus. Preferred Qualifications Experience building self-service platforms for development teams. Familiarity with observability and monitoring tools (Prometheus, Grafana, Datadog, CloudWatch). Background in financial services or other highly regulated environments. About Goldman Sachs At Goldman Sachs, we commit ...

Senior SRE: GCP & Kubernetes, Automation Lead

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
degree of autonomy and ownership. You will design, implement, and operate cloud-native infrastructure on GCP using Kubernetes, Terraform, and Helm, while championing automation, observability, and best practices across the SDLC. #J-18808-Ljbffr ...

Senior DevOps Engineer - Cloud, Kubernetes & AI Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
work to improve developer productivity and platform resilience. You will design cloud and containerized infrastructure, implement CI/CD pipelines, and advance IaC, observability, and deployment tooling in close collaboration with engineers and product teams. #J-18808-Ljbffr ...

Remote Platform Engineer: AI, Cloud & CI/CD Automation

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
design and operate pipelines, IaC, and cloud services across AWS/Azure and Kubernetes, aligning with GitOps and security best practices. You will implement observability, self‐service tooling, and automation while collaborating with developers to ship reliable software and AI workloads at scale. #J-18808-Ljbffr ...