476 to 500 of 2,341 Observability Jobs in London

AI Delivery Tech Lead, Song Service

Location
Greater London, England, United Kingdom
services Lead engineering teams across complex, multi‐system agentic programmes Set technical standards, quality frameworks and architecture governance Ensure systems are designed for scalability, observability and compliance Collaborate closely with Forward Deployed Engineers, Experience Designers and Strategy leads Grow people, capability and the practice Mentor engineers and grow technical capability ...

AI Delivery Tech Lead (CL8)

Location
Greater London, England, United Kingdom
engineering teams across complex, multi-system agentic programmes Contribute to technical standards, quality frameworks and architecture governance Ensure systems are designed for scalability, observability and compliance Collaborate closely with Forward Deployed Engineers, Experience Designers and Strategy leads Grow people, capability and the practice Support more junior engineers and develop technical ...

Lead Data Engineer

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
practices and follows secure-by-design principle. Experience of working with agile frameworks like Scrum, Kanban. Background working in financial industryExperience of modern observability platforms e.g. DataDogAbout the London Stock Exchange GroupLSEG's vision is to be the most trusted expert in global financial markets. This is achieved through leading ...

Sr. Software Engineer, Inference

Location
Greater London, England, United Kingdom
alongside deep familiarity with networked systems and performance optimisation. Hands-on experience with Kubernetes at production scale, including automated CI/CD and modern observability stacks (e.g., Prometheus, Grafana, OpenTelemetry). Practical, working knowledge of inference internals: batching strategies, caching, mixed precision (BF16/FP8), and streaming token delivery. Proven ...

Presales Architect - CSP & MSP

Location
Greater London, England, United Kingdom
meeting sustainability goals. We do this from Day 0 provisioning of application infrastructure supporting both green, and brown, field management through Day 2+ observability of hybrid cloud environments. We provide world discovery and monitoring, event and incident management, remediation, and automation, powered by AI. We help our enterprise ...

SDN Presales Architect – CSP & MSP – UK

Location
Greater London, England, United Kingdom
meeting sustainability goals. We do this from Day 0 provisioning of application infrastructure supporting both green, and brown, field management through Day 2+ observability of hybrid cloud environments. We provide world discovery and monitoring, event and incident management, remediation, and automation, powered by AI. We help our enterprise ...

DBA Lead

Location
City Of London, England, United Kingdom
availability, performance, security, and operational stability. The team manages MySQL, SQL Server, PostgreSQL, Oracle, and cloud database technologies across hybrid environments while driving automation, observability, disaster recovery readiness, and continuous improvement initiatives. Working closely with engineering, infrastructure, security, and DevOps teams, DBAs play a key role in delivering resilient …/or Terragrunt CI/CD pipeline implementation using GitHub Actions Git-centric development workflows Automation experience using: Python, Bash/Shell Scripting, Ansible Observability & Reliability Engineering Experience with enterprise monitoring and observability platforms including Grafana/Datadog/ELK Azure Monitor CloudWatch Understanding of SRE concepts including: Service Level ...

Cloud SRE: Automation, Kubernetes & Observability

Location
Greater London, England, United Kingdom
LexisNexis Risk Solutions seeks a Site Reliability Engineer to design, build and operate cloud infrastructure across AWS and Azure. You will enable engineering teams through infrastructure as code, automation, and platform reliability practices. This hands ...

Senior Linux DevOps Engineer

Hiring Organisation
RedTech Recruitment Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
hands-on, Linux-focused DevOps role where you will take ownership of large-scale production environments, working across Linux systems, automation, Kubernetes, containerisation, networking, observability and cloud infrastructure. Location: London, hybrid working with a minimum of 2 days per week in the office Salary: Up to £100,000 per annum … networking, security and performance troubleshooting within Linux environments Hands-on experience operating containerised workloads using Docker and Kubernetes Experience with monitoring, logging and observability technologies such as Prometheus, Grafana, Loki, OpenTelemetry or the ELK Stack Experience with Infrastructure as Code and automation tooling such as Terraform and Ansible Experience building ...

SRE Engineer – FinTech Reliability, Observability & Cloud

Location
Greater London, England, United Kingdom
Hamilton Barnes Associates Limited is seeking a Site Reliability Engineer to work at the intersection of software engineering and infrastructure. You'll develop internal platforms, tooling, and automation across Linux, distributed systems, and cloud-native ...

Senior Site Reliability Engineer - Cloud & Observability

Location
Greater London, England, United Kingdom
A leading global financial markets company is seeking a Senior Engineer in Site Reliability. This role involves maintaining service level objectives, enhancing system reliability, and automating to ensure scalability. With a strong focus on cloud ...

Senior DevOps Platform Engineer

Hiring Organisation
LEAP29
Location
London, UK
Employment Type
Full-time
platform reliability, scalability and security through automation and engineering best practicesSupport cloud migrations from traditional data centre environments into modern cloud-native platformsImplement monitoring, observability and logging solutions across complex distributed systemsWork closely with development, security and architecture teams to improve software delivery processesRequired ExperienceStrong commercial experience working … financial services, government or healthcareExperience with GitOps practices and tools such as ArgoCDKnowledge of service mesh technologies including Istio or similarExperience with monitoring and observability platforms such as Prometheus, Grafana, Dynatrace, New Relic, Splunk or ELKExperience supporting hybrid cloud environments across AWS, Azure, GCP or private cloudExperience with security tooling ...

Senior Machine Learning Engineer (MLOps)

Hiring Organisation
ASOS
Location
London, UK
Employment Type
Full-time
search relevance, personalisation and emerging AI applications. This is a highly engineering-focused role with an emphasis on cloud-native systems, platform architecture, automation, observability and operational excellence. What You'll Be Doing: Design and build scalable machine learning platforms and infrastructure supporting model training, deployment and serving. Develop highly … data products. Design batch and real-time inference architectures using modern cloud-native technologies. Improve reliability, resilience and performance across ML workloads through monitoring, observability and automation. Build tooling and frameworks that enable data scientists and ML engineers to deploy models safely and efficiently. Own production services, infrastructure and operational ...

Java Backend Developer

Location
City Of London, England, United Kingdom
Investigate complex production incidents and perform root cause analysis. Implement permanent fixes for recurring production issues. Drive improvements in application stability, performance, monitoring, and observability . Participate in solution design, technical discussions, and code reviews. Develop unit and integration tests and ensure appropriate test coverage. Support application releases, deployments … Kubernetes and containerized applications . Experience with Microservices and distributed application architecture. Experience with Docker and cloud-native application development. Familiarity with monitoring and observability tools such as Splunk, Dynatrace, Datadog, Prometheus, or Grafana. Experience with CI/CD tools such as Jenkins, GitLab CI, GitHub Actions, or Azure DevOps. ...

Senior Technical Consultant - DevOps

Hiring Organisation
CloudBees
Location
London, UK
Employment Type
Full-time
Proven experience advising customers on DevOps, software delivery modernization, platform engineering, or cloud transformation. Strong grasp of CI/CD, Infrastructure as Code, GitOps, observability, security, and cloud-native architecture, with hands-on experience designing and implementing these at enterprise scale. Hands-on experience with a major public cloud … architectures. Experience with platform strategy, DevEx, or internal developer platforms (IDPs).Familiarity with AI-enabled development, agentic workflows, LLMs, or AI governance. Experience with observability/telemetry platforms (OpenTelemetry, Splunk, Dynatrace, Datadog, AppDynamics, Grafana).Experience in regulated or large-scale enterprise environments. Thought leadership via publications, conference speaking, or community ...

Senior Software Developer (Python)

Location
Greater London, England, United Kingdom
product development, including ingestion, transformation, model deployment, and visualisation workflows. Ensure software solutions meet security, reliability, scalability, and performance requirements. Support operational excellence through observability, monitoring, troubleshooting, and incident resolution. Contribute to technical design reviews, architecture discussions, and engineering best practices. Mentor junior developers and contribute to fostering a strong … Pandas, Apache Beam, Spark, or similar frameworks. Strong SQL skills and experience with analytical databases such as BigQuery, PostgreSQL, or ClickHouse. Familiarity with observability and monitoring tools such as Prometheus, Grafana, Splunk, or OpenTelemetry. Experience deploying and supporting large-scale distributed systems in Kubernetes environments. Knowledge of graph analytics, network ...

Cloud and Platform Engineer-Consultant-AI and Digital Factory

Location
Greater London, England, United Kingdom
team’s needs• Be responsible for infrastructure and platform provisioning, configuration, build, deployment, and monitoring using Infrastructe-as-Code• Define and implement what good observability practices look like for clients, track reliability metrics (SLIs/SLOs/error budgets) and participate in incident response where relevant to your specialism• Build … agile team environmentsSite Reliability Engineering• Defining and operating against SLIs/SLOs/error budgets• Incident response, on-call practices, and post-incident review• Observability and monitoring: Grafana, Dynatrace, CloudWatch, Prometheus, Datadog, OpenTelemetryGeneral• Scripting (bash/shell)• Testing tooling: Selenium, Cucumber, etc.• Ability to flexibly support a variety of technologies ...

Senior Java Developer

Hiring Organisation
Intelligent Resourcing Solutions Ltd
Location
London, UK
Employment Type
Full-time
Job Description Job Description We are seeking an experienced Senior Java Developer to join our development team in a 100% hands-on software development role. The successful candidate will be responsible for designing, developing, enhancing ...

Senior DevOps Engineer - Cloud Automation & AI Tools

Location
Greater London, England, United Kingdom
operate RX Cloud Platforms across AWS and Azure from Richmond, London. You will partner with software engineers to improve cloud infrastructure, deployment automation, observability and platform reliability. The role emphasizes automation-first priorities, AI-assisted tooling, and collaboration with agile teams to drive incident resolution and continuous improvement. #J ...

Platform Engineer: Kubernetes, Cloud & Automation

Location
Greater London, England, United Kingdom
harden infrastructure, apply IaC with Terraform/Ansible/Helm, and ensure secure, scalable operations across GCP, AWS and Azure, while contributing to observability and performance tuning. #J-18808-Ljbffr ...

Cloud DevOps Engineer — Flexible Hours & AI Automation

Location
Greater London, England, United Kingdom
will build, automate, secure, and support RX Cloud Platforms on AWS and Azure, collaborating with software and cloud teams to boost deployment automation, observability, and platform reliability. You will implement IaC with Terraform, contribute to CI/CD pipelines using GitHub Actions, and apply security best practices while improving monitoring ...

Hybrid London: DevOps & SRE Engineer, Cloud Reliability

Location
Greater London, England, United Kingdom
London office, a hybrid role based in London with at least 3 days on site. You will help shape reliability, scalability and observability of our cloud‐native platform, collaborating across the engineering org to build a high‐performance, well‐monitored system. You will drive incident response, CI/CD improvements ...

Senior Azure Data & AI Platform Engineer

Location
City Of London, England, United Kingdom
teams. The candidate will work hands-on with Terraform, Bicep, Azure DevOps or GitHub Actions, and ensure best practices for security, cost control, and observability across data platforms #J-18808-Ljbffr ...

Secure Cloud Platform Engineer | DevSecOps & SRE

Location
Greater London, England, United Kingdom
national security. Join agile, multi-disciplinary teams focused on CI/CD, Infrastructure as Code and live service support, with emphasis on security, observability and incident readiness to ensure #J-18808-Ljbffr ...

Infrastructure Lead

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
other teams, with documentation and examples. In addition, you should have experience working with Kubernetes in production environments at scale, and be familiar with observability tools such as Prometheus and Grafana. Strong Linux Server Administration and Configuration Management skills, as well as some networking experience, are also required. The ideal ...