26 to 50 of 3,940 Observability Jobs

Senior Specialist Engineer (Specialist Site Reliability Engineer SRE)

Hiring Organisation
National Health Service
Location
London, United Kingdom
Salary
£ 70 K
scalable, and perform optimally in production environments. The role will monitor and manage these aspects while taking responsibility for multiple cloud infrastructure services. Observability of systems will be key to prioritising the operational service improvements and performance improvements to meet and exceed SLOs (Service Level Objectives).Main duties … solving skills to identify bottlenecks with an engineering mindsetEnsure systems can handle current and future workloads through automation and capacity planningContinuously improve services through observability, and identify ways to improve observability practicesFollow SRE principles. Guide and educate stakeholders to adopt implemented principlesProvide technical documentation for engineers. Providing training, where appropriateWorking ...

DBA Lead

Location
City Of London, England, United Kingdom
availability, performance, security, and operational stability. The team manages MySQL, SQL Server, PostgreSQL, Oracle, and cloud database technologies across hybrid environments while driving automation, observability, disaster recovery readiness, and continuous improvement initiatives. Working closely with engineering, infrastructure, security, and DevOps teams, DBAs play a key role in delivering resilient …/or Terragrunt CI/CD pipeline implementation using GitHub Actions Git-centric development workflows Automation experience using: Python, Bash/Shell Scripting, Ansible Observability & Reliability Engineering Experience with enterprise monitoring and observability platforms including Grafana/Datadog/ELK Azure Monitor CloudWatch Understanding of SRE concepts including: Service Level ...

Java Backend Developer

Location
City Of London, England, United Kingdom
Investigate complex production incidents and perform root cause analysis. Implement permanent fixes for recurring production issues. Drive improvements in application stability, performance, monitoring, and observability . Participate in solution design, technical discussions, and code reviews. Develop unit and integration tests and ensure appropriate test coverage. Support application releases, deployments … Kubernetes and containerized applications . Experience with Microservices and distributed application architecture. Experience with Docker and cloud-native application development. Familiarity with monitoring and observability tools such as Splunk, Dynatrace, Datadog, Prometheus, or Grafana. Experience with CI/CD tools such as Jenkins, GitLab CI, GitHub Actions, or Azure DevOps. ...

Vice President, Credit Technology Engineering

Hiring Organisation
Ares Management
Location
London, UK
Employment Type
Full-time
credit technology ecosystem. Define architecture standards for distributed systems, APIs, event-driven architectures, and microservices. Drive engineering decisions related to scalability, resiliency, performance, observability, and operational excellence. Evaluate new technologies and architectural patterns to improve engineering effectiveness and platform capabilities. Application Development & Technology Delivery Lead development across: Credit deal management … technologies. Implement containerized solutions utilizing Kubernetes and cloud platform services. Drive modernization efforts focused on cloud adoption, scalability, resilience, and operational efficiency. Ensure proper observability, monitoring, logging, and operational support across production environments. Integration & Event-Driven Architecture Design and implement API-first architectures and integration frameworks. Leverage messaging platforms, event ...

Senior Software Developer (Python)

Location
Greater London, England, United Kingdom
product development, including ingestion, transformation, model deployment, and visualisation workflows. Ensure software solutions meet security, reliability, scalability, and performance requirements. Support operational excellence through observability, monitoring, troubleshooting, and incident resolution. Contribute to technical design reviews, architecture discussions, and engineering best practices. Mentor junior developers and contribute to fostering a strong … Pandas, Apache Beam, Spark, or similar frameworks. Strong SQL skills and experience with analytical databases such as BigQuery, PostgreSQL, or ClickHouse. Familiarity with observability and monitoring tools such as Prometheus, Grafana, Splunk, or OpenTelemetry. Experience deploying and supporting large-scale distributed systems in Kubernetes environments. Knowledge of graph analytics, network ...

Senior DevOps Engineer

Location
City of Edinburgh, Scotland, United Kingdom
integration platforms. Drive incident management, root cause analysis, and post-incident reviews. Improve platform availability, scalability, and performance through automation and engineering improvements. Monitoring & Observability Implement enterprise monitoring, alerting, logging, and observability solutions. Grafana ELK/Elastic Stack OpenTelemetry Proactively identify issues before they impact customers and engineering teams. Develop ...

Cloud and Platform Engineer-Consultant-AI and Digital Factory

Location
Manchester, England, United Kingdom
team’s needs• Be responsible for infrastructure and platform provisioning, configuration, build, deployment, and monitoring using Infrastructe-as-Code• Define and implement what good observability practices look like for clients, track reliability metrics (SLIs/SLOs/error budgets) and participate in incident response where relevant to your specialism• Build … agile team environmentsSite Reliability Engineering• Defining and operating against SLIs/SLOs/error budgets• Incident response, on-call practices, and post-incident review• Observability and monitoring: Grafana, Dynatrace, CloudWatch, Prometheus, Datadog, OpenTelemetryGeneral• Scripting (bash/shell)• Testing tooling: Selenium, Cucumber, etc.• Ability to flexibly support a variety of technologies ...

Cloud and Platform Engineer-Consultant-AI and Digital Factory

Location
Greater London, England, United Kingdom
team’s needs• Be responsible for infrastructure and platform provisioning, configuration, build, deployment, and monitoring using Infrastructe-as-Code• Define and implement what good observability practices look like for clients, track reliability metrics (SLIs/SLOs/error budgets) and participate in incident response where relevant to your specialism• Build … agile team environmentsSite Reliability Engineering• Defining and operating against SLIs/SLOs/error budgets• Incident response, on-call practices, and post-incident review• Observability and monitoring: Grafana, Dynatrace, CloudWatch, Prometheus, Datadog, OpenTelemetryGeneral• Scripting (bash/shell)• Testing tooling: Selenium, Cucumber, etc.• Ability to flexibly support a variety of technologies ...

Cloud and Platform Engineer-Consultant-AI and Digital Factory

Location
Newcastle upon Tyne, England, United Kingdom
team’s needs• Be responsible for infrastructure and platform provisioning, configuration, build, deployment, and monitoring using Infrastructe-as-Code• Define and implement what good observability practices look like for clients, track reliability metrics (SLIs/SLOs/error budgets) and participate in incident response where relevant to your specialism• Build … agile team environmentsSite Reliability Engineering• Defining and operating against SLIs/SLOs/error budgets• Incident response, on-call practices, and post-incident review• Observability and monitoring: Grafana, Dynatrace, CloudWatch, Prometheus, Datadog, OpenTelemetryGeneral• Scripting (bash/shell)• Testing tooling: Selenium, Cucumber, etc.• Ability to flexibly support a variety of technologies ...

DevOps Engineer

Hiring Organisation
Sanderson Recruitment
Location
London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
Up to £425 per day + Inside IR-35
legacy platforms and drive adoption of cloud-native technologies Manage and optimise production Kubernetes environments, ensuring scalability, performance and resilience Implement monitoring, alerting and observability solutions using Prometheus and Grafana Embed security best practices throughout the software delivery lifecycle Develop automation tooling using Ansible, Python and Bash Work closely with ...

Software Engineer III - SRE

Hiring Organisation
JP Morgan Chase
Location
Bournemouth, Dorset, United Kingdom
Salary
£ 70 K
e.g., Python, Java/Spring Boot, .NET).Experience with infrastructure as code tools (Terraform, CloudFormation, etc.).Strong knowledge of Linux/Unix systems.Experience with observability and monitoring tools (e.g., Grafana, Dynatrace, Prometheus, Datadog, Splunk).Familiarity with CI/CD tools (e.g., Jenkins, GitLab) and container orchestration (e.g., Docker, Kubernetes ...

Software Engineer III - SRE

Hiring Organisation
JP Morgan Chase
Location
Bournemouth, Dorset, UK
Employment Type
Full-time
Python, Java/Spring Boot, .NET).Experience with infrastructure as code tools (Terraform, CloudFormation, etc.).Strong knowledge of Linux/Unix systems. Experience with observability and monitoring tools (e.g., Grafana, Dynatrace, Prometheus, Datadog, Splunk).Familiarity with CI/CD tools (e.g., Jenkins, GitLab) and container orchestration (e.g., Docker, Kubernetes ...

Senior Lead Software Engineer - Python / Go

Location
Auchentibber, Scotland, United Kingdom
eliminate platform bottlenecks. Define and promote paved paths and self-service workflows for developers. Implement real-time telemetry pipelines and workflows for platform observability and analytics. Champion adoption of productivity tools through documentation, training, and engagement with the developer community. Standardize use of AI-assisted coding tools and AI-powered ...

Senior Lead Platform Engineer - Python

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
eliminate platform bottlenecks. Define and promote paved paths and self-service workflows for developers. Implement real-time telemetry pipelines and workflows for platform observability and analytics. Champion adoption of productivity tools through documentation, training, and engagement with the developer community. Standardize use of AI-assisted coding tools and AI-powered ...

Principal Cloud Engineer (Terraform), London

Location
Greater London, England, United Kingdom
Apply data quality and validation frameworks to ensure accuracy, completeness, and freshness of cost and usage data across all cloud providers; instrument pipelines with observability tooling to surface issues proactively. Build and maintain reusable data assets - curated datasets, aggregations, and data marts - that power FinOps dashboards, showback/chargeback reporting ...

Senior Software Engineer

Hiring Organisation
AECOM
Location
London, UK
Employment Type
Full-time
deployment pipelines such as GitHub Actions. Depth in distributed or event-driven systems, asynchronous processing, or data-intensive and high-throughput applications. Experience with observability and performance tuning, including tools such as Prometheus, Grafana, ELK or Datadog, or optimising CPU-, GPU- or distributed ML workloads. Experience in startups, scaleups ...

Senior Java Developer

Hiring Organisation
Intelligent Resourcing Solutions Ltd
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
GBP Annual
Job Description We are seeking an experienced Senior Java Developer to join our development team in a 100% hands-on software development role . The successful candidate will be responsible for designing, developing, enhancing, and ...

Monitoring & Observability Engineer (Dynatrace)

Location
Greater London, England, United Kingdom
Monitoring & Observability Engineer (Dynatrace) Location: UK - London, UK - Reading, UK - Hatfield, UK - Nottingham, UK - Manchester, UK - Milton Keynes, UK - Birmingham | Job-ID: 214264 | Contract type: Standard | Business Unit: IT Consulting Life on the team At Computacenter, you’ll be joining a world-class team of over 1,000 skilled professionals … across the UK, Germany, France, and India, delivering complex, enterprise-grade IT solutions and consultancy across infrastructure, cloud, and modern operations. As a Monitoring & Observability Engineer, you'll work in high-impact delivery teams that support some of the world’s most well-known organisations. You’ll play ...

Monitoring & Observability Engineer (Dynatrace)

Hiring Organisation
Computacenter
Location
London, UK
Employment Type
Full-time
across the UK, Germany, France, and India, delivering complex, enterprise-grade IT solutions and consultancy across infrastructure, cloud, and modern operations. As a Monitoring & Observability Engineer, you'll work in high-impact delivery teams that support some of the world's most well-known organisations. You'll play … reliability across their IT estates—contributing to their operational success through proactive insight and incident prevention. What you'll doDesign, implement, and manage observability solutions using industry-leading tools such as Dynatrace (primary), Grafana, and SplunkCollect and analyse telemetry data (metrics, logs, traces, events) to diagnose and resolve system ...

Observability Engineer - Assistant Vice President

Location
Greater London, England, United Kingdom
Site Reliability Engineer (SRE) - Assistant Vice President is a technical professional responsible for the hands‐on execution, technical implementation, and deployment of SRE and observability principles in a complex, critical, and large-scale multi-disciplinary environment. In this role, you will apply a deep understanding of multiple technology domains … technical contributor, you will drive the migration of applications from existing monitoring tools (Geneos ITRS, Prometheus, ELK, Splunk, AppDynamics, etc.) to Google Cloud Observability (GCO) and Grafana using OpenTelemetry (OTel) as the instrumentation standard. You will act as a hands‐on technical authority, authoring reusable deployment solutions, configuring telemetry collectors ...

DevOps Engineer

Hiring Organisation
Opus Recruitment Solutions
Location
Newcastle upon Tyne, Shiremoor, Tyne & Wear, United Kingdom
Employment Type
Permanent
Salary
£45000 - £58000/annum
Code. Develop automation solutions to improve delivery speed, reliability, and scalability. Support containerised and cloud-native environments using Kubernetes and Docker. Implement monitoring, logging, observability, and alerting solutions. Promote DevSecOps practices, security controls, and operational excellence. Collaborate with engineers, architects, testers, and product teams to deliver high-quality software. Lead … Bicep. Knowledge of Docker, Kubernetes, OpenShift, and container orchestration. Scripting skills in Python, Bash, PowerShell, Node.js, or TypeScript. Experience with monitoring and observability tools such as Grafana, Prometheus, Splunk, ELK, or CloudWatch. Strong understanding of DevSecOps, networking, cloud architecture, and Agile delivery. Leadership Responsibilities Provide technical leadership across DevOps ...

Lead DevOps Engineer

Location
Greater London, England, United Kingdom
partnership with engineering, infrastructure, security, and operations teams.It is an opportunity to improve reliability, scalability, security, and delivery while advancing DevOps, platform engineering, observability, and AI-enabled infrastructure tooling.**Responsibilities*** Lead the design and evolution of scalable, secure, and highly available trading infrastructure.* Provide technical direction, mentorship, and guidance … assisted engineering tools to accelerate development, automation, troubleshooting, documentation, and operational workflows.* Identify, design, and help implement AI-enabled operational capabilities for infrastructure automation, observability, incident response, and platform engineering.* Contribute to the strategy for safe, practical, and secure adoption of AI across infrastructure and DevOps practices.* Lead and participate ...

Platform Engineer

Location
Greater London, England, United Kingdom
data and AI workflows. It’s an excellent opportunity for an experienced Platform/DevOps Engineer to work with cloud, Kubernetes, CI/CD, observability, and emerging AI infrastructure while helping establish scalable, secure, and reliable engineering practices. This is an opportunity to join an innovative, progressive, and collaborative team. … agent orchestration AI Evaluation & Quality: Eval harnesses and golden datasets, LLM-as-judge and human-in-the-loop review, regression suites, and red-teaming Observability & Monitoring: Prometheus, Grafana, Datadog, Splunk, Elastic/ELK, OpenTelemetry, including GenAI tracing and token, latency, and cost telemetry Platform Security & Policy-as-Code: HashiCorp Vault ...

Software Engineering Tech Lead (SRE + AI)

Location
Greater London, England, United Kingdom
monitor, diagnose, and auto-remediate global SaaS infrastructure. What you'll do Technical Leadership & Architecture: Define the technical roadmap and architecture for AI-assisted observability, automated incident response, and self-healing cloud infrastructure. Agentic Workflows & Tooling: Design and build production-grade AI agents, MCP tool integrations, and deterministic evaluation pipelines … automation. Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments. Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated RCA. Preferred Qualifications AI & Agentic Systems: Hands-on experience building LLM pipelines, AI Agents, Model ...

Software Engineer

Location
Greater London, England, United Kingdom
diagnose, and auto-remediate global SaaS infrastructure. What You'll Do Technical Design & Architecture: Design and implement high-resilience software systems for AI-assisted observability, automated incident response, and self-healing cloud infrastructure. Agentic Workflows & Tooling: Design, build, and maintain production-grade AI agents, MCP tool integrations, and deterministic evaluation … automation. Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments. Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated root cause analysis (RCA). Preferred Qualifications AI & Agentic Systems: E xperience building LLM pipelines ...