26 to 50 of 3,821 Permanent Observability Jobs

DBA Lead

Location
City Of London, England, United Kingdom
availability, performance, security, and operational stability. The team manages MySQL, SQL Server, PostgreSQL, Oracle, and cloud database technologies across hybrid environments while driving automation, observability, disaster recovery readiness, and continuous improvement initiatives. Working closely with engineering, infrastructure, security, and DevOps teams, DBAs play a key role in delivering resilient …/or Terragrunt CI/CD pipeline implementation using GitHub Actions Git-centric development workflows Automation experience using: Python, Bash/Shell Scripting, Ansible Observability & Reliability Engineering Experience with enterprise monitoring and observability platforms including Grafana/Datadog/ELK Azure Monitor CloudWatch Understanding of SRE concepts including: Service Level ...

Java Backend Developer

Location
City Of London, England, United Kingdom
Investigate complex production incidents and perform root cause analysis. Implement permanent fixes for recurring production issues. Drive improvements in application stability, performance, monitoring, and observability . Participate in solution design, technical discussions, and code reviews. Develop unit and integration tests and ensure appropriate test coverage. Support application releases, deployments … Kubernetes and containerized applications . Experience with Microservices and distributed application architecture. Experience with Docker and cloud-native application development. Familiarity with monitoring and observability tools such as Splunk, Dynatrace, Datadog, Prometheus, or Grafana. Experience with CI/CD tools such as Jenkins, GitLab CI, GitHub Actions, or Azure DevOps. ...

Vice President, Credit Technology Engineering

Hiring Organisation
Ares Management
Location
London, UK
Employment Type
Full-time
credit technology ecosystem. Define architecture standards for distributed systems, APIs, event-driven architectures, and microservices. Drive engineering decisions related to scalability, resiliency, performance, observability, and operational excellence. Evaluate new technologies and architectural patterns to improve engineering effectiveness and platform capabilities. Application Development & Technology Delivery Lead development across: Credit deal management … technologies. Implement containerized solutions utilizing Kubernetes and cloud platform services. Drive modernization efforts focused on cloud adoption, scalability, resilience, and operational efficiency. Ensure proper observability, monitoring, logging, and operational support across production environments. Integration & Event-Driven Architecture Design and implement API-first architectures and integration frameworks. Leverage messaging platforms, event ...

Senior Software Developer (Python)

Location
Greater London, England, United Kingdom
product development, including ingestion, transformation, model deployment, and visualisation workflows. Ensure software solutions meet security, reliability, scalability, and performance requirements. Support operational excellence through observability, monitoring, troubleshooting, and incident resolution. Contribute to technical design reviews, architecture discussions, and engineering best practices. Mentor junior developers and contribute to fostering a strong … Pandas, Apache Beam, Spark, or similar frameworks. Strong SQL skills and experience with analytical databases such as BigQuery, PostgreSQL, or ClickHouse. Familiarity with observability and monitoring tools such as Prometheus, Grafana, Splunk, or OpenTelemetry. Experience deploying and supporting large-scale distributed systems in Kubernetes environments. Knowledge of graph analytics, network ...

Senior DevOps Engineer

Location
City of Edinburgh, Scotland, United Kingdom
integration platforms. Drive incident management, root cause analysis, and post-incident reviews. Improve platform availability, scalability, and performance through automation and engineering improvements. Monitoring & Observability Implement enterprise monitoring, alerting, logging, and observability solutions. Grafana ELK/Elastic Stack OpenTelemetry Proactively identify issues before they impact customers and engineering teams. Develop ...

Cloud and Platform Engineer-Consultant-AI and Digital Factory

Location
Manchester, England, United Kingdom
team’s needs• Be responsible for infrastructure and platform provisioning, configuration, build, deployment, and monitoring using Infrastructe-as-Code• Define and implement what good observability practices look like for clients, track reliability metrics (SLIs/SLOs/error budgets) and participate in incident response where relevant to your specialism• Build … agile team environmentsSite Reliability Engineering• Defining and operating against SLIs/SLOs/error budgets• Incident response, on-call practices, and post-incident review• Observability and monitoring: Grafana, Dynatrace, CloudWatch, Prometheus, Datadog, OpenTelemetryGeneral• Scripting (bash/shell)• Testing tooling: Selenium, Cucumber, etc.• Ability to flexibly support a variety of technologies ...

Cloud and Platform Engineer-Consultant-AI and Digital Factory

Location
Greater London, England, United Kingdom
team’s needs• Be responsible for infrastructure and platform provisioning, configuration, build, deployment, and monitoring using Infrastructe-as-Code• Define and implement what good observability practices look like for clients, track reliability metrics (SLIs/SLOs/error budgets) and participate in incident response where relevant to your specialism• Build … agile team environmentsSite Reliability Engineering• Defining and operating against SLIs/SLOs/error budgets• Incident response, on-call practices, and post-incident review• Observability and monitoring: Grafana, Dynatrace, CloudWatch, Prometheus, Datadog, OpenTelemetryGeneral• Scripting (bash/shell)• Testing tooling: Selenium, Cucumber, etc.• Ability to flexibly support a variety of technologies ...

Cloud and Platform Engineer-Consultant-AI and Digital Factory

Location
Newcastle upon Tyne, England, United Kingdom
team’s needs• Be responsible for infrastructure and platform provisioning, configuration, build, deployment, and monitoring using Infrastructe-as-Code• Define and implement what good observability practices look like for clients, track reliability metrics (SLIs/SLOs/error budgets) and participate in incident response where relevant to your specialism• Build … agile team environmentsSite Reliability Engineering• Defining and operating against SLIs/SLOs/error budgets• Incident response, on-call practices, and post-incident review• Observability and monitoring: Grafana, Dynatrace, CloudWatch, Prometheus, Datadog, OpenTelemetryGeneral• Scripting (bash/shell)• Testing tooling: Selenium, Cucumber, etc.• Ability to flexibly support a variety of technologies ...

Software Engineer III - SRE

Hiring Organisation
JP Morgan Chase
Location
Bournemouth, Dorset, United Kingdom
Salary
£ 70 K
e.g., Python, Java/Spring Boot, .NET).Experience with infrastructure as code tools (Terraform, CloudFormation, etc.).Strong knowledge of Linux/Unix systems.Experience with observability and monitoring tools (e.g., Grafana, Dynatrace, Prometheus, Datadog, Splunk).Familiarity with CI/CD tools (e.g., Jenkins, GitLab) and container orchestration (e.g., Docker, Kubernetes ...

Software Engineer III - SRE

Hiring Organisation
JP Morgan Chase
Location
Bournemouth, Dorset, UK
Employment Type
Full-time
Python, Java/Spring Boot, .NET).Experience with infrastructure as code tools (Terraform, CloudFormation, etc.).Strong knowledge of Linux/Unix systems. Experience with observability and monitoring tools (e.g., Grafana, Dynatrace, Prometheus, Datadog, Splunk).Familiarity with CI/CD tools (e.g., Jenkins, GitLab) and container orchestration (e.g., Docker, Kubernetes ...

Senior Lead Software Engineer - Python / Go

Location
Auchentibber, Scotland, United Kingdom
eliminate platform bottlenecks. Define and promote paved paths and self-service workflows for developers. Implement real-time telemetry pipelines and workflows for platform observability and analytics. Champion adoption of productivity tools through documentation, training, and engagement with the developer community. Standardize use of AI-assisted coding tools and AI-powered ...

Senior Lead Platform Engineer - Python

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
eliminate platform bottlenecks. Define and promote paved paths and self-service workflows for developers. Implement real-time telemetry pipelines and workflows for platform observability and analytics. Champion adoption of productivity tools through documentation, training, and engagement with the developer community. Standardize use of AI-assisted coding tools and AI-powered ...

Principal Cloud Engineer (Terraform), London

Location
Greater London, England, United Kingdom
Apply data quality and validation frameworks to ensure accuracy, completeness, and freshness of cost and usage data across all cloud providers; instrument pipelines with observability tooling to surface issues proactively. Build and maintain reusable data assets - curated datasets, aggregations, and data marts - that power FinOps dashboards, showback/chargeback reporting ...

Senior Software Engineer

Hiring Organisation
AECOM
Location
London, UK
Employment Type
Full-time
deployment pipelines such as GitHub Actions. Depth in distributed or event-driven systems, asynchronous processing, or data-intensive and high-throughput applications. Experience with observability and performance tuning, including tools such as Prometheus, Grafana, ELK or Datadog, or optimising CPU-, GPU- or distributed ML workloads. Experience in startups, scaleups ...

Monitoring & Observability Engineer (Dynatrace)

Location
Greater London, England, United Kingdom
Monitoring & Observability Engineer (Dynatrace) Location: UK - London, UK - Reading, UK - Hatfield, UK - Nottingham, UK - Manchester, UK - Milton Keynes, UK - Birmingham | Job-ID: 214264 | Contract type: Standard | Business Unit: IT Consulting Life on the team At Computacenter, you’ll be joining a world-class team of over 1,000 skilled professionals … across the UK, Germany, France, and India, delivering complex, enterprise-grade IT solutions and consultancy across infrastructure, cloud, and modern operations. As a Monitoring & Observability Engineer, you'll work in high-impact delivery teams that support some of the world’s most well-known organisations. You’ll play ...

Monitoring & Observability Engineer (Dynatrace)

Hiring Organisation
Computacenter
Location
London, UK
Employment Type
Full-time
across the UK, Germany, France, and India, delivering complex, enterprise-grade IT solutions and consultancy across infrastructure, cloud, and modern operations. As a Monitoring & Observability Engineer, you'll work in high-impact delivery teams that support some of the world's most well-known organisations. You'll play … reliability across their IT estates—contributing to their operational success through proactive insight and incident prevention. What you'll doDesign, implement, and manage observability solutions using industry-leading tools such as Dynatrace (primary), Grafana, and SplunkCollect and analyse telemetry data (metrics, logs, traces, events) to diagnose and resolve system ...

Observability Engineer - Assistant Vice President

Location
Greater London, England, United Kingdom
Site Reliability Engineer (SRE) - Assistant Vice President is a technical professional responsible for the hands‐on execution, technical implementation, and deployment of SRE and observability principles in a complex, critical, and large-scale multi-disciplinary environment. In this role, you will apply a deep understanding of multiple technology domains … technical contributor, you will drive the migration of applications from existing monitoring tools (Geneos ITRS, Prometheus, ELK, Splunk, AppDynamics, etc.) to Google Cloud Observability (GCO) and Grafana using OpenTelemetry (OTel) as the instrumentation standard. You will act as a hands‐on technical authority, authoring reusable deployment solutions, configuring telemetry collectors ...

DevOps Engineer

Hiring Organisation
Opus Recruitment Solutions
Location
Newcastle upon Tyne, Shiremoor, Tyne & Wear, United Kingdom
Employment Type
Permanent
Salary
£45000 - £58000/annum
Code. Develop automation solutions to improve delivery speed, reliability, and scalability. Support containerised and cloud-native environments using Kubernetes and Docker. Implement monitoring, logging, observability, and alerting solutions. Promote DevSecOps practices, security controls, and operational excellence. Collaborate with engineers, architects, testers, and product teams to deliver high-quality software. Lead … Bicep. Knowledge of Docker, Kubernetes, OpenShift, and container orchestration. Scripting skills in Python, Bash, PowerShell, Node.js, or TypeScript. Experience with monitoring and observability tools such as Grafana, Prometheus, Splunk, ELK, or CloudWatch. Strong understanding of DevSecOps, networking, cloud architecture, and Agile delivery. Leadership Responsibilities Provide technical leadership across DevOps ...

Lead DevOps Engineer

Location
Greater London, England, United Kingdom
partnership with engineering, infrastructure, security, and operations teams.It is an opportunity to improve reliability, scalability, security, and delivery while advancing DevOps, platform engineering, observability, and AI-enabled infrastructure tooling.**Responsibilities*** Lead the design and evolution of scalable, secure, and highly available trading infrastructure.* Provide technical direction, mentorship, and guidance … assisted engineering tools to accelerate development, automation, troubleshooting, documentation, and operational workflows.* Identify, design, and help implement AI-enabled operational capabilities for infrastructure automation, observability, incident response, and platform engineering.* Contribute to the strategy for safe, practical, and secure adoption of AI across infrastructure and DevOps practices.* Lead and participate ...

Platform Engineer

Location
Greater London, England, United Kingdom
data and AI workflows. It’s an excellent opportunity for an experienced Platform/DevOps Engineer to work with cloud, Kubernetes, CI/CD, observability, and emerging AI infrastructure while helping establish scalable, secure, and reliable engineering practices. This is an opportunity to join an innovative, progressive, and collaborative team. … agent orchestration AI Evaluation & Quality: Eval harnesses and golden datasets, LLM-as-judge and human-in-the-loop review, regression suites, and red-teaming Observability & Monitoring: Prometheus, Grafana, Datadog, Splunk, Elastic/ELK, OpenTelemetry, including GenAI tracing and token, latency, and cost telemetry Platform Security & Policy-as-Code: HashiCorp Vault ...

Software Engineering Tech Lead (SRE + AI)

Location
Greater London, England, United Kingdom
monitor, diagnose, and auto-remediate global SaaS infrastructure. What you'll do Technical Leadership & Architecture: Define the technical roadmap and architecture for AI-assisted observability, automated incident response, and self-healing cloud infrastructure. Agentic Workflows & Tooling: Design and build production-grade AI agents, MCP tool integrations, and deterministic evaluation pipelines … automation. Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments. Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated RCA. Preferred Qualifications AI & Agentic Systems: Hands-on experience building LLM pipelines, AI Agents, Model ...

Software Engineer

Location
Greater London, England, United Kingdom
diagnose, and auto-remediate global SaaS infrastructure. What You'll Do Technical Design & Architecture: Design and implement high-resilience software systems for AI-assisted observability, automated incident response, and self-healing cloud infrastructure. Agentic Workflows & Tooling: Design, build, and maintain production-grade AI agents, MCP tool integrations, and deterministic evaluation … automation. Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments. Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated root cause analysis (RCA). Preferred Qualifications AI & Agentic Systems: E xperience building LLM pipelines ...

Sr. Site Reliability Engineer/ SWE

Location
Ham, England, United Kingdom
platform strategy. In this role, you’ll ensure our development platform and tools let engineers focus on innovation instead of infrastructure. You’ll promote observability best practices and automate resolution of recurring issues, working closely with software engineering teams to support security, availability, and performance. Responsibilities include triaging issues, collaborating … implement, and maintain systems for high availability, scalability, and performance. Monitor and improve application reliability through proactive measures and incident response. Develop and maintain observability solutions (metrics, logging, tracing). Participate in on‐call rotations and drive root cause analysis for incidents. Collaboration & Continuous Improvement Partner with engineering teams ...

Sr. Site Reliability Engineer/ SWE

Hiring Organisation
Visa
Location
Basingstoke, Hampshire, UK
Employment Type
Full-time
platform strategy. In this role, you'll ensure our development platform and tools let engineers focus on innovation instead of infrastructure. You'll promote observability best practices and automate resolution of recurring issues, working closely with software engineering teams to support security, availability, and performance. Responsibilities include triaging issues, collaborating … implement, and maintain systems for high availability, scalability, and performance. Monitor and improve application reliability through proactive measures and incident response. Develop and maintain observability solutions (metrics, logging, tracing). Participate in on-call rotations and drive root cause analysis for incidents. Collaboration & Continuous Improvement Partner with engineering teams ...

Full Stack Lead Architect, Vice President

Location
Belfast City District, Northern Ireland, United Kingdom
frameworks, AI Agents, and intelligent automation solutions. Architect event-driven systems using Kafka, Solace, and distributed caching technologies. Establish engineering standards, SDLC best practices, observability, resiliency, and security controls. Collaborate with Architecture, Product, Data Engineering, Infrastructure, Security, and Business teams globally. Lead technical reviews, solution governance, and critical production issue …/Financial Services experience. Experience with Data Lake, Lakehouse, and Analytics platforms. Knowledge of OAuth2, JWT, Zero Trust, and API Security standards. Experience with observability platforms such as Grafana, Dynatrace, Splunk, Prometheus, or OpenTelemetry. Leadership Expectations Provide technical leadership across multiple teams and strategic initiatives. Drive architecture decisions and engineering ...