1,876 to 1,900 of 5,338 Observability Jobs

Lead Site Reliability Engineer (Dynatrace)

Location
London, United Kingdom
looking for an experienced Site Reliability Engineer/Observability Engineer with deep Dynatrace expertise to join a major technology and platform engineering programme. This is not a role for someone who has simply used Dynatrace dashboards. We're looking for an engineer who has been involved in the implementation, configuration … technical SME within complex production environments. What we're looking for Strong hands-on Dynatrace implementation and administration experience Experience designing and implementing observability/monitoring solutions end-to-end Strong SRE and production engineering background Experience configuring instrumentation, metrics, alerting and monitoring Understanding of technologies such as OneAgent, ActiveGate ...

Lead Site Reliability Engineer (Dynatrace)

Hiring Organisation
SF Partners
Location
South West England, United Kingdom
Employment Type
Full-Time
Salary
£80,000 - £100,000 per annum
looking for an experienced Site Reliability Engineer/Observability Engineer with deep Dynatrace expertise to join a major technology and platform engineering programme. This is not a role for someone who has simply used Dynatrace dashboards. We're looking for an engineer who has been involved in the implementation, configuration … technical SME within complex production environments. What we're looking for Strong hands-on Dynatrace implementation and administration experience Experience designing and implementing observability/monitoring solutions end-to-end Strong SRE and production engineering background Experience configuring instrumentation, metrics, alerting and monitoring Understanding of technologies such as OneAgent, ActiveGate ...

Fullstack Engineer

Location
Bracknell, England, United Kingdom
Design and maintain scalable Go-based microservices Build and support REST and gRPC APIs Develop integrations and event-driven solutions across distributed systems Improve observability, reliability, scalability, and security Deploy and support applications in AWS Work with Docker, Kubernetes, and CI/CD pipelines Participate in production support … Nice to have: AWS experience Nice to have: Kubernetes and container orchestration Nice to have: Event-driven architectures and messaging platforms Nice to have: Observability, monitoring, and distributed tracing Nice to have: Experience working in a SaaS product organisation Core Competencies Demonstrates expertise in building modern web applications using React ...

Senior SRE Engineer

Hiring Organisation
SF Partners
Location
Birmingham, West Midlands (County), United Kingdom
Employment Type
Permanent
Salary
£100000 - £110000/annum great training and progression opp
largest and most complex enterprise platforms. Working within a highly skilled, multi-disciplinary engineering team, this Senior SRE Engineer will take technical ownership of observability and reliability across large-scale AWS environments, helping engineering teams improve platform performance, resilience and availability. A key focus of the role will … enterprise environments - End-to-end Dynatrace implementation experience - Experience designing and deploying Dynatrace across complex cloud environments - Strong understanding of APM, infrastructure monitoring and observability - Dynatrace configuration, dashboards, alerting and performance monitoring - Experience integrating Dynatrace into AWS and wider engineering/tooling ecosystems - Strong understanding of SLIs, SLOs, availability, reliability ...

Site Reliability Engineer with Python

Hiring Organisation
BC Forward
Location
Charlotte, North Carolina, United States
Employment Type
Permanent
Salary
USD Hourly
Site Reliability Engineer with Python to join our dynamic team. The ideal candidate will have strong experience in Python development, Linux administration, infrastructure automation, observability, and cloud-native platforms and a proven ability to improve reliability, enhance observability, and drive operational efficiency at scale. Responsibilities: Own reliability and operational health … core platform services and components. Design, implement, and maintain automation for infrastructure and operational workflows. Improve platform and application observability using modern monitoring and logging tools. Lead incident response, troubleshoot production issues, and drive post-incident remediation. Partner with core engineering teams on platform modernization and cloud migration initiatives. Harden ...

SRE Engineer

Location
Greater London, England, United Kingdom
automate the deployment of our software Automate the provisioning and management of our infrastructure using Infrastructure as Code (IaC) tools Define, implement, and maintain observability solutions for our applications to ensure we can proactively detect system degradation, easily understand system state, and quickly diagnose issues Diagnose and resolve production issues … must have and one should be good at coding in terraform CICD Tools hands on : Jenkins , GitHub , GitHub Actions, Cloud Deployment pipelines Observability Tools - Splunk//Graphana/Datadog and Distributed Tracing, ELF, & Dynatrace Problem-Solving : Proven ability to troubleshoot complex issues in distributed systems and debug problems effectively. ...

Machine Learning Operations Engineer

Location
Greater London, England, United Kingdom
build and operate the platform capabilities that take machine-learning models from experimentation into reliable production services You’ll own the automation, deployment, observability and operational controls around the ML lifecycle, working closely with research engineers, software engineers, platform teams and product teams This is not a research role. … model metadata and reproducibility across research and production Build reusable tooling and platform capabilities that support multiple models and engineering teams Model serving and observability Deploy and operate batch and online inference services in containerised cloud environments Define and meet availability, latency, throughput and recovery objectives for ML services Monitor ...

Software Engineer, Real-Time

Location
Nottingham, England, United Kingdom
that show how reliably and quickly market data reaches customers. You will work with experienced engineers to develop production software, measure performance and enable observability of a globally distributed platform. Prior market-data or observability experience is not required; we are looking for strong engineering fundamentals, curiosity and a willingness … hands‐on development role for an engineer at an early stage of their career who wants to build experience in mission-critical distributed systems, observability, cloud-native engineering and real‐time financial technology. What You’ll Be Doing Build and improve software components that aggregate, correlate and present data from ...

Director of Site Reliability Engineering

Location
Greater London, England, United Kingdom
robust incident management frameworks and lead major incident response activities for critical systems Implement blameless postmortems and deliver systemic improvements across production environments Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets … operational tooling to reduce manual processes Requirements Strong background in Site Reliability Engineering, DevOps, or platform operations in complex, distributed environments Expertise in observability platforms, troubleshooting distributed systems, and telemetry‐driven insights Hands‐on experience with automation, Infrastructure as Code (Terraform or CloudFormation), and CI/CD practices Deep understanding ...

Software Engineer, Real-Time

Hiring Organisation
London Stock Exchange Group
Location
Nottingham, Nottinghamshire, United Kingdom
Salary
£ 70 K
that show how reliably and quickly market data reaches customers. You will work with experienced engineers to develop production software, measure performance and enable observability of a globally distributed platform. Prior market-data or observability experience is not required; we are looking for strong engineering fundamentals, curiosity and a willingness … hands-on development role for an engineer at an early stage of their career who wants to build experience in mission-critical distributed systems, observability, cloud-native engineering and real-time financial technology.WHAT YOU’LL BE DOINGBuild and improve software components that aggregate, correlate and present data from over ...

Software Engineer, Real-Time

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
that show how reliably and quickly market data reaches customers. You will work with experienced engineers to develop production software, measure performance and enable observability of a globally distributed platform. Prior market-data or observability experience is not required; we are looking for strong engineering fundamentals, curiosity and a willingness … hands-on development role for an engineer at an early stage of their career who wants to build experience in mission-critical distributed systems, observability, cloud-native engineering and real-time financial technology. WHAT YOU'LL BE DOINGBuild and improve software components that aggregate, correlate and present data from over ...

Senior DevOps Engineer

Location
East Midlands, England, United Kingdom
/CD pipelines and deployment orchestration. Support Kubernetes and OpenShift platform troubleshooting and optimisation. Deliver secure and compliant infrastructure solutions. Implement monitoring, logging, and observability tooling across environments. Collaborate with engineering, architecture, and delivery teams to improve deployment efficiency and platform reliability. Champion automation-first approaches to infrastructure and application … designing and maintaining enterprise-scale CI/CD pipelines . Strong understanding of cloud security and secure delivery practices. Experience implementing monitoring, logging, and observability solutions. Ability to define technical standards, governance, and reusable deployment frameworks. Experience working within large-scale enterprise transformation programmes. Desirable Skills Experience within highly regulated ...

Senior Platform Engineer: AI-Ready Infra & Security

Location
Slough, England, United Kingdom
scalable platform features, automate operations, and build self-service workflows. The role emphasizes strong Python, Linux, Terraform/Ansible, Docker and Kubernetes proficiency, plus observability with Prometheus, Grafana and OpenTelemetry. #J-18808-Ljbffr ...

Senior Azure DevOps Engineer: Kubernetes & CI/CD

Location
Greater London, England, United Kingdom
London to build and improve containerised environments using Docker and Kubernetes. You will work in a hybrid model across Azure, Terraform, CI/CD, observability and automation. The role requires strong Linux experience and hands-on expertise in deploying enterprise-scale Azure solutions, IaC, and large-scale CI/ ...

DevOps Lead

Hiring Organisation
TurleyWay Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent
manage and optimise containerised environments using Kubernetes. Oversee source control, CI/CD pipelines and development workflows within GitHub and build and maintain monitoring, observability and alerting solutions using Grafana. To be considered you be able to demonstrate proven experience in a DevOps Lead, Senior DevOps Engineer, or Platform Engineering … expertise in Kubernetes and container orchestration technologies, administering and managing GitHub and CI/CD pipelines. Solid knowledge of Grafana and modern monitoring/observability practices and extensive database exposure, including performance, administration, and optimisation of enterprise database environments. In return we offer a competitive basic salary plus bonus scheme ...

DevOps Manager

Hiring Organisation
TurleyWay Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent
manage and optimise containerised environments using Kubernetes. Oversee source control, CI/CD pipelines and development workflows within GitHub and build and maintain monitoring, observability and alerting solutions using Grafana. To be considered you be able to demonstrate proven experience in a DevOps Lead, Senior DevOps Engineer, or Platform Engineering … expertise in Kubernetes and container orchestration technologies, administering and managing GitHub and CI/CD pipelines. Solid knowledge of Grafana and modern monitoring/observability practices and extensive database exposure, including performance, administration, and optimisation of enterprise database environments. In return we offer a competitive basic salary plus bonus scheme ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
Remote, UK
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real-world processes. … Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., Design evaluation, monitoring, and observability to ensure reliability and readiness for production. js and TypeScript. Build modern, responsive React frontends to make AI useful in customer workflows. Architect cloud-native systems on AWS (Lambda, ECS/Fargate, API Gateway ...

Software Engineer, Cloud Infrastructure

Location
Greater London, England, United Kingdom
high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure. About the Role All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable … solving novel problems in performance and scalability. Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience. Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale. Developer Productivity: Create tools, environments, and workflows that ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
Pickmere, Cheshire, UK
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real-world processes. … Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., Design evaluation, monitoring, and observability to ensure reliability and readiness for production. js and TypeScript. Build modern, responsive React frontends to make AI useful in customer workflows. Architect cloud-native systems on AWS (Lambda, ECS/Fargate, API Gateway ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
Cullompton, Devon, UK
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real-world processes. … Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., Design evaluation, monitoring, and observability to ensure reliability and readiness for production. js and TypeScript. Build modern, responsive React frontends to make AI useful in customer workflows. Architect cloud-native systems on AWS (Lambda, ECS/Fargate, API Gateway ...

Remote Senior AI Software Engineer

Location
Warwickshire, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real-world processes. … Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., Design evaluation, monitoring, and observability to ensure reliability and readiness for production. js and TypeScript. Build modern, responsive React frontends to make AI useful in customer workflows. Architect cloud-native systems on AWS (Lambda, ECS/Fargate, API Gateway ...

Remote Senior AI Software Engineer

Location
Bushey, Hertfordshire, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real-world processes. … Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., Design evaluation, monitoring, and observability to ensure reliability and readiness for production. js and TypeScript. Build modern, responsive React frontends to make AI useful in customer workflows. Architect cloud-native systems on AWS (Lambda, ECS/Fargate, API Gateway ...

Senior Software Engineer, ML Infrastructure Roku, Inc.

Location
Cambridge, England, United Kingdom
conversational AI experiences used across millions of Roku devices. The team works across fulfilment ranking, model delivery, offline and online evaluation, low-latency services, observability and product quality. Its published work includes shared model-serving and MLOps paths, automated evaluation and retraining, caching and telemetry, and agent-assisted release … agent, including tool routing, retrieval, guardrails and answer caching. Design caching as an intentional latency and cost lever for high-volume services. Build observability for ML and LLM systems, including latency attribution, quality metrics, tracing and per-request cost. Improve the reliability and operability of distributed systems, and lead ...

Platform Engineer - Common Platform

Location
Greater London, England, United Kingdom
develop automation and platform tooling Designing and implementing AWS-native infrastructure and services Building and maintaining standard CI/CD pipelines and workflows Supporting observability and monitoring capabilities across the platform Managing Kubernetes upgrades and platform improvements Supporting production incidents and resolving complex infrastructure issues Reviewing and approving infrastructure … design Enterprise security and governance Compliance frameworks and organisational policies Implementing platform changes across distributed teams Cloud cost optimisation Helm and Kubernetes management tooling Observability and monitoring You’ll be someone who enjoys building platforms and solving engineering problems rather than simply maintaining existing infrastructure. Strong communication and collaboration skills ...

AI-Driven Cloud & Platform Engineer for Digital Factory

Location
Manchester, England, United Kingdom
implement CI/CD automation, infrastructure-as-code, and self-service tooling, leveraging Docker, Kubernetes, and multiple cloud providers. You’ll define observability and reliability practices, collaborate with architects, and mentor teammates. You should bring hands-on experience in cloud, DevOps, and SRE disciplines, with a willingness to work across ...