2,301 to 2,325 of 5,534 Observability Jobs

Senior SRE: Cloud Platform Reliability & Observability

Location
Cambridge, England, United Kingdom
reliability, performance, and scalability of its cloud platforms from its Cambridge, UK operations. You will work across DevOps, engineering, and platform teams to implement observability, reliability frameworks, and automated deployment practices. The role emphasizes incident response, capacity planning, and continuous improvement, with a focus on IaC, SOO principles, and cross ...

Senior Azure SRE — Platform Reliability & Observability

Location
Southampton, England, United Kingdom
Senior Site Reliability Engineer to join its Azure-focused Cloud Platform Engineering team in Southampton on a hybrid basis. The role focuses on observability, reliability, security and scalable cloud platforms, with leadership across operations and engineering teams. The ideal candidate has extensive Azure and Kubernetes experience, strong automation skills ...

Senior SRE: Architecting Reliability & Observability

Location
Nottingham, England, United Kingdom
reliability from day one, and may be called for major incidents. This role demands proactive leadership, hands-on expertise in AWS/Azure, Linux, observability, and cost optimization, with a focus on scalable, secure platforms and a culture of continuous improvement. #J-18808-Ljbffr ...

Senior SRE Technical Lead — Reliability & Observability

Location
Greater London, England, United Kingdom
seeking a Technical Lead SRE in Greater London. In this role, you will enhance the reliability engineering capabilities, collaborating with various teams to establish observability standards and ensure operational excellence. The ideal candidate will have over 10 years of experience in SRE or related fields, strong AWS and Kubernetes skills ...

Senior Observability Solution Architect – Pre-Sales

Location
Greater London, England, United Kingdom
leading observability platform in Greater London is seeking an experienced Solution Engineer to join their team. This role involves collaborating with account executives on technical sales cycles, delivering impactful presentations, and overseeing technical aspects of the process. The ideal candidate will have a minimum of 5 years in a customer ...

Senior Platform SRE: Reliable, Scalable Observability

Location
Greater London, England, United Kingdom
team to own availability and performance of mission-critical services, and to lead on-call and incident response. You will build tooling, improve observability, and drive platform scalability in collaboration with product teams, while lowering costs and improving developer productivity. The ideal candidate has around 8 years of distributed systems ...

Senior Product Manager, Observability New UK

Location
United Kingdom
partnering with engineering, design, and go-to-market to turn customer and operational problems into shippable outcomes. As a Senior Technical Product Manager for Observability, you own the platform that gives customers and internal operators real-time visibility into their GPU fleet: the telemetry pipeline that scrapes data from physical … infrastructure, the aggregation and storage layer, and the observability surfaces (logs, metrics, and traces) that enable fleet management, incident response, and alerting at scale. You partner daily with Fleet Software, Network Engineering, Data Centre Operations, and customer teams to make fleet health visible, actionable, and reliable as Nscale scales from ...

Observability & AIOps Engineer (all genders)

Hiring Organisation
Lam Research
Location
Villach, Kärnten, Austria
Employment Type
Permanent
Salary
EUR Annual
innovative information system solutions and services. Together, we support users globally with data, information, and systems to achieve their business objectives. Aufgaben As an Observability & AIOps Engineer, you will build the intelligence layer that enables enterprise systems and AI agents to understand operational health in real time. You will transform … autonomous operations by enabling AI agents to perceive, reason about, and respond to complex system behavior. What you'll do Design and implement enterprise observability strategies across cloud, infrastructure, applications, and platforms. Build telemetry pipelines that ingest logs, metrics, traces, events, and topology information. Deploy and tune observability platforms ...

Cloud SRE: Automation, Kubernetes & Observability

Location
Greater London, England, United Kingdom
LexisNexis Risk Solutions seeks a Site Reliability Engineer to design, build and operate cloud infrastructure across AWS and Azure. You will enable engineering teams through infrastructure as code, automation, and platform reliability practices. This hands ...

DevOps Platform Engineer (Hybrid • Azure & Observability)

Location
Aberdeen City, Scotland, United Kingdom
Aize AS in Aberdeen is hiring a DevOps Engineer to build, deploy and operate cloud infrastructure for our Internal Development Platform. You will focus on networking, containers, and infrastructure as code to move features from ...

SRE Engineer – FinTech Reliability, Observability & Cloud

Location
Greater London, England, United Kingdom
Hamilton Barnes Associates Limited is seeking a Site Reliability Engineer to work at the intersection of software engineering and infrastructure. You'll develop internal platforms, tooling, and automation across Linux, distributed systems, and cloud-native ...

Senior Site Reliability Engineer - Cloud & Observability

Location
Greater London, England, United Kingdom
A leading global financial markets company is seeking a Senior Engineer in Site Reliability. This role involves maintaining service level objectives, enhancing system reliability, and automating to ensure scalability. With a strong focus on cloud ...

Senior Cloud Platform SRE & Observability Lead

Location
Southampton, England, United Kingdom
NICE in Southampton is seeking a Site Reliability Engineer with over 6 years of experience in SRE. The role entails managing production systems and developing automation solutions to enhance performance and reliability. Candidates must have ...

Site Reliability Engineer

Location
Manchester, England, United Kingdom
Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices. Full-time Closes 05/08/2026 You will have strong software engineering skills, approaching system reliability and observability as a software problem — protecting, providing … optimise system health, while engineering automation and tooling for effective service management. Collaboration is key, working across multiple functions to embed reliability and observability best practices throughout the software development life cycle. Your contributions will ensure our systems meet user demands and foster a culture of continuous improvement. This role ...

Software Engineer, SRE

Location
United Kingdom
Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices. Full-time Closes 30/09/2026 You will have strong software engineering skills, approaching system reliability and observability as a software problem — protecting, providing … optimise system health, while engineering automation and tooling for effective service management. Collaboration is key, working across multiple functions to embed reliability and observability best practices throughout the software development life cycle. Your contributions will ensure our systems meet user demands and foster a culture of continuous improvement. This role ...

Software Engineer, SRE

Location
Manchester, England, United Kingdom
Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices. Full-time Closes 30/09/2026 You will have strong software engineering skills, approaching system reliability and observability as a software problem — protecting, providing … optimise system health, while engineering automation and tooling for effective service management. Collaboration is key, working across multiple functions to embed reliability and observability best practices throughout the software development life cycle. Your contributions will ensure our systems meet user demands and foster a culture of continuous improvement. This role ...

Data Reliability Engineer

Hiring Organisation
Ashdown Group
Location
London, UK
Employment Type
Full-time
work from home 2 days per week. This is a high-impact role focused on improving data quality, reducing incidents, and building scalable observability across a modern enterprise data platform. You'll help ensure data across the organisation is accurate, reliable, and trusted for critical business decision-making. … style roles, with strong SQL and Python skills and experience working in modern cloud-based data environments. Hands-on experience with data observability tools such as Grafana, Monte Carlo, or Acceldata, and data governance/quality platforms like Informatica, Collibra or Microsoft Purview is highly desirable. Experience within the Azure ...

Senior Software Engineer – Agentic Development Enablement

Location
Greater London, England, United Kingdom
such as Claude Code and GitHub Copilot Design and implement guardrails, controls, and engineering patterns for AI-assisted development Contribute to endpoint and platform observability, telemetry, and policy enforcement Define how controls work consistently across local development environments and CI/CD pipelines Explore changes to development environments, including containerised … production Requirements Strong hands-on background as a software engineer Broad technical understanding across developer tooling, cloud platforms, operating systems, desktop environments, security controls, observability, telemetry, and CI/CD Experience working on developer workflows and engineering ways of working, not only end-user application delivery Ability to work ...

Site Reliability Engineer with Python

Hiring Organisation
BC Forward
Location
Pennington, New Jersey, United States
Employment Type
Permanent
Salary
USD Hourly
Global Markets. The ideal candidate will have strong experience in Python, Django, REST APIs, MySQL, Linux, Infrastructure-as-Code, CI/CD, and observability and a proven ability to increase platform reliability, improve observability, automate operations, and lead incident response at scale. Responsibilities: Own reliability and operational health of core … Build, enhance, and maintain internal tools that support platform operations. Lead incident response, root-cause analysis, and post-incident improvements. Improve application and platform observability using Dynatrace and OpenTelemetry. Partner with Quartz core teams on platform modernization and cloud migrations. Harden production systems and evolve support tooling for large-scale ...

Director, Site Reliability Engineering

Location
Manchester, England, United Kingdom
engineered into every service throughout its lifecycle. Working alongside the Director of Site Reliability Operations, this leader will define the engineering standards, automation, observability, production readiness, and resilience capabilities that enable world‐class operational performance. While Site Reliability Operations owns the day‐to‐day operation of production services, the Site … global Site Reliability Engineering organization. This role owns the engineering strategy, governance, architecture, and technical practices that improve service reliability through software engineering, observability, automation, resilience engineering, production engineering, and operational readiness. Rather than operating production systems on a day‐to‐day basis, the Site Reliability Engineering organization develops ...

Front End developer

Location
Greater London, England, United Kingdom
applications. You’ll work across Python services and React/Vite frontends , collaborating closely with engineering teams to improve application reliability, CI/CD, observability, security and developer tooling. Key experience: Docker and CI/CD pipelines Azure cloud experience is essential; AWS is a bonus Good knowledge of Linux … networking, security and observability Experience supporting scalable application platforms and cloud deployments Able to work independently and establish reusable engineering patterns and best practices A great opportunity for someone who enjoys working across development, cloud and platform engineering. Additional Information #TalanUK #J-18808-Ljbffr ...

Cloud / Infrastructure Engineer

Location
Whiteley, England, United Kingdom
Partner with application engineering teams to improve deployment workflows and reduce operational risk Operational Readiness and Collaboration Integrate infrastructure with monitoring, logging, alerting, and observability systems Support production operations and participate in incident response activities related to infrastructure and deployments Collaborate with backend and data engineering teams to enable scalable … reliable service operation Develop and maintain operational documentation, procedures, and support runbooks Utilize observability data to identify issues and drive corrective actions Required Qualifications 5+ years of experience in cloud, infrastructure, or DevOps engineering roles Hands-on ownership of AWS infrastructure in production environments Strong experience with Infrastructure-as-Code ...

Agent Engineer

Location
Greater London, England, United Kingdom
users to create user-friendly solutions for complex business processes Participate in technical design reviews, planning sessions, and code reviews Contribute to infrastructure and observability practices with the Engineering team Continuously improve quality, reliability, and usability across internal platforms Support Elliptic's mission to make crypto markets safer, more transparent … framework experience such as NestJS or Express is nice to have Terraform or infrastructure-as-code experience is nice to have Datadog or similar observability platform experience is nice to have DynamoDB or other NoSQL database experience at scale is nice to have Distributed or event-driven architecture experience, including ...

Principal Software Engineer - Full Stack - AI

Location
York and North Yorkshire, England, United Kingdom
full stack, guiding the development of responsive frontend applications (React/TypeScript) and robust, scalable backend services (Python, Java, Kotlin, or Node.js). LLM Observability & Reliability: Establish robust LLM observability, evaluations, and caching, implementing latency optimisations and comprehensive monitoring (logging, usage tracking, agent behaviour). Operational Excellence: Champion high availability … performance optimisation, and observability across frontends and backend microservices, focusing on practices that maintain platform reliability and optimise MTTD and MTTR. Mentorship & Collaboration: Elevate the engineering organisation by mentoring senior and junior engineers, conducting rigorous code and system design reviews, and partnering with product managers to translate product visions into ...

Azure Systems Engineer

Location
Greater London, England, United Kingdom
integrations and production environments Implementing Infrastructure as Code (IaC) and automation solutions Supporting identity, access management, security and governance across Azure Improving platform monitoring, observability, resilience and performance Supporting SQL Server and Azure SQL environments, including access, backups and troubleshooting Contributing to vulnerability management, incident response and operational security Strong … Entra ID, IAM, RBAC and privileged access Understanding of secure application delivery, configuration and secrets management Experience with Azure Monitor, Application Insights or equivalent observability tooling Knowledge of SQL Server/Azure SQL , including access management, backups and troubleshooting Strong understanding of TCP/IP, DNS, routing, firewalls and private ...