2,251 to 2,275 of 4,614 Observability Jobs

Non-Functional Test Specialist

Location
United Kingdom
Hands-on experience in planning, preparing & executing Resilience/Disaster Recovery/Continuity/Failover testing/Performance Exposure to tools supporting resilience & operational observability (Splunk, Prometheus, Kafka, Chaos tooling etc) Understanding of high-availability architectures, infrastructure redundancy & backup/restore strategies Are familiar with working within large scale & complex ...

Tech Lead

Location
Manchester, England, United Kingdom
engineering culture Improve delivery flow, remove blockers and help the team ship value quickly Champion engineering best practices across testing, CI/CD, observability and operational excellence Work closely with Product, Design and other business stakeholders to turn priorities into successful outcomes Encourage pragmatic adoption of AI tools to improve ...

Sr. Engineer - Platform

Location
Greater London, England, United Kingdom
Your Role We’re hiring a Senior Engineer to join our Platform Engineering team, working with squads dedicated to site reliability, cloud infrastructure, observability and production operations. You’ll help drive operational consistency and ensure that Hudl engineers can build on a highly available, scalable and secure platform. ...

AI Principal Architect

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
agents, orchestration frameworks, or enterprise AI platforms. Deep understanding of AI architecture considerations, including data readiness, model selection, security, responsible AI, governance, integration, scalability, observability, and operations. Experience leading architecture for complex client pursuits, large transformation programs, or multi-tower technology solutions. Bachelor's degree or equivalent work experience. Preferred ...

Software Engineer - Front End - Studio (Core)

Location
Greater London, England, United Kingdom
back-end leads to shape roadmaps and turn ambiguous problems into clear technical plans. Own the operational excellence of your domains, ensuring reliability, observability, and a high SLA. Clearly explain the nuances of system design and front-end paradigms to engineers and stakeholders alike. The Team You will join ...

Customs AI Integration & Automation Engineer (UK)

Location
United Kingdom
allowed to do. Agent orchestration frameworks and standards (LangGraph, MCP, Temporal or equivalent), and experience with long-running or multi-step agent workflows. Observability and cost control for inference at volume. What success looks like in your first year By month three: You understand our customs processes well enough ...

Remote Principal Software Engineer

Location
United Kingdom
live and breathe this approach ourselves: we release new versions of Gearset multiple times a day and we continually invest in improving our own observability and infrastructure tools. This means we can identify and react to issues quickly and delight our users by getting improvements to them as fast ...

Lead Site Reliability Engineer, Athena Core

Location
Greater London, England, United Kingdom
appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations. A strong understanding of Linux, Proficient knowledge and experience in observability such as white and black box monitoring, service level objective alerting, and telemetry collection Proficient with continuous integration and continuous delivery practices and tooling Proficient ...

Customs AI Integration & Automation Engineer (UK)

Location
Felixstowe, England, United Kingdom
allowed to do. Agent orchestration frameworks and standards (LangGraph, MCP, Temporal or equivalent), and experience with long-running or multi-step agent workflows. Observability and cost control for inference at volume. What success looks like in your first year By month three: You understand our customs processes well enough ...

Remote Principal Software Engineer

Location
Ipswich, Suffolk, United Kingdom
live and breathe this approach ourselves: we release new versions of Gearset multiple times a day and we continually invest in improving our own observability and infrastructure tools. This means we can identify and react to issues quickly and delight our users by getting improvements to them as fast ...

Remote Principal Software Engineer

Location
Glasgow, Lanarkshire, United Kingdom
live and breathe this approach ourselves: we release new versions of Gearset multiple times a day and we continually invest in improving our own observability and infrastructure tools. This means we can identify and react to issues quickly and delight our users by getting improvements to them as fast ...

Remote Principal Software Engineer

Location
Worksop, Nottinghamshire, United Kingdom
live and breathe this approach ourselves: we release new versions of Gearset multiple times a day and we continually invest in improving our own observability and infrastructure tools. This means we can identify and react to issues quickly and delight our users by getting improvements to them as fast ...

Remote Principal Software Engineer

Location
Bristol, Gloucestershire, United Kingdom
live and breathe this approach ourselves: we release new versions of Gearset multiple times a day and we continually invest in improving our own observability and infrastructure tools. This means we can identify and react to issues quickly and delight our users by getting improvements to them as fast ...

Remote Principal Software Engineer

Location
New Milton, Hampshire, United Kingdom
live and breathe this approach ourselves: we release new versions of Gearset multiple times a day and we continually invest in improving our own observability and infrastructure tools. This means we can identify and react to issues quickly and delight our users by getting improvements to them as fast ...

Remote Principal Software Engineer

Location
Dunfermline, Fife, United Kingdom
live and breathe this approach ourselves: we release new versions of Gearset multiple times a day and we continually invest in improving our own observability and infrastructure tools. This means we can identify and react to issues quickly and delight our users by getting improvements to them as fast ...

Remote Principal Software Engineer

Location
Bracknell, Berkshire, United Kingdom
live and breathe this approach ourselves: we release new versions of Gearset multiple times a day and we continually invest in improving our own observability and infrastructure tools. This means we can identify and react to issues quickly and delight our users by getting improvements to them as fast ...

Remote Principal Software Engineer

Location
Beverley, East Yorkshire, United Kingdom
live and breathe this approach ourselves: we release new versions of Gearset multiple times a day and we continually invest in improving our own observability and infrastructure tools. This means we can identify and react to issues quickly and delight our users by getting improvements to them as fast ...

Remote Principal Software Engineer

Location
Stone, Staffordshire, United Kingdom
live and breathe this approach ourselves: we release new versions of Gearset multiple times a day and we continually invest in improving our own observability and infrastructure tools. This means we can identify and react to issues quickly and delight our users by getting improvements to them as fast ...

Software Engineer (Backend)

Location
Greater London, England, United Kingdom
Rust work or demonstrable learning. We don't expect you to know our exact stack Solid backend fundamentals: concurrency, data modelling, transactions, reliability and observability Experience designing and consuming APIs (REST/gRPC/GraphQL) and integrating with services and data stores A strong testing mindset and familiarity with ...

Remote Principal Software Engineer

Location
Tain, Inverness-shire, United Kingdom
live and breathe this approach ourselves: we release new versions of Gearset multiple times a day and we continually invest in improving our own observability and infrastructure tools. This means we can identify and react to issues quickly and delight our users by getting improvements to them as fast ...

Senior DevOps Engineer

Location
West End, England, United Kingdom
leads continuous improvement initiatives that improve delivery flow, reduce toil, and strengthen developer experience. Advises on tooling, methodologies, and engineering processes. Designs robust observability frameworks using logs, metrics, tracing, and actionable alerting to ensure cross‐system visibility. Guides others in troubleshooting complex, multi‐layer issues, drives improvements for reliability ...

Solutions Architect - Data Enablement

Hiring Organisation
Conferma Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
architecture governance. Design cloud-native solutions on Microsoft Azure, balancing performance, resilience, scalability and cost. Define and validate non-functional requirements, including security, observability and reliability. Drive platform modernisation through cloud-native technologies, automation, AI-assisted engineering and API-first, event-driven architecture. Produce and maintain architecture artefacts, including ADRs ...

Senior Software Engineer

Location
Cramlington, England, United Kingdom
priorities Confident, senior-level communicator able to work autonomously and influence without authority Experience mentoring or supporting the development of other engineers Familiarity with observability tooling (e.g. Grafana, Prometheus/Mimir) is a plus Pragmatic approach to trade-offs between speed, technical debt, and long-term maintainability A genuine self ...

AI Platform Lead (Director)

Location
Greater London, England, United Kingdom
platforms such as Databricks, SageMaker, or MLflow Exposure to hybrid cloud architectures and large scale data processing ecosystems Knowledge of enterprise data platforms, observability, and cost management tooling Leadership Expectations Set technical direction and influence enterprise wide platform strategy and adoption Lead high performing engineering teams and foster a culture ...

Site Reliability Engineer (SRE) - Glasgow, UK

Location
Glasgow, Scotland, United Kingdom
infrastructure* Ensure high system availability performance scalability and reliability* Participate in incident management root cause analysis and problem resolution* Implement proactive monitoring alerting and observability solutions* Reduce operational overhead through automation and selfhealing mechanisms* Support production releases and deployment activitiesAWS Cloud Engineering* Design deploy and manage AWSbased infrastructure and services ...