2,376 to 2,400 of 5,534 Observability Jobs

Director of Platform and Reliability

Location
Greater London, England, United Kingdom
business continuity and incident readiness, including recovery objectives and resilience testing Design, build and modernise production infrastructure Champion Terraform and Infrastructure as Code Improve observability, service ownership, SLOs, incident response, on-call practices and reliability engineering Improve CI/CD, deployment workflows, developer environments and internal platform tooling Partner with … away from the technology. You’ll shape the long-term platform, reliability, security and operational-resilience whilst taking on complex challenges including legacy modernisation, observability, cloud cost and third-party risk. Your work will have a direct impact on a business helping more people achieve homeownership, while giving you significant ...

Secure Cloud Platform Engineer | DevSecOps & SRE

Location
Cheltenham, England, United Kingdom
national security. Join agile, multi-disciplinary teams focused on CI/CD, Infrastructure as Code and live service support, with emphasis on security, observability and incident readiness to ensure #J-18808-Ljbffr ...

Secure Cloud Platform Engineer | DevSecOps & SRE

Location
Manchester, England, United Kingdom
national security. Join agile, multi-disciplinary teams focused on CI/CD, Infrastructure as Code and live service support, with emphasis on security, observability and incident readiness to ensure #J-18808-Ljbffr ...

Remote Developer Support Engineer (London)

Hiring Organisation
Braintrust
Location
Paisley, Scotland, United Kingdom
About the company Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare … Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve. About the Role At Braintrust, surprisingly good developer support is one of our most important strategic advantages. Our customers are developers building LLM-powered applications, and they move fast. ...

Remote Developer Support Engineer (London)

Location
Caernarfon, Caernarfonshire, United Kingdom
About the company Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare … Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve. About the Role At Braintrust, surprisingly good developer support is one of our most important strategic advantages. Our customers are developers building LLM-powered applications, and they move fast. ...

Remote Developer Support Engineer (London)

Location
Glasgow, Lanarkshire, United Kingdom
About the company Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare … Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve. About the Role At Braintrust, surprisingly good developer support is one of our most important strategic advantages. Our customers are developers building LLM-powered applications, and they move fast. ...

Remote Developer Support Engineer (London)

Location
Falkirk, Stirlingshire, United Kingdom
About the company Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare … Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve. About the Role At Braintrust, surprisingly good developer support is one of our most important strategic advantages. Our customers are developers building LLM-powered applications, and they move fast. ...

Remote Developer Support Engineer (London)

Location
Mauchline, Ayrshire, United Kingdom
About the company Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare … Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve. About the Role At Braintrust, surprisingly good developer support is one of our most important strategic advantages. Our customers are developers building LLM-powered applications, and they move fast. ...

Remote Developer Support Engineer (London)

Location
Cramlington, Northumberland, United Kingdom
About the company Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare … Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve. About the Role At Braintrust, surprisingly good developer support is one of our most important strategic advantages. Our customers are developers building LLM-powered applications, and they move fast. ...

Remote Developer Support Engineer (London)

Location
Llanidloes, Radnorshire, United Kingdom
About the company Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare … Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve. About the Role At Braintrust, surprisingly good developer support is one of our most important strategic advantages. Our customers are developers building LLM-powered applications, and they move fast. ...

Remote Developer Support Engineer (London)

Location
Bedford, Bedfordshire, United Kingdom
About the company Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare … Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve. About the Role At Braintrust, surprisingly good developer support is one of our most important strategic advantages. Our customers are developers building LLM-powered applications, and they move fast. ...

Remote Developer Support Engineer (London)

Location
Stonehaven, Aberdeenshire, United Kingdom
About the company Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare … Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve. About the Role At Braintrust, surprisingly good developer support is one of our most important strategic advantages. Our customers are developers building LLM-powered applications, and they move fast. ...

Remote Developer Support Engineer (London)

Location
Montrose, Angus, United Kingdom
About the company Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare … Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve. About the Role At Braintrust, surprisingly good developer support is one of our most important strategic advantages. Our customers are developers building LLM-powered applications, and they move fast. ...

Remote Developer Support Engineer (London)

Location
Lyndhurst, Hampshire, United Kingdom
About the company Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare … Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve. About the Role At Braintrust, surprisingly good developer support is one of our most important strategic advantages. Our customers are developers building LLM-powered applications, and they move fast. ...

Infrastructure Support Engineer

Hiring Organisation
Blues Point Ltd
Location
Shoreham-by-Sea, West Sussex, United Kingdom
Employment Type
Full-Time
Salary
£50,000 - £55,000 per annum
Manage Active Directory, SSL certificates and secure communications Support networking across AWS, including VPCs, IGWs, NAT Gateways and security groups Contribute to monitoring and observability using Grafana Support infrastructure maintenance, migrations, patching, audit and compliance activity Work collaboratively within an Agile/sprint-based environment There will be occasional … Code experience Experience with containers and orchestration PostgreSQL experience Networking and load-balancing knowledge Experience with SSL certificates and endpoint management Grafana/observability experience Good communication skills and the ability to explain technical issues to non-technical audiences A proactive approach and the ability to work independently ...

Vice President, Production Services Application Support

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
urgency, recover priority incidents under pressure, and maintain core support coverage across on-site and offshore support hours. Use SQL scripting, automation, monitoring, and observability tools to improve operational resilience, service health, reliability, and incident response. To be successful in this role, were seeking the following: Excellent SQL scripting skills. … solutions for alert correlation, anomaly detection, predictive monitoring, and service optimisation. Strong understanding of Site Reliability Engineering (SRE) principles, including service health, reliability, availability, observability, incident reduction, and continuous service improvement. Experience with SRE practices such as monitoring and alert tuning, incident management, post-incident reviews, root cause analysis ...

Secure Cloud Platform Engineer | DevSecOps & SRE

Location
Greater London, England, United Kingdom
national security. Join agile, multi-disciplinary teams focused on CI/CD, Infrastructure as Code and live service support, with emphasis on security, observability and incident readiness to ensure #J-18808-Ljbffr ...

Engineering Architect

Location
Bath, England, United Kingdom
will own and evolve the technical foundations that underpin product and service delivery, including engineering standards, development frameworks, CI/CD pipelines, testing harnesses, observability, and operational practices. Your focus will be on enabling a fast-moving, AI-first delivery team to consistently deliver high-quality, scalable, and supportable solutions … varied role, you will: own and version technical foundations across all seven harness dimensions (Security, performance, coding style, UI/UX baselines, logging and observability, testing) so that every deliverable inherits secure, observable and tested defaults build and maintain the shared ADO CI/CD pipelines and release governance ...

Senior AI Engineer (AI Platform)

Location
Greater London, England, United Kingdom
will also contribute to the production foundations needed to operate AI capabilities reliably, including LLMOps, model access patterns, prompt and agent lifecycle practices, evaluation, observability and secure enterprise integration. This is a hands‐on engineering role where the capabilities you build will be used by other teams across ASOS, helping … latency monitoring, alerting, scaling considerations and operational readiness Applying CI/CD and software engineering best practices to AI platform and agentic components Embedding observability by default, ensuring AI systems are measurable, debuggable and auditable through logs, metrics and traces Working with Cloud Infrastructure and Security teams to design secure ...

Senior DevOps Engineer

Location
Greater London, England, United Kingdom
Senior DevOps Engineer to support the team that keeps ThreatAware's platform running and shipping. You'll own AWS infrastructure, CI/CD pipelines, observability, and platform reliability — and define the developer workflow end-to-end, pioneering how AI can streamline every step. You report to the CTO. … cost optimisation. You know this space and you care about getting it right. Define the developer workflow and keep improving it. Build the tooling, observability, and automation that let engineers focus on code, not friction. Pioneer AI-assisted DevOps — explore and implement ways AI can streamline builds, deployments, monitoring ...

Oracle BSS Stack Lead

Location
Newbury, England, United Kingdom
Management controls. Drive automation across deployment, monitoring and operational processes. Implement continuous integration and continuous delivery capabilities using Jenkins, GitHub and Azure DevOps. Define observability standards using Dynatrace, Splunk and other enterprise monitoring tools. Reduce repetitive operational activity through sustainable automation. Build trusted working relationships with customer stakeholders, business teams … clearly to technical teams, business stakeholders and senior decision-makers. Experience with cloud technologies, containers and Kubernetes would be beneficial. Knowledge of monitoring and observability platforms would be beneficial. Familiarity with Oracle OSS components would be advantageous but is not essential. Not a perfect fit? Concerned you may not meet ...

Software Engineer, GPU Infrastructure- ChatGPT Engineering

Location
Greater London, England, United Kingdom
large-scale GPU infrastructure supporting ChatGPT inference. Build internal platforms, tooling, and AI-powered agents that automate fleet operations and reduce operational overhead. Improve observability, reliability, and operational efficiency across thousands of GPUs. Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response. Identify infrastructure bottlenecks and implement … software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure. Understand infrastructure observability, monitoring, capacity planning, and incident management. Enjoy identifying cross-team pain points and building reusable platforms that improve developer productivity. Are comfortable working across software ...

VF3 VOIS Managed Service Operations Lead

Location
Newbury, England, United Kingdom
with Change Management controls. Drive automation across deployment, monitoring and operational processes. Support CI/CD adoption using Jenkins, GitHub and Azure DevOps. Establish observability standards using Dynatrace, Splunk and other enterprise monitoring platforms. Reduce repetitive operational work through sustainable automation and process improvement. Build trusted relationships with senior client … improvements. Strong commercial awareness with the ability to balance customer outcomes, operational performance, and contractual commitments. Knowledge of DevOps, Site Reliability Engineering (SRE), automation, observability, and modern operational practices. Experience working with third-party vendors and managing complex stakeholder ecosystems across multiple functions. Desirable experience with Oracle BSS, OSS, integration ...

Senior Infrastructure Platform Engineer Veeam / VMware / Hyper-V

Hiring Organisation
100% IT Recruitment Ltd
Location
United Kingdom
Employment Type
Permanent, Work From Home
Salary
£60,000
major incident recovery, particularly around backup, restore and platform availability. Drive automation to reduce manual processes and improve operational efficiency. Develop and maintain monitoring, observability and alerting using tools such as Grafana. Coordinate the day-to-day priorities of the Platform Engineering team. Maintain engineering standards, documentation and operational procedures. … Azure DevOps-focused role. Experience supporting highly available production infrastructure. Strong troubleshooting and problem-solving skills across enterprise infrastructure. Experience with monitoring and observability platforms such as Grafana. Experience automating operational tasks using PowerShell, scripting or similar technologies. Excellent understanding of backup, disaster recovery and platform resilience. Ability to coordinate ...

SRE Technical Lead

Hiring Organisation
83zero Limited
Location
Wokingham, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
senior technical escalation point for major incidents and high-risk releases. Lead blameless post-incident reviews and ensure measurable service improvements. Define and oversee observability, monitoring and capacity management practices. Ensure SRE approaches align with security, governance and compliance requirements. Mentor and coach senior engineers, helping to improve SRE maturity … OpenShift. Experience designing and supporting hybrid and multi-cloud platforms. Experience with service mesh technologies such as Istio. Strong hands-on experience with observability tooling including Prometheus, Grafana, Loki, Tempo and OpenTelemetry. Infrastructure as Code and GitOps expertise using tools such as Helm, Kustomize, ArgoCD and Tekton. Experience building ...