2,401 to 2,425 of 4,597 Observability Jobs

Cloud Native SRE Engineer | Kubernetes & IaC | Hybrid

Location
Sheffield, England, United Kingdom
Bumper is seeking a DevOps/SRE Engineer to join our Hybrid Sheffield team. You’ll help shape reliability, scalability and observability of our cloud-native platform, building, deploying and monitoring systems across AWS/Azure. You’ll collaborate with engineers, lead incident response, and drive automation to reduce toil ...

Software Engineer, Backend (BI Platform)

Location
Greater London, England, United Kingdom
infrastructure our BI platforms run on. This is a hands-on software and platform engineering role: services and integrations, identity and access automation, observability, reliability and developer experience. BI is the domain you serve (primarily Looker and Sigma, and the systems around them), but the problems are backend and platform … Build the integrations between BI platforms and the wider ecosystem: warehouses, identity providers, orchestration, CI/CD and internal developer tooling Improve the availability, observability, performance, scalability and security of the systems you own, and carry them in production Use AI tooling as part of how you engineer, and where ...

Senior Lead Data Platform Engineer - Java/Python

Location
Greater London, England, United Kingdom
develop secure, scalable production code in Java, Python, or other modern languages, and collaborate with agile teams to advance cloud-native data platforms and observability practices across the organization. #J-18808-Ljbffr ...

AWS Cloud DevOps Engineer - Automation & CI/CD

Location
Manchester, England, United Kingdom
based infrastructure across hybrid UK environments. You will implement IaC with Terraform or CloudFormation, automate CI/CD pipelines and monitor performance with observability best practices. In this role you’ll work in an Agile setup, collaborate with stakeholders, and contribute to securing cloud workloads while continually improving deployment reliability. ...

Remote Platform Engineer: AI, Cloud & CI/CD Automation

Location
Greater London, England, United Kingdom
design and operate pipelines, IaC, and cloud services across AWS/Azure and Kubernetes, aligning with GitOps and security best practices. You will implement observability, self‐service tooling, and automation while collaborating with developers to ship reliable software and AI workloads at scale. #J-18808-Ljbffr ...

AWS Cloud DevOps Engineer - Automation & CI/CD

Location
West of England, England, United Kingdom
based infrastructure across hybrid UK environments. You will implement IaC with Terraform or CloudFormation, automate CI/CD pipelines and monitor performance with observability best practices. In this role you’ll work in an Agile setup, collaborate with stakeholders, and contribute to securing cloud workloads while continually improving deployment reliability. ...

Senior DevOps Engineer: Kubernetes & Python

Location
Greater London, England, United Kingdom
/OpenShift, architect container-based microservices, and enforce security and compliance across the stack. Expect hands-on work with GitOps, CI/CD, and observability tools to maintain high performance in production. The role emphasizes troubleshooting complex container workloads, optimizing resilience, and collaborating with #J-18808-Ljbffr ...

Director, Cloud Infrastructure

Location
Greater London, England, United Kingdom
Cloud Infrastructure to lead that work. This person will own the platform foundations Sanity engineers build on every day: cloud infrastructure, Kubernetes, networking, routing, observability, CI/CD, deployment paths, incident response, and the standards that make production ownership work across product teams. The scale is real. Content Lake alone … edge, gateway, caching, object storage, and GCP infrastructure. Some of the work is already in motion: moving Varnish and Mead onto Fastly, tightening observability, finishing our developer on‐call rollout, and making production readiness a normal part of shipping. This is a leadership role for someone who can still ...

Senior Network SRE: Cloud Reliability & IaC

Location
Greater London, England, United Kingdom
focus on cloud automation, IaC, and governance across our AWS infra, contributing to highly available services for millions of users. You will own automation, observability, and incident response, driving cross-functional projects in an Agile environment. Strong Linux skills, Python/Bash, and Terraform expertise are essential for success. #J ...

SRE Technical Lead - SC Cleared

Hiring Organisation
F5 consultants
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
error budgets and reliability standards Provide technical leadership across complex Kubernetes and OpenShift environments Drive automation across monitoring, incident response, recovery and remediation Shape observability, capacity management and operational readiness practices Act as a senior technical escalation point for major incidents and high-risk releases Lead reliability reviews, toil reduction … OpenShift Proven SRE/reliability engineering experience within large-scale production environments Experience defining or working with SLAs, SLOs and error budgets Strong observability experience with tools such as Prometheus, Grafana, Loki, Tempo or OpenTelemetry Strong Infrastructure as Code and GitOps experience - ideally Helm, Kustomize, ArgoCD and/or Tekton ...

Global Cloud Networking & Security Architect

Location
Greater London, England, United Kingdom
multi-cloud strategy, implement Zero Trust, and build automation platforms using Terraform, Python, APIs, and CI/CD pipelines, advancing AI-enabled operations and observability across the #J-18808-Ljbffr ...

AI Platform & Site Reliability Engineering Consultant

Hiring Organisation
Akkodis
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£88000 - £96000/annum
include: Define and embed SRE engagement models aligned to modern engineering and traditional ITSM/ITIL practices Establish SLIs, SLOs, and Error Budgets Shape observability strategies using metrics, logs, and traces Design incident response models and post-incident learning loops Reduce toil through automation and engineering excellence Deliver SRE capability … Looking For Extensive experience in SRE, cloud operations, or DevOps Proven consulting or advisory background Experience with AWS, Azure, or GCP Strong observability and incident management expertise Ability to obtain UK SC clearance Modis International Ltd acts as an employment agency for permanent recruitment and an employment business ...

Sr. DevOps Engineer - OpenShift & Azure Automation Lead (Hybrid)

Location
East Midlands, England, United Kingdom
governance, design secure platform automation, and drive CI/CD pipelines across OpenShift, Kubernetes, and Azure. Several enterprise-scale initiatives require robust automation and observability expertise. BPSS is required to start; SC clearance can be processed during onboarding. Competitive market rates offered. #J-18808-Ljbffr ...

Senior Engineering Manager, Developer Experience

Location
Greater London, England, United Kingdom
About the team The Developer Experience team owns the internal platform that every the company engineer touches daily: CI/CD pipelines, observability tooling, our developer portal, and an emerging AI platform. It's a high-visibility role: the work you lead directly shapes the productivity of hundreds of engineers … looks like at the company as we scale. What you'll do Lead and develop a growing team of 5+ highly motivated engineers across observability, CI/CD, developer portal (Backstage), and FinOps tooling — setting clear priorities and establishing strong ways of working. Own and evolve the technical roadmap across ...

Senior Vice President, Full-Stack Engineer

Location
Manchester, England, United Kingdom
underlying workflow engine (e.g., Camunda) to enable extensibility, portability, and enterprise-scale orchestration. Drive delivery excellence across workflow and decisioning platforms, embedding observability, resilience, auditability, and performance at scale, while establishing engineering standards across CI/CD, testing, security, and data architecture. Qualifications Bachelor’s or Master’s degree … platform design and optimisation. Proven track record of delivering production-grade platforms, embedding engineering excellence across test automation, CI/CD, observability, resilience, and traceability while driving continuous improvement of SDLC practices at scale. Hands-on technical leader who can actively contribute to solution design and critical builds, while defining ...

Senior Backend Engineer - Databases - Analytics | UK | Remote

Location
Pathhead, Scotland, United Kingdom
United Kingdom (Remote) Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimise telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud … brings AI to observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. ...

Database Reliability Engineer

Location
Southampton, England, United Kingdom
tune performance across hundreds of instances. Architect Cross‐Cloud Portability: use CNPG and cloud‐native patterns to keep our database layer provider‐agnostic. Evolve Observability & Monitoring: build proactive monitoring and alerting to detect regressions before they affect customers. Support Replication & Mobility: enable data streaming and zero‐downtime migration strategies … provision infrastructure, avoiding manual implementations. Distributed Systems enthusiast: enjoy the challenge of multi‐tenant, multi‐region, multi‐cloud scenarios with rigorous data integrity. Security & Observability mindset: build deep observability (Prometheus/Grafana/OpenTelemetry/Humio) and guardrails for secure operation. Engineering via code: deliver backend services in Java with ...

NOC Engineer

Location
West of England, England, United Kingdom
back online with speed and precision Writing automation that removes repetitive manual work and makes the platform more robust Enhancing the monitoring, alerting and observability tooling used across our cloud estate Partnering with software, platform, cloud and security teams to raise the bar on reliability and operational standards Taking part … Working with AWS infrastructure Running workloads on Terraform and Docker Supporting live production systems and handling incidents Scripting in Python, Bash or Go Using observability and monitoring tools like Grafana, Prometheus, Datadog, Splunk or CloudWatch A solid grasp of core networking concepts, including DNS, TCP/IP and load balancing ...

Lead AI Engineer IRC302971

Location
Greater London, England, United Kingdom
clients transform their digital platforms and products. As an AI Engineer, you’ll own the development of the core agentic frameworks, evaluation pipelines, and observability tooling that let it operate with safety, trust, and intelligence at scale. You won’t just be integrating AI into a product … CrewAI is a plus. Production Python engineering : Clean, modular, testable, maintainable code — you care about system reliability as much as model output. Evaluation & observability : Practical experience instrumenting tracing (LangSmith, Arize) and building CI/CD pipelines built specifically for LLMs. Engineering foundations : Distributed systems, scalable data pipelines (Kafka ...

ai engineer for an AI platform

Location
Greater London, England, United Kingdom
latency monitoring, alerting, scaling considerations and operational readiness Apply CI/CD and software engineering best practices to AI platform and agentic components Embed observability through logs, metrics and traces Partner with Cloud Infrastructure and Security teams to design secure, scalable and cost-effective Azure environments Требования Significant experience … APIs and service-oriented systems Experience with CI/CD pipelines, automated testing and versioned deployments for AI or platform components Practical experience with observability tooling, including logging, metrics, tracing and alerting Experience using telemetry to improve reliability and performance Comfortable working in cloud environments Strong collaboration skills ...

Infrastructure & Developer Platform Lead

Location
Greater London, England, United Kingdom
cloud architecture strategies into implementation plans and engineering roadmaps. Build and maintain reliable, secure and self‐service engineering environments.Drive automation across provisioning, deployment, testing, observability and operational processes. Ensure infrastructure meets agreed requirements for security, operability, compliance and resilience.Improve developer productivity through reusable platform services, engineering tooling and automation. Support … native software delivery through platform capabilities, governance, automation and operational controls.Drive reliability, security, observability and operational readiness across platform services. Manage technical risks, infrastructure incidents, dependencies and operational blockers.Provide technical leadership, mentoring and engineering guidance while remaining actively involved in platform engineering decisions. Build platform capability, operational knowledge and engineering ...

NOC Engineer

Hiring Organisation
Spectrum IT Recruitment
Location
Bristol, United Kingdom
Employment Type
Permanent
Salary
£50000/annum
back online with speed and precision Writing automation that removes repetitive manual work and makes the platform more robust Enhancing the monitoring, alerting and observability tooling used across our cloud estate Partnering with software, platform, cloud and security teams to raise the bar on reliability and operational standards Taking part … Working with AWS infrastructure Running workloads on Terraform and Docker Supporting live production systems and handling incidents Scripting in Python, Bash or Go Using observability and monitoring tools like Grafana, Prometheus, Datadog, Splunk or CloudWatch A solid grasp of core networking concepts, including DNS, TCP/IP and load balancing ...

NOC Engineer

Hiring Organisation
Spectrum IT Recruitment
Location
Basingstoke, Hampshire, United Kingdom
Employment Type
Permanent
Salary
£50000/annum Bonus, Pension, Healthcare
back online with speed and precision Writing automation that removes repetitive manual work and makes the platform more robust Enhancing the monitoring, alerting and observability tooling used across our cloud estate Partnering with software, platform, cloud and security teams to raise the bar on reliability and operational standards Taking part … Working with AWS infrastructure Running workloads on Terraform and Docker Supporting live production systems and handling incidents Scripting in Python, Bash or Go Using observability and monitoring tools like Grafana, Prometheus, Datadog, Splunk or CloudWatch A solid grasp of core networking concepts, including DNS, TCP/IP and load balancing ...

Software Engineer III - Fullstack (Java, React, Python and AI) Engineer

Location
Greater London, England, United Kingdom
with product, marketing, and partners to translate requirements into well-designed technical solutions Improve engineering excellence through code reviews, test automation, CI/CD, observability, performance tuning, and operational best practices Ensure solutions meet security, privacy, and compliance expectations including consent management, data minimization, and access controls Required Qualifications, Capabilities … with campaign management and attribution systems Background in AI-enabled content generation, experimentation, or workflow automation Skills in test automation, CI/CD, and observability tools Experience optimizing performance and scalability Familiarity with consent management and data minimization practices Ability to drive innovation in personalization and measurement Experience working ...