2,901 to 2,925 of 4,317 Permanent Observability Jobs

Site Reliability Engineer - Hybrid, Observability

Location
United Kingdom
bet365 Group is seeking a Site Reliability Engineer to strengthen the stability, performance and resilience of the systems behind every click and live change across our global product. In this full-time role, you will ...

Site Reliability Engineer I: Cloud-Native & Observability

Location
Greater London, England, United Kingdom
Axon is seeking an experienced Site Reliability Engineer for its Real Time Operations in London. You will contribute to building reliable cloud-native services, collaborate with RTO engineering teams, and enable product teams to scale ...

Infrastructure Software Engineer — Cloud & Observability

Location
Greater London, England, United Kingdom
A leading fintech company in the United Kingdom is seeking a Software Engineer specialized in Infrastructure to support deployment and maintenance of cutting-edge software. The role involves building reliable and scalable applications, developing tools ...

Site Reliability Engineer – Cloud Reliability & Observability

Location
Fenny Stratford, England, United Kingdom
Connells Group UK is seeking an experienced Site Reliability Engineer (SRE) to join the Group Technology Team in Milton Keynes. The role focuses on building and operating ConnellsX, our Azure-based internal developer platform, ensuring ...

Software Engineer II, AI Enablement Team

Location
Greater London, England, United Kingdom
Build cost observability and usage-tracking tooling across LLM and agent workloads Design and ship internal tooling and guardrails for safe, productive agentic coding workflows Evaluate AI tools, models, and vendors and produce build-vs-buy recommendations Write and maintain production-quality code and automated tests Review peers' code … particular technology stack or language Clear written and verbal communication Generalist attitude toward learning and growth Experience with LLM/agentic tooling, AI cost observability, or internal developer platforms is a bonus Core Competencies Demonstrates expertise in building cost observability and usage-tracking tooling for AI workloads, with a strong ...

Senior QA Engineer

Hiring Organisation
SRG
Location
Warrington, Cheshire, United Kingdom
Employment Type
Full-Time
Salary
£45,000 - £50,000 per annum
both fast and reliable. You'll help move quality earlier into the process (shift-left), while also using real production insights to improve decisions (observability-led quality). There's strong scope to influence how QA operates within the team, from testing strategy through to continuous improvement. What … considered early in design and development Leading exploratory testing to uncover issues real users might experience Using monitoring, metrics, and logs to drive observability-led quality and improve production outcomes Identifying risks early and helping the team make informed decisions Improving QA processes, standards, and ways of working across ...

Senior Infrastructure & Operations Engineer (Kubernetes / Platform Reliability)

Location
Greater London, England, United Kingdom
production systems at scale and who focuses on making infrastructure predictable and stable. You’ll work across Kubernetes, networking, CI/CD, Cloudflare, and observability to create a platform engineers can trust. What You’ll Do Design, deploy, and maintain production Kubernetes clusters. Own cluster reliability, upgrades, security, and performance. … Build and operate monitoring, logging, and alerting pipelines. Ensure full-stack observability across infrastructure and services. Design and maintain CI/CD pipelines that are fast, reproducible, and safe. Improve deployment strategies (rollouts, canaries, rollbacks). Automate infrastructure provisioning and configuration. Investigate and resolve production incidents. Improve system resilience, redundancy ...

Senior Software Engineer

Location
United Kingdom
environment Work with AWS serverless technologies, including Lambda, SQS, EventBridge and API Gateway Work with MongoDB and document databases Contribute to CI/CD, observability, incident management and production support Work closely with Product, Engineering and Operations to turn business requirements into scalable solutions Mentor engineers and contribute to engineering … standards and best practice Strong Node.js & TypeScript Strong system design and architecture AWS & Serverless MongoDB/document databases Distributed systems, observability and incident management CI/CD and production operations Experience working in a regulated environment, ideally FinTech or financial services Comfortable owning services within a build-and-run environment ...

Senior Product Manager (SaaS)

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
comfort with technical details are must-haves. You'll be on the technical side too, having experience with containerised platforms using Kubernetes, databases, and observability tools such as Prometheus and OpenTelemetry. This is a chance to shape the future of observability and security, build products people count ...

Remote Senior Mobile Engineer - Instrumentation SDK (iOS) UK Remote

Location
Clevedon, Somerset, United Kingdom
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. ...

Remote Senior Mobile Engineer - Instrumentation SDK (iOS) UK Remote

Location
Bakewell, Derbyshire, United Kingdom
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. ...

Platform Engineer, Kubernetes & Automation — AI Cloud

Location
United Kingdom
seeking a Platform Engineer to operate and advance a cloud-native platform powering AI workloads. You will work on Kubernetes clusters, automation, and observability to deliver reliable services across our AI infrastructure. You will collaborate with software, infra, and SRE teams, mentor peers, and help define standards for deployment ...

Senior Platform & Cloud Engineer – Azure, DevOps

Location
Greater London, England, United Kingdom
cloud solutions. You will partner with architects and other engineers to deliver cloud adoption, environment design, and operational readiness, while embedding security, compliance and observability throughout. #J-18808-Ljbffr ...

Platform Engineer

Location
Greater London, England, United Kingdom
direction and build the systems, tooling and processes that the wider engineering team relies on. You’ll work across production infrastructure, Linux performance, observability, deployments, developer experience and internal tooling, with significant freedom to decide what needs improving and take ownership of delivering it. Responsibilities Own and improve infrastructure supporting … real-time, 24/7 production systems Build reliable deployment, rollback and operational workflows Improve observability across metrics, logging, dashboards, tracing and alerting Develop tooling and automation that allows engineers to ship faster and more safely Improve CI/CD, build processes, test environments and configuration workflows Work ...

Senior Python Data Platform Engineer for Finance Data Lakes

Location
Greater London, England, United Kingdom
components with Python in our London office. You will own ETL pipelines, data lakes/lakehouses, and distributed systems while improving CI/CD, observability, and infrastructure. Experience with financial/market data is essential, as is a strong background in databases and data-intensive systems. #J-18808-Ljbffr ...

Cloud-Native Backend Engineer for Data Processing

Location
Greater London, England, United Kingdom
with a focus on reliability and performance. You will collaborate with product managers and researchers to design scalable systems, use IaC, and contribute to observability with Prometheus and Loki. The team values curiosity and ownership, shipping robust software from Canary Wharf. #J-18808-Ljbffr ...

DataOps Engineer - Secure, Automated Data Pipelines

Location
United Kingdom
architect and deliver DataOps capabilities, focusing on repeatability, governance, and scalable data applications in air-gapped environments. You will implement and manage monitoring and observability to ensure data quality across the data flow, using tools like Airflow, Docker, Terraform and Kubernetes, with a strong emphasis on CI/ ...

Fullstack Engineer

Location
Cheltenham, England, United Kingdom
looking for a Fullstack Engineer with the following skills: The Fundamentals: Modern JavaScript GitLab CI/CD Observability Technologies Java GitLab Version Control Database Technologies Web API Integration Apache NiFi Message Broker Technologies OpenShift Web Auth Distributed Architecture Kubernetes SecDevOps Cloud Native Technologies Required Clearance: DV or recently active ...

Platform Engineer: Azure Cloud Infra, CI/CD & Kubernetes

Location
Greater London, England, United Kingdom
release systems, primarily Azure, Terraform and Ansible. This hands-on role focuses on reliable, scalable, and secure cloud environments, continuous integration and deployment, observability, cost optimisation, and collaboration with engineering teams to enable rapid, safe software delivery. #J-18808-Ljbffr ...

Remote Senior Mobile Engineer - Instrumentation SDK (iOS) UK Remote

Location
Loanhead, Midlothian, United Kingdom
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. ...

Automation-First QA Engineer for API & Backend

Location
England, United Kingdom
design and maintain automated test suites (Karate, Cucumber with RestAssured) in Java, integrate tests into CI/CD, and contribute to contract testing and observability with New Relic. #J-18808-Ljbffr ...

MLOps & DevOps Engineer for Production Platform

Location
Manchester, England, United Kingdom
specialist software team in Manchester. You will build and operate delivery pipelines, environments, and the model-serving path with IaC, CI/CD, observability, cost control, and production support across the platform. You will bring production experience, strong Terraform and Kubernetes skills, CI/CD ownership, and the ability ...

Lead AI Infra SRE: Scale, Reliability & Mentorship

Location
Gloucester, England, United Kingdom
improving automation to support AI/HPC workloads. You will configure and operate resilient Linux systems (Ubuntu), refine performance, and contribute to the observability stack with Prometheus and Grafana. #J-18808-Ljbffr ...

Hybrid AI Platform Engineer — MLOps/DevOps

Location
City Of London, England, United Kingdom
collaborate with data science and engineering teams to deploy AI workloads and ensure reliable infrastructure. The role focuses on building CI/CD pipelines, observability, and IaC automation, while optimizing performance and cost. You will ensure security and high availability for production systems. #J-18808-Ljbffr ...