351 to 375 of 3,893 Permanent Observability Jobs

Senior Architect, ADC/Quant

Location
Greater London, England, United Kingdom
compliance requirements typical of the London financial services sector.* Resilience & SRE: Advanced knowledge of building fault-tolerant architectures leveraging modern SRE principles, robust observability, and cloud-native resilience patterns.* Matrix Leadership: Exceptional stakeholder and client communication skills, with a track record of influencing cross-functional engineering pods, product managers ...

Lead Solution Architect

Hiring Organisation
BP
Location
London, United Kingdom
Salary
£ 70 K
intelligent scheduling, anomaly detection, automated P&L attribution, predictive maintenance and agentic workflow orchestration (e.g., AWS AgentCore).Define and govern DevOps, platform engineering and observability standards, including CI/CD pipelines, infrastructure-as-code, containerisation (Docker, Kubernetes), monitoring, alerting and incident response architecture.People, Community & GovernanceMentor and develop the architecture community ...

Lead Solution Architect

Hiring Organisation
BP
Location
London, UK
Employment Type
Full-time
intelligent scheduling, anomaly detection, automated P&L attribution, predictive maintenance and agentic workflow orchestration (e.g., AWS AgentCore).Define and govern DevOps, platform engineering and observability standards, including CI/CD pipelines, infrastructure-as-code, containerisation (Docker, Kubernetes), monitoring, alerting and incident response architecture. People, Community & GovernanceMentor and develop the architecture ...

Senior Software Engineer Ref. 3839

Location
Cheltenham, England, United Kingdom
particularly interested in candidates with experience in technologies and practices such as C++, Golang, Python, Rust, CMake, HELM, YAML, JSON, containerisation, VCPkg, observability, DevOps, Kubernetes, systems administration and Linux. However, we also welcome applications from candidates whose experience aligns with any of the technologies and disciplines outlined above. ...

Senior Software Engineer Ref. 3839

Location
Manchester, England, United Kingdom
particularly interested in candidates with experience in technologies and practices such as C++, Golang, Python, Rust, CMake, HELM, YAML, JSON, containerisation, VCPkg, observability, DevOps, Kubernetes, systems administration and Linux. However, we also welcome applications from candidates whose experience aligns with any of the technologies and disciplines outlined above. ...

Software Engineer — Observability Instrumentation

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
high-impact research - designing systems that scale, accelerate discovery and support innovation across the firm. Take the next step in your career. The roleThe Observability Engineering Team manages access to G-Research's telemetry platforms, ensuring our engineering teams can effectively produce and consume telemetry for their services. … looking for a technically strong, customer-focused Software Engineer to help make observability easier to adopt across the organisation. This role focuses on the producer side: instrumentation patterns, OpenTelemetry SDKs and the collector configurations that help teams emit consistent, high-quality telemetry. This role is suited to someone who enjoys ...

Platform Engineer – Monitoring, Observability & SIEM (MONSO)

Location
Greater London, England, United Kingdom
Platform Engineer – Monitoring, Observability & SIEM (MONSO) For our SPEAR Technology (Security, Platform Engineering, Automation and Runtime) division in London we are looking to hire a: Platform Engineer – Monitoring, Observability & SIEM (MONSO) Like solving puzzles with an inquisitive mind? Think outside the box and challenge the status quo? Prefer simplicity over … proactive ownership? Then consider joining Berenberg’s SPEAR Technology programme. SPEAR consists of our CyberSecurity team and several platform engineering teams responsible for Monitoring, Observability, Kubernetes, Developer Platform, Network, and Datacentre Infrastructure. Due to each team’s compact size, all team members are subject matter experts offering an excellent environment ...

Platform Engineer

Location
Greater London, England, United Kingdom
efficiently and securely. Working closely with software developers, architects, and delivery teams, you will help establish best practices around cloud infrastructure, CI/CD, observability, security, and application reliability. The role combines hands-on engineering with the opportunity to influence platform standards and development practices across multiple projects. Key Responsibilities … with modern front-end technologies, including React and Vite. Develop and improve CI/CD pipelines, automation, and developer tooling. Implement monitoring, logging, and observability solutions to improve system reliability and performance. Manage and optimise cloud infrastructure, containerised environments, and platform configurations. Ensure security and operational best practices are embedded ...

AWS DevOps Engineer

Location
Greater London, England, United Kingdom
Amazon EKS, automating infrastructure with Terraform, orchestrating containers via EKS, building robust CI/CD with GitOps, and implementing strong monitoring/observability + AWS EKS IAM concepts to support secure, high-performance, cost-optimised systems. You will work closely with developers and SREs in an automation-first, collaborative culture. … Support live production AWS/EKS environments, including on-call rotation, incident response, root cause analysis, and blameless post-mortems. Automate provisioning, configuration, scaling, observability, and secure workload access across AWS services. Collaborate on improving system reliability, observability, security (DevSecOps/EKS IAM), and AWS cost optimization. Qualifications: Strong Linux ...

Senior DevOps Engineer - AVP

Location
Belfast City District, Northern Ireland, United Kingdom
operations on a global scale. In this role, you will apply deep technical expertise across CI/CD pipelines, container orchestration, cloud infrastructure, and observability to deliver resilient, high-quality software systems. Your work will directly shape how Citi's engineering teams build, ship, and monitor production services across … engineering teams. Build and manage containerized workloads on Kubernetes and OpenShift using Helm, ensuring systems are scalable, reliable, and production ready. Architect and maintain observability solutions — including log aggregation with Splunk and Elastic/Kibana, and metrics monitoring with Prometheus and Grafana — to give engineering teams real-time visibility into ...

Director of Platform Engineering

Location
Greater London, England, United Kingdom
ITRS, we make society’s critical technology work. Our mission is to deliver automated and holistic IT observability solutions that safeguard critical applications and enable innovation. We are the only monitoring and observability platform designed for the most demanding and regulated industries — trusted by 90% of Tier 1 capital markets … this team is closely aligned with platform engineering practices, customer delivery priorities and operational standards. The ITRS engineering teams are building a next-generation observability platform with the capability to collect, store and analyse the vast amount of data generated by banks and financial institutions. Requirements As Director of Platform ...

Platform Engineer Graduate Considered

Hiring Organisation
RedTech Recruitment Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£55,000
hands-on role with Kubernetes at its core, offering the opportunity to work across container orchestration, CI/CD, Infrastructure as Code, observability and cloud-native technologies. Location: Cambridge, 2 days per week in the office Salary: £40,000 - £60,000 per annum + benefits Requirements for Platform Engineer Graduate … efficiently at scale Work closely with software engineering teams to understand infrastructure requirements and encourage scalable, cloud-native development practices Improve the reliability, observability and performance of Kubernetes clusters Build and maintain CI/CD pipelines to support efficient software delivery Develop and improve Infrastructure as Code using technologies such ...

Member of technical staff (Infrastructure) - Paris

Location
Greater London, England, United Kingdom
Company’s agent platform including client-facing APIs and agent runtimes within various deployment scenarios (multi-tenant and on-prem). Setup and maintain observability and monitoring strategies. Requirements: MUST HAVE Observability and monitoring (Datadog, Prometheus, Grafana, ...) Good knowledge of a modern programming language (ideally Python or JS/ ...

Sr Lead AI Platform Engineer

Location
Auchentibber, Scotland, United Kingdom
Owns the design and build of the team's platform: deployment pipelines, model serving, containerisation, orchestration, and environment management Sets the standard for reliability, observability, and operational excellence across the team's production AI/ML services Builds the tooling and paved paths that let AI engineers ship agentic … record of building deployment and release automation Experience serving, scaling, and monitoring ML models or data-intensive services in production (MLOps) Practical experience with observability tooling (metrics, logging, tracing) and production incident response Experience operating ML/LLM workloads in production (LLMOps, inference reliability, cost/performance management) Strong communication ...

Senior ML Ops Engineer

Hiring Organisation
Harnham - Data & Analytics Recruitment
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£75,000 - £85,000 per annum
production-grade orchestration pipelines for model training, inference, monitoring, and retraining. Developing CI/CD pipelines to automate testing, deployment, and operational processes. Implementing observability, monitoring, and alerting frameworks to ensure platform reliability and performance. Driving best practice around Responsible AI, governance, transparency, and auditability. Collaborating with engineering, delivery … practices using Azure DevOps or similar tooling. Experience with infrastructure-as-code technologies such as Terraform, Bicep, or Pulumi. Knowledge of monitoring and observability tooling such as Prometheus and Grafana. Experience with orchestration platforms including Dagster, Airflow, Prefect, or similar. BENEFITS The successful Senior MLOps Engineer will receive the following ...

Senior Platform Engineer (Python)

Location
Slough, England, United Kingdom
workload scheduling, configuration, and runtime resource access. Experience building CI/CD pipelines with GitLab CI, GitHub Actions, Jenkins, or similar tools. Experience with observability tools such as Prometheus, Grafana, and OpenTelemetry. Good understanding of platform security: secrets management, IAM, network isolation, and dependency/supply-chain risks. Strong troubleshooting … teams to identify recurring pain points and turn them into scalable platform features. Automate manual infrastructure operations and build reliable self-service workflows. Implement observability through metrics, structured logging, and alerting. Own platform production issues from troubleshooting and root cause analysis to permanent resolution. Manage and scale on-prem compute ...

Camunda Architect / Lead Architect (Camunda 8 Preferred)

Location
Greater London, England, United Kingdom
strategy across Camunda SaaS and self-managed models Collaborate with engineering and DevOps teams on: CI/CD pipelines containerised deployments Kubernetes-based environments observability and platform monitoring Technical Leadership Provide technical leadership and architectural guidance across engineering teams and delivery squads Mentor engineers and support the uplift of workflow … Camunda 8 Experience with Camunda SaaS and self-managed deployments at scale Exposure to cloud platforms such as: AWS Azure GCP Experience with observability tooling such as: Prometheus Grafana OpenTelemetry Experience with: enterprise integration patterns API management workflow/task UI customisation Experience in regulated industries such as: Banking Insurance ...

DevOps Engineer - SC Cleared

Hiring Organisation
Opus Recruitment Solutions Ltd
Location
Martock, Somerset, United Kingdom
Employment Type
Full-Time
Salary
£350.00 - £500.00 per day
CI. Automate deployment, configuration and operational processes to improve efficiency and reliability. Support containerised workloads and cloud-native deployment patterns. Implement monitoring, logging and observability solutions to improve platform visibility and operational performance. Drive infrastructure resilience, scalability, security and cost optimisation initiatives. Support incident response, troubleshooting and root cause analysis … Code expertise. Experience building and maintaining CI/CD pipelines, ideally using GitLab CI. Strong Docker and containerisation experience. Experience with monitoring and observability tools such as Prometheus, Grafana and CloudWatch. Knowledge of AWS security best practices, IAM and Well-Architected principles. Experience implementing automation and reducing manual operational overhead. ...

Operations Engineering Lead

Hiring Organisation
Willis Towers Watson
Location
London, UK
Employment Type
Full-time
base platform operations team. You will help set direction for shared infrastructure and engineering operations - the production hosting, security perimeter, release pipelines, observability, and developer tooling that every engineering squad depends on to ship safely. This is a role that includes Line Management responsibility for your team … networking, scalability, performance, and cost-efficiency across production environmentsOversee Infrastructure as Code practices (e.g. Terraform) and ensure environments are consistent, auditable, and secureShape monitoring, observability and alerting across the engineering organisation (e.g. Datadog, CloudWatch) so issues are detected and resolved before customers are impacted4Lead incident management practices - on-call, triage ...

Lead DevOps Engineer

Location
Greater London, England, United Kingdom
secure supply chain configurations. Establish IaC and GitOps standards, adding automated testing to every infrastructure change. Prototype agentic infrastructure components, including deployment and observability platforms in service meshes. Contribute to the Kong AI Gateway, including Dataplane deployments, ACM/SSL integration, and observability via DataDog Champion DevSecOps maturity by embedding … Data, and AI teams to shape DevOps and AI platform architectures with regulatory compliance. Stay ahead of CNCF and AI ecosystem innovations, from eBPF observability to agent‐aware orchestration. What We’re Looking For We value diverse backgrounds and paths to expertise. If you bring solid experience in several ...

Operations Engineering Lead

Hiring Organisation
WTW
Location
Greater London, United Kingdom
Employment Type
Full Time
base platform operations team. You will help set direction for shared infrastructure and engineering operations - the production hosting, security perimeter, release pipelines, observability, and developer tooling that every engineering squad depends on to ship safely. This is a role that includes Line Management responsibility for your team … performance, and cost-efficiency across production environments Oversee Infrastructure as Code practices (e.g. Terraform) and ensure environments are consistent, auditable, and secure Shape monitoring, observability and alerting across the engineering organisation (e.g. Datadog, CloudWatch) so issues are detected and resolved before customers are impacted4 Lead incident management practices - on-call ...

Devops SRE

Location
Greater London, England, United Kingdom
shared Kubernetes services such as CoreDNS , cert‐manager , Dynatrace , Cloudability , and Infoblox . Familiarity with OPA Gatekeeper for policy enforcement and tenant isolation. Security, Observability & Performance Strong security mindset with a proven track record of designing secure, resilient cloud‐native systems. Experience implementing observability stacks including Prometheus , Dynatrace , and OpenTelemetry ...

Staff Software Engineer - Databases SRE | UK | Remote

Hiring Organisation
Grafana Labs
Location
United Kingdom
Salary
£ 70 K
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. ...

Senior Cloud Architect (all genders)

Hiring Organisation
Lam Research
Location
Villach, Kärnten, Austria
Employment Type
Permanent
Salary
EUR Annual
guardrails. Developer Platform & Experience Build an internal developer platform (IDP) that abstracts complexity and provides self-service: project scaffolding, environment creation, CI/CD, observability, and secrets management. Curate platform catalogs/registries (Terraform module registry, container base images, reusable pipeline templates). Partner with product/engineering teams … Policy, OPA, Conftest) and continuous compliance. Build secure-by-default blueprints: network segmentation, private endpoints, encryption, vulnerability mgmt, and SBOM/SLSA practices. SRE, Observability, and Operations Embed SRE practices: SLOs, error budgets, incident response, postmortems, chaos/gamedays. Standardize observability: logs, metrics, traces, dashboards, and alerts (e.g., Azure Monitor ...