901 to 925 of 2,373 Observability Jobs in London

Cloud Support Engineer

Location
Greater London, England, United Kingdom
providing technical depth, structure and calm during high-pressure client situations Develop and refine tools, dashboards and automation to improve support delivery, observability and onboarding Identify recurring issues, propose and lead solutions that improve platform stability, reduce effort and enhance client experience Provide mentoring, training and technical oversight … production troubleshooting Strong Linux systems knowledge, including filesystems, networking and system internals Programming skills in Golang and Python and experience with infrastructure tools or observability stacks (e.g. Grafana, Prometheus, EFK) Confidence in working with cloud-native platforms and tools (e.g. Kubernetes, Terraform, AWS/GCP, Docker) Excellent communication skills under ...

Global Banking & Markets - Software Engineer - Vice President - London London · United Kingdom [...]

Location
Greater London, England, United Kingdom
standard for years to come. What You Will Do Design, build, and operate high‐availability, multi‐region, cloud‐native services with security and comprehensive observability (metrics, distributed tracing, structured logging) built in at every layer. Develop event‐driven architectures, multi‐stage processing pipelines, and optimized data paths for high‐throughput … patterns (retry, dead‐letter queues, error isolation). Cloud & Infrastructure : Cloud platforms (GCP, AWS), container orchestration (Kubernetes, Docker), and JVM tuning for containerized workloads. Observability & Operations : Application instrumentation (metrics, distributed tracing, structured logging) and production support in high‐availability environments. Data & Performance : Data modeling, SQL/NoSQL databases, caching strategies ...

Forward Deployed Engineer - Platform Engineer

Hiring Organisation
Kyndryl
Location
London, UK
Employment Type
Full-time
guardrails, with security as a first-class concern (policy-as-code/OPA, access controls, secrets management, compliance-driven engineering) Instrument platforms for observability (Grafana, Prometheus, OpenTelemetry) Capture deployment learnings and share best practices to inform core platform frameworks Contribute code, automated blueprints, and feedback to core platform teams … secure network access to endpoints Solid grasp of CI/CD pipelines, version control (Git & GitHub), and microservices/API architectures Practical experience with observability tooling (Grafana, Prometheus, OpenTelemetry) Working knowledge of generative AI platforms: LLM hosting, LLM gateways (e.g. LiteLLM, Portkey, Kong AI Gateway), and MLOps/LLMOps practices ...

Senior Fullstack Engineer (Python + React.js)

Location
Greater London, England, United Kingdom
minimal downtime. Write unit and integration tests to maintain code reliability and ensure high- quality releases. Continuously monitor and optimize backend performance using observability tools such as Datadog, Cloud Watch or similar. Participate in design discussions and decision-making to enhance system robustness and scalability. Maintain technical documentation to ensure … handling asynchronous communication. Experience with Infrastructure as Code (IaC) tools like Terraform or CloudFormation for managing cloud infrastructure. Knowledge of observability and monitoring tools, such as Cloud Watch or Datadog, to track and troubleshoot system performance. Familiarity with serverless architectures (e.g., AWS Lambda) and event‐-driven programming paradigms. Exposure ...

DevOps Team Manager

Hiring Organisation
Bromcom Computers Plc
Location
Bromley, London, United Kingdom
Employment Type
Permanent
technical quality while enabling engineers to own their work. Set and maintain engineering standards for Azure architecture, Azure DevOps, Bicep/ARM, deployment patterns, observability, resilience, security and operational support. Challenge designs and changes using risk, maintainability, failure-mode, rollback and supportability thinking; involve senior engineers and technical leadership where … access follows least-privilege principles, is reviewed regularly and is supported by effective joiner-mover-leaver, break-glass and segregation-of-duties controls. Own observability standards across Azure Monitor and Grafana, security and vulnerability follow-up, and cloud cost and FinOps accountability for the Azure estate. Stakeholder & Cross-Team Influence ...

AIML Software Engineer, AI for Science

Location
City Of London, England, United Kingdom
infrastructure as code. Strong problem-solving and debugging skills, and experience working in cluster settings or cloud-based environments. Experience operating production services — monitoring, observability and alerting, and diagnosing and resolving issues in live systems. Experience designing and administering SQL databases — schema design, query performance, and day-to-day operational … including defining and working to service-level objectives (SLOs/SLIs). Experience with incident response and post-incident review, and with building the observability that supports it. Infrastructure-as-code (e.g. Terraform) for provisioning and maintaining cloud environments. Experience developing and administering workloads on Kubernetes (e.g. GKE). Familiarity ...

Platform Engineer – Monitoring, Observability & SIEM (MONSO)

Location
Greater London, England, United Kingdom
Platform Engineer – Monitoring, Observability & SIEM (MONSO) For our SPEAR Technology (Security, Platform Engineering, Automation and Runtime) division in London we are looking to hire a: Platform Engineer – Monitoring, Observability & SIEM (MONSO) Like solving puzzles with an inquisitive mind? Think outside the box and challenge the status quo? Prefer simplicity over … proactive ownership? Then consider joining Berenberg’s SPEAR Technology programme. SPEAR consists of our CyberSecurity team and several platform engineering teams responsible for Monitoring, Observability, Kubernetes, Developer Platform, Network, and Datacentre Infrastructure. Due to each team’s compact size, all team members are subject matter experts offering an excellent environment ...

Platform Engineer - Monitoring, Observability & SIEM (MONSO)

Hiring Organisation
Berenberg
Location
London, UK
Employment Type
Full-time
Platform Engineer – Monitoring, Observability & SIEM (MONSO) Persönliche Daten Land Vereinigtes Königreich Stadt London Art der Anstellung Professional Arbeitszeit Vollzeit Vertragsart Unbefristet Offene Stellen 1 Beschreibung & Anforderungen For our SPEAR Technology (Security, Platform Engineering, Automation and Runtime) division in London we are looking to hire a: Platform Engineer – Monitoring, Observability & SIEM … proactive ownership? Then consider joining Berenberg's SPEAR Technology programme. SPEAR consists of our CyberSecurity team and several platform engineering teams responsible for Monitoring, Observability, Kubernetes, Developer Platform, Network, and Datacentre Infrastructure. Due to each team's compact size, all team members are subject matter experts offering an excellent environment ...

Senior Platform & Cloud Engineer – Azure, DevOps

Location
Greater London, England, United Kingdom
cloud solutions. You will partner with architects and other engineers to deliver cloud adoption, environment design, and operational readiness, while embedding security, compliance and observability throughout. #J-18808-Ljbffr ...

Cloud-Native Backend Engineer for Data Processing

Location
Greater London, England, United Kingdom
with a focus on reliability and performance. You will collaborate with product managers and researchers to design scalable systems, use IaC, and contribute to observability with Prometheus and Loki. The team values curiosity and ownership, shipping robust software from Canary Wharf. #J-18808-Ljbffr ...

Senior GenAI Platform Engineer - Real-Time GPU Inference

Location
Greater London, England, United Kingdom
serving, inference, and training pipelines, collaborating across DoorDash, Wolt, and Deliveroo. You will set direction for GPU autoscaling, end-to-end serving stacks, and observability while mentoring engineers and shaping production-grade capabilities for AI-driven products. #J-18808-Ljbffr ...

Platform Engineer: Azure Cloud Infra, CI/CD & Kubernetes

Location
Greater London, England, United Kingdom
release systems, primarily Azure, Terraform and Ansible. This hands-on role focuses on reliable, scalable, and secure cloud environments, continuous integration and deployment, observability, cost optimisation, and collaboration with engineering teams to enable rapid, safe software delivery. #J-18808-Ljbffr ...

Hybrid AI Platform Engineer — MLOps/DevOps

Location
City Of London, England, United Kingdom
collaborate with data science and engineering teams to deploy AI workloads and ensure reliable infrastructure. The role focuses on building CI/CD pipelines, observability, and IaC automation, while optimizing performance and cost. You will ensure security and high availability for production systems. #J-18808-Ljbffr ...

Platform Engineer: Core Backend & Developer Tools

Location
Greater London, England, United Kingdom
shared backend services, frameworks, and tooling that improve how internal teams develop, deploy, and operate software. The role emphasizes platform engineering, API standards, and observability across services. The ideal candidate has extensive backend experience, strong Python skills, and familiarity with FastAPI, Docker, Kubernetes, and CI/CD in cloud contexts. ...

Lead Platform Engineer: GCP, Terraform & CI/CD

Location
Greater London, England, United Kingdom
will design, build and operate scalable infrastructure, lead a small team of engineers, and shape platform standards across engineering teams. You will manage security, observability, and compliance while enabling product teams to ship rapidly. The role blends hands-on engineering with leadership, mentoring, and roadmap responsibility. #J-18808-Ljbffr ...

Senior DevOps Engineer: Cloud, Kubernetes & CI/CD

Location
Greater London, England, United Kingdom
automation initiatives while collaborating with developers and operations teams. You will build reproducible infrastructure, manage Kubernetes-based workloads, and drive CI/CD and observability improvements across AWS deployments. A fit combines hands-on cloud and container skills with strong collaboration. #J-18808-Ljbffr ...

Senior Payments & Incentives Engineer — Scale FinTech

Location
Greater London, England, United Kingdom
high-scale campaigns. The role emphasizes collaboration with product, finance, and revenue operations, plus on-call incident management and a focus on operational reliability, observability, and SDLC in an Agile environment. #J-18808-Ljbffr ...

Senior DevEx & Platform Engineer - AWS/K8s

Location
Greater London, England, United Kingdom
/Terraform-based platforms, and guide reliability initiatives at a global scale. You will lead on-call, privacy-focused infrastructure, and end-to-end observability, collaborating across distributed teams and driving security-first engineering practices. #J-18808-Ljbffr ...

Lead AI Security Engineer — Agentic AI (Python/AWS) Remote

Location
Greater London, England, United Kingdom
hands-on development, secure delivery, and building reusable AI-native components within a modern security automation program. A strong emphasis on CI/CD, observability, and security is required. #J-18808-Ljbffr ...

Senior Platform Engineer: Hybrid Cloud & HashiCorp Expert

Location
Greater London, England, United Kingdom
approaches and treat infrastructure as software to support engineering and research teams. The role emphasizes Terraform and Ansible for deployment pipelines, strong security and observability, and collaboration with software engineers, quantitative #J-18808-Ljbffr ...

Senior Database Platform Engineer: Cloud & SQL Mastery

Location
Greater London, England, United Kingdom
modernise towards PostgreSQL and cloud-native services, and collaborate across Platform, Data, and Operations teams to ensure resilient data services. The role emphasises automation, observability, security and incident response, with remote-first flexibility and occasional travel to London for planning #J-18808-Ljbffr ...

Senior Quality Engineer: Lead Testing & Automation

Location
City Of London, England, United Kingdom
quality and testing strategy, write test charters, and partner with product and engineering to shift-left quality and prevent defects. You will enhance our observability stack, engage in risk-based testing, mentor teammates, and contribute to automation, using Docker, C#/.NET, Azure DevOps and cloud technologies. Join ...

Cloud‐Native SRE for CTL Platform – Associate

Location
City Of London, England, United Kingdom
Foundation Engineering - SRE Platforms – Associate in London to help build, run and improve the CTL cloud-native platform. You will focus on reliability, scalability, observability and incident response for a mission-critical trading system. Strong programming in Java or Go, expertise in distributed systems, and experience with Prometheus, Grafana ...

Senior Cloud Platform DevOps Engineer — Hybrid UK

Location
City Of London, England, United Kingdom
engineer to drive a cloud-forward, AI-driven SaaS platform. You will build and scale secure cloud-native serverless infrastructure, focusing on deployment pipelines, observability, and security to support generative AI workflows. You will work in an agile environment alongside artists and pipeline teams, enabling rapid feature delivery while maintaining ...

Remote Cloud-Native Engineering Manager

Location
Greater London, England, United Kingdom
global 24x7 operation. You will shepherd architectural direction, implement serverless and container-based solutions on AWS, and drive devops, CI/CD, and observability practices across the teams. The role requires 10+ years in software engineering with at least 5 years in management, a strong background ...