876 to 900 of 2,340 Observability Jobs in London

Senior Platform Engineer

Location
Greater London, England, United Kingdom
people freedom while keeping us in control. What you'll be doing from day one Owning our infrastructure as code in Terraform, plus alerting, observability (Prometheus, Grafana) and reliability, including load testing and disaster recovery exercises. Making CI/CD faster (GitHub Actions, ArgoCD) and taking obstacles out of engineers … ideally multi-region. Hands-on with GCP, with Azure a bonus. You understand how model consumption works through each cloud. Deep experience with Terraform, observability and CI/CD, and you write solid Python (Elixir is a bonus, or you're keen to learn). A clear communicator ...

Principal Site Reliability Engineer, Infrastructure Observability

Location
Greater London, England, United Kingdom
opportunity to grow and make a difference in ways that matter to you. Role Summary In this role as Principal Site Reliability Engineer, Infrastructure Observability you will help formulate, develop, and implement a team of Site Reliability Engineers (SREs) focused on the observability, sustainability, scalability, measurability and recoverability … Proficiency with understanding and explaining incident situations and their recovery plans to prevent recurrence Knowledge/experience driving dashboard standardization across the ecosystem for observability, APM and infrastructure monitoring, and application‐specific logging Knowledge/experience with observability tools such as New Relic, SolarWinds DPA, Elastic Stack, Prometheus, Grafana, Splunk ...

Senior Backend Engineer (.NET & Python)

Location
Greater London, England, United Kingdom
ship, contributing to testing, troubleshooting and continuous improvement. Collaborate with Product, Design and Engineering teams to deliver customer-focused solutions. Improve platform reliability, observability, performance and security. Help evolve Benifex's AI‐assisted development practices and support other engineers across the team. What are we looking for? Commercial experience developing … backend applications using C#, .NET/.NET Core, Python and SQL. Production GenAI experience, including technologies such as MCP, RAG, agent orchestration, evaluations and observability tooling. An understanding of responsible AI, software quality, performance and security best practices. Experience delivering features end-to-end within modern engineering environments. Exposure ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
integrations, and human‐in‐the‐loop controls evolve across the stack. Continuously identify and exploit opportunities to improve performance, reliability, and user experience, using observability and analysis to find signals in noisy systems. Navigate confidently across legacy and greenfield contexts, applying AI tooling pragmatically to modernise where it matters most. … across both human-written and AI-generated code, embedding security validation into the development pipeline rather than treating it as an afterthought. Own system observability and reliability, moving from reactive alerting to proactive, model-assisted incident prevention. Participate in on-call rotation, bringing the same rigour to incident response that ...

Platform Engineer, SDO London, United Kingdom

Location
Greater London, England, United Kingdom
optimise cloud systems for performance, reliability, and cost efficiency. Assist in managing containerised workloads and orchestration platforms (e.g. Docker, Kubernetes). Implement and maintain observability tools (logging, metrics, alerting) to ensure system health and rapid incident response. Work with engineering teams to ensure infrastructure meets application requirements and supports scalable … cloud architecture, and system reliability. Strong troubleshooting and problem‐solving skills. Desirable: Experience with containerisation and orchestration (Docker, Kubernetes). Familiarity with monitoring and observability tools (e.g. Prometheus, Grafana, ELK). Experience working with Linux systems and shell scripting. Programming or scripting experience (e.g. Python, Bash, Go, or similar). ...

Platform Engineer, SDO

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
optimise cloud systems for performance, reliability, and cost efficiency. Assist in managing containerised workloads and orchestration platforms (e.g. Docker, Kubernetes).Implement and maintain observability tools (logging, metrics, alerting) to ensure system health and rapid incident response. Work with engineering teams to ensure infrastructure meets application requirements and supports scalable software … networking, cloud architecture, and system reliability. Strong troubleshooting and problem-solving skills. Desirable: Experience with containerisation and orchestration (Docker, Kubernetes).Familiarity with monitoring and observability tools (e.g. Prometheus, Grafana, ELK).Experience working with Linux systems and shell scripting. Programming or scripting experience (e.g. Python, Bash, Go, or similar).Understanding ...

Senior Network Site Reliability Engineer

Location
Greater London, England, United Kingdom
Build and lead all aspects of our CI/CD pipelines (GitLab/GitHub) and provide automated solutions for IaC deployment. Implement automation and observability across infrastructure resources. Bring your own ideas for improving our infrastructure stack and implement them. Take part in our operational rotation to react and resolve … teams Nice to have Experience with containers and Kubernetes Familiarity with AWS governance and security controls, including SCPs and IAM policies Experience improving reliability, observability, performance, or incident response processes Exposure to large-scale CDN, edge, or traffic-routing environments Experience working in globally distributed infrastructure or platform teams Basic ...

Cloud Support Engineer

Location
Greater London, England, United Kingdom
providing technical depth, structure and calm during high-pressure client situations Develop and refine tools, dashboards and automation to improve support delivery, observability and onboarding Identify recurring issues, propose and lead solutions that improve platform stability, reduce effort and enhance client experience Provide mentoring, training and technical oversight … production troubleshooting Strong Linux systems knowledge, including filesystems, networking and system internals Programming skills in Golang and Python and experience with infrastructure tools or observability stacks (e.g. Grafana, Prometheus, EFK) Confidence in working with cloud-native platforms and tools (e.g. Kubernetes, Terraform, AWS/GCP, Docker) Excellent communication skills under ...

Global Banking & Markets - Software Engineer - Vice President - London London · United Kingdom [...]

Location
Greater London, England, United Kingdom
standard for years to come. What You Will Do Design, build, and operate high‐availability, multi‐region, cloud‐native services with security and comprehensive observability (metrics, distributed tracing, structured logging) built in at every layer. Develop event‐driven architectures, multi‐stage processing pipelines, and optimized data paths for high‐throughput … patterns (retry, dead‐letter queues, error isolation). Cloud & Infrastructure : Cloud platforms (GCP, AWS), container orchestration (Kubernetes, Docker), and JVM tuning for containerized workloads. Observability & Operations : Application instrumentation (metrics, distributed tracing, structured logging) and production support in high‐availability environments. Data & Performance : Data modeling, SQL/NoSQL databases, caching strategies ...

Forward Deployed Engineer - Platform Engineer

Hiring Organisation
Kyndryl
Location
London, UK
Employment Type
Full-time
guardrails, with security as a first-class concern (policy-as-code/OPA, access controls, secrets management, compliance-driven engineering) Instrument platforms for observability (Grafana, Prometheus, OpenTelemetry) Capture deployment learnings and share best practices to inform core platform frameworks Contribute code, automated blueprints, and feedback to core platform teams … secure network access to endpoints Solid grasp of CI/CD pipelines, version control (Git & GitHub), and microservices/API architectures Practical experience with observability tooling (Grafana, Prometheus, OpenTelemetry) Working knowledge of generative AI platforms: LLM hosting, LLM gateways (e.g. LiteLLM, Portkey, Kong AI Gateway), and MLOps/LLMOps practices ...

Senior Fullstack Engineer (Python + React.js)

Location
Greater London, England, United Kingdom
minimal downtime. Write unit and integration tests to maintain code reliability and ensure high- quality releases. Continuously monitor and optimize backend performance using observability tools such as Datadog, Cloud Watch or similar. Participate in design discussions and decision-making to enhance system robustness and scalability. Maintain technical documentation to ensure … handling asynchronous communication. Experience with Infrastructure as Code (IaC) tools like Terraform or CloudFormation for managing cloud infrastructure. Knowledge of observability and monitoring tools, such as Cloud Watch or Datadog, to track and troubleshoot system performance. Familiarity with serverless architectures (e.g., AWS Lambda) and event‐-driven programming paradigms. Exposure ...

DevOps Team Manager

Hiring Organisation
Bromcom Computers Plc
Location
Bromley, London, United Kingdom
Employment Type
Permanent
technical quality while enabling engineers to own their work. Set and maintain engineering standards for Azure architecture, Azure DevOps, Bicep/ARM, deployment patterns, observability, resilience, security and operational support. Challenge designs and changes using risk, maintainability, failure-mode, rollback and supportability thinking; involve senior engineers and technical leadership where … access follows least-privilege principles, is reviewed regularly and is supported by effective joiner-mover-leaver, break-glass and segregation-of-duties controls. Own observability standards across Azure Monitor and Grafana, security and vulnerability follow-up, and cloud cost and FinOps accountability for the Azure estate. Stakeholder & Cross-Team Influence ...

AIML Software Engineer, AI for Science

Location
City Of London, England, United Kingdom
infrastructure as code. Strong problem-solving and debugging skills, and experience working in cluster settings or cloud-based environments. Experience operating production services — monitoring, observability and alerting, and diagnosing and resolving issues in live systems. Experience designing and administering SQL databases — schema design, query performance, and day-to-day operational … including defining and working to service-level objectives (SLOs/SLIs). Experience with incident response and post-incident review, and with building the observability that supports it. Infrastructure-as-code (e.g. Terraform) for provisioning and maintaining cloud environments. Experience developing and administering workloads on Kubernetes (e.g. GKE). Familiarity ...

Platform Engineer – Monitoring, Observability & SIEM (MONSO)

Location
Greater London, England, United Kingdom
Platform Engineer – Monitoring, Observability & SIEM (MONSO) For our SPEAR Technology (Security, Platform Engineering, Automation and Runtime) division in London we are looking to hire a: Platform Engineer – Monitoring, Observability & SIEM (MONSO) Like solving puzzles with an inquisitive mind? Think outside the box and challenge the status quo? Prefer simplicity over … proactive ownership? Then consider joining Berenberg’s SPEAR Technology programme. SPEAR consists of our CyberSecurity team and several platform engineering teams responsible for Monitoring, Observability, Kubernetes, Developer Platform, Network, and Datacentre Infrastructure. Due to each team’s compact size, all team members are subject matter experts offering an excellent environment ...

Platform Engineer - Monitoring, Observability & SIEM (MONSO)

Hiring Organisation
Berenberg
Location
London, UK
Employment Type
Full-time
Platform Engineer – Monitoring, Observability & SIEM (MONSO) Persönliche Daten Land Vereinigtes Königreich Stadt London Art der Anstellung Professional Arbeitszeit Vollzeit Vertragsart Unbefristet Offene Stellen 1 Beschreibung & Anforderungen For our SPEAR Technology (Security, Platform Engineering, Automation and Runtime) division in London we are looking to hire a: Platform Engineer – Monitoring, Observability & SIEM … proactive ownership? Then consider joining Berenberg's SPEAR Technology programme. SPEAR consists of our CyberSecurity team and several platform engineering teams responsible for Monitoring, Observability, Kubernetes, Developer Platform, Network, and Datacentre Infrastructure. Due to each team's compact size, all team members are subject matter experts offering an excellent environment ...

Senior Platform & Cloud Engineer – Azure, DevOps

Location
Greater London, England, United Kingdom
cloud solutions. You will partner with architects and other engineers to deliver cloud adoption, environment design, and operational readiness, while embedding security, compliance and observability throughout. #J-18808-Ljbffr ...

Cloud-Native Backend Engineer for Data Processing

Location
Greater London, England, United Kingdom
with a focus on reliability and performance. You will collaborate with product managers and researchers to design scalable systems, use IaC, and contribute to observability with Prometheus and Loki. The team values curiosity and ownership, shipping robust software from Canary Wharf. #J-18808-Ljbffr ...

Senior GenAI Platform Engineer - Real-Time GPU Inference

Location
Greater London, England, United Kingdom
serving, inference, and training pipelines, collaborating across DoorDash, Wolt, and Deliveroo. You will set direction for GPU autoscaling, end-to-end serving stacks, and observability while mentoring engineers and shaping production-grade capabilities for AI-driven products. #J-18808-Ljbffr ...

Platform Engineer: Azure Cloud Infra, CI/CD & Kubernetes

Location
Greater London, England, United Kingdom
release systems, primarily Azure, Terraform and Ansible. This hands-on role focuses on reliable, scalable, and secure cloud environments, continuous integration and deployment, observability, cost optimisation, and collaboration with engineering teams to enable rapid, safe software delivery. #J-18808-Ljbffr ...

Hybrid AI Platform Engineer — MLOps/DevOps

Location
City Of London, England, United Kingdom
collaborate with data science and engineering teams to deploy AI workloads and ensure reliable infrastructure. The role focuses on building CI/CD pipelines, observability, and IaC automation, while optimizing performance and cost. You will ensure security and high availability for production systems. #J-18808-Ljbffr ...

Platform Engineer: Core Backend & Developer Tools

Location
Greater London, England, United Kingdom
shared backend services, frameworks, and tooling that improve how internal teams develop, deploy, and operate software. The role emphasizes platform engineering, API standards, and observability across services. The ideal candidate has extensive backend experience, strong Python skills, and familiarity with FastAPI, Docker, Kubernetes, and CI/CD in cloud contexts. ...

Lead Platform Engineer: GCP, Terraform & CI/CD

Location
Greater London, England, United Kingdom
will design, build and operate scalable infrastructure, lead a small team of engineers, and shape platform standards across engineering teams. You will manage security, observability, and compliance while enabling product teams to ship rapidly. The role blends hands-on engineering with leadership, mentoring, and roadmap responsibility. #J-18808-Ljbffr ...

Senior DevOps Engineer: Cloud, Kubernetes & CI/CD

Location
Greater London, England, United Kingdom
automation initiatives while collaborating with developers and operations teams. You will build reproducible infrastructure, manage Kubernetes-based workloads, and drive CI/CD and observability improvements across AWS deployments. A fit combines hands-on cloud and container skills with strong collaboration. #J-18808-Ljbffr ...

Senior Payments & Incentives Engineer — Scale FinTech

Location
Greater London, England, United Kingdom
high-scale campaigns. The role emphasizes collaboration with product, finance, and revenue operations, plus on-call incident management and a focus on operational reliability, observability, and SDLC in an Agile environment. #J-18808-Ljbffr ...

Senior DevEx & Platform Engineer - AWS/K8s

Location
Greater London, England, United Kingdom
/Terraform-based platforms, and guide reliability initiatives at a global scale. You will lead on-call, privacy-focused infrastructure, and end-to-end observability, collaborating across distributed teams and driving security-first engineering practices. #J-18808-Ljbffr ...