1,201 to 1,225 of 2,378 Observability Jobs in London

Devops SRE

Location
Greater London, England, United Kingdom
shared Kubernetes services such as CoreDNS , cert‐manager , Dynatrace , Cloudability , and Infoblox . Familiarity with OPA Gatekeeper for policy enforcement and tenant isolation. Security, Observability & Performance Strong security mindset with a proven track record of designing secure, resilient cloud‐native systems. Experience implementing observability stacks including Prometheus , Dynatrace , and OpenTelemetry ...

Platform Engineer

Location
Greater London, England, United Kingdom
efficiently and securely. Working closely with software developers, architects, and delivery teams, you will help establish best practices around cloud infrastructure, CI/CD, observability, security, and application reliability. The role combines hands-on engineering with the opportunity to influence platform standards and development practices across multiple projects. Key Responsibilities … with modern front-end technologies, including React and Vite. Develop and improve CI/CD pipelines, automation, and developer tooling. Implement monitoring, logging, and observability solutions to improve system reliability and performance. Manage and optimise cloud infrastructure, containerised environments, and platform configurations. Ensure security and operational best practices are embedded ...

Camunda Architect / Lead Architect (Camunda 8 Preferred)

Location
Greater London, England, United Kingdom
strategy across Camunda SaaS and self-managed models Collaborate with engineering and DevOps teams on: CI/CD pipelines containerised deployments Kubernetes-based environments observability and platform monitoring Technical Leadership Provide technical leadership and architectural guidance across engineering teams and delivery squads Mentor engineers and support the uplift of workflow … Camunda 8 Experience with Camunda SaaS and self-managed deployments at scale Exposure to cloud platforms such as: AWS Azure GCP Experience with observability tooling such as: Prometheus Grafana OpenTelemetry Experience with: enterprise integration patterns API management workflow/task UI customisation Experience in regulated industries such as: Banking Insurance ...

Operations Engineering Lead

Hiring Organisation
Willis Towers Watson
Location
London, UK
Employment Type
Full-time
base platform operations team. You will help set direction for shared infrastructure and engineering operations - the production hosting, security perimeter, release pipelines, observability, and developer tooling that every engineering squad depends on to ship safely. This is a role that includes Line Management responsibility for your team … networking, scalability, performance, and cost-efficiency across production environmentsOversee Infrastructure as Code practices (e.g. Terraform) and ensure environments are consistent, auditable, and secureShape monitoring, observability and alerting across the engineering organisation (e.g. Datadog, CloudWatch) so issues are detected and resolved before customers are impacted4Lead incident management practices - on-call, triage ...

Site Reliability Engineer (remote working)

Location
Greater London, England, United Kingdom
integration of development artefacts with wider enterprise platforms Supporting production and non-production environments and taking ownership of incidents when requiredUsing monitoring and observability to identify potential reliability and performance issues Working with Engineering Managers, Technical Leads and Delivery Managers to drive technical initiatives forward Providing technical guidance and mentoring … containerised environments Azure DevOps and CI/CD Infrastructure as Code - Terraform, Bicep or ARM Azure identity, secrets and access management Monitoring and observability Automation and scripting Microservices and cloud-native environments Production support and incident management Experience with Grafana, Azure Monitor, Log Analytics, Application Insights, PowerShell or ServiceNow would ...

Cloud Platform Engineer

Location
Greater London, England, United Kingdom
both human and machine/agent identities; secure by default, least privilege, secrets management and continuous compliance. Apply SRE practices - SLOs/SLIs, observability, capacity planning, resilience and blameless incident management - to keep the platform reliable and cost-efficient. Partner with data engineering to design and optimise data pipelines, data … workloads and data ‐ including secrets management, least‐privilege IAM and machine/workload identity. Solid programming/scripting ability (e.g. Python, Go) and strong observability, reliability and cost‐optimisation practices. Desirable requirements: Experience working as a Site Reliability Engineer (SRE) with SLOs/SLIs, error budgets and incident management. ...

Embedded Engineer - DC Cleared

Hiring Organisation
VIQU IT Recruitment
Location
London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
£600 - 700 per day
Work alongside data, AI and platform engineers so models, pipelines and infrastructure land as one product. Improve the day to day: CI/CD, observability, security hardening and documentation. Write code you would be happy to inherit, and hold that same standard in code review. Work with the consultancy … secrets management. Putting AI or ML models into production, including LLM features. Infrastructure as code with Terraform, Ansible or similar, plus configuration management. Observability and incident response: metrics, logging and tracing. Message queues, event-driven architectures or data pipelines. Kubernetes: deploying, running or debugging workloads, including K3s, RKE2 or OpenShift. ...

Embedded Engineer - DC Cleared

Hiring Organisation
VIQU IT Recruitment
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£600.00 - £700.00 per day
Work alongside data, AI and platform engineers so models, pipelines and infrastructure land as one product. Improve the day to day: CI/CD, observability, security hardening and documentation. Write code you would be happy to inherit, and hold that same standard in code review. Work with the consultancy … secrets management. Putting AI or ML models into production, including LLM features. Infrastructure as code with Terraform, Ansible or similar, plus configuration management. Observability and incident response: metrics, logging and tracing. Message queues, event-driven architectures or data pipelines. Kubernetes: deploying, running or debugging workloads, including K3s, RKE2 or OpenShift. ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
architecture and technical decision‐making as our platform evolves Champion engineering excellence through automated testing, CI/CD, code reviews and pair programming Use observability and monitoring to understand service health, troubleshoot complex distributed systems and improve reliability Mentor and support other engineers, sharing knowledge and helping raise technical capability … modern CI/CD practices Strong understanding of automated testing, code quality, security and software engineering best practices Experience troubleshooting distributed systems and using observability to improve performance and reliability Strong communication and collaboration skills, with experience contributing to technical decisions Experience mentoring and supporting other engineers through code reviews ...

Senior Data Engineer (Data Platform)

Location
Greater London, England, United Kingdom
other parts of Teya - Guaranteeing proper data governance following best practices, while not adding 100 steps of bureaucracy - Improving data reliability, quality, and observability across key datasets, while ensuring data pipelines are simple to set up. - Building, maintaining, and improving data models for large and complex datasets, while ensuring they … experience provisioning and managing cloud infrastructure with Terraform - Experience with CI/CD pipelines, automated testing and Git-based development workflows - Familiarity with observability practices, including logging, metrics, alerting and production troubleshooting. - Strong grasp of software engineering principles and best practices - Experience contributing to or leading data warehouse architecture ...

Deployed Architect, Professional Services (London)

Location
Greater London, England, United Kingdom
together. LangChain is a place where your contributions can shape how this technology shows up in the real world. Today, our platform includes LangSmith (Observability, Evaluation, Deployment, Fleet, and Sandboxes), our open source frameworks (LangChain, LangGraph, and Deep Agents), and the newly launched LangSmith Engine for autonomous agent improvement. … backup strategies, and sizing Experience designing high-availability and disaster recovery solutions Strong understanding of networking, security (SSO/RBAC, TLS, secrets management), and observability (Prometheus, Grafana, Datadog) Experience with CI/CD pipelines for infrastructure and applications Agent Engineering & Development: 1+ years of experience building production AI/ ...

Senior Python Developer

Hiring Organisation
Tech 4
Location
City of London, London, United Kingdom
Employment Type
Permanent
integrations, and human-in-the-loop controls evolve across the stack. Continuously identify and exploit opportunities to improve performance, reliability, and user experience, using observability and analysis to find signals in noisy systems. Navigate confidently across legacy and greenfield contexts, applying AI tooling pragmatically to modernise where it matters most. … evolve architecture pragmatically. Strong system design fundamentals across scalability, performance, and distributed systems, including API design (REST, GraphQL). Hands-on experience with observability tooling (Datadog, Grafana, or similar) and a data-informed approach to system health and reliability. Solid SQL and data management skills, with an appreciation ...

Senior Python Developer (PYTHON/REACT/AZURE)

Hiring Organisation
Tech4 Limited
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP 80,000 - 110,000 Annual
integrations, and human-in-the-loop controls evolve across the stack. Continuously identify and exploit opportunities to improve performance, reliability, and user experience, using observability and analysis to find signals in noisy systems. Navigate confidently across Legacy and greenfield contexts, applying AI tooling pragmatically to modernise where it matters most. … evolve architecture pragmatically. Strong system design fundamentals across scalability, performance, and distributed systems, including API design (REST, GraphQL). Hands-on experience with observability tooling (Datadog, Grafana, or similar) and a data-informed approach to system health and reliability. Solid SQL and data management skills, with an appreciation ...

Snr Lead Software Engineer - Developer Experience Engineering

Location
Greater London, England, United Kingdom
Senior Lead Software Engineer at JPMorganChase within the International Consumer Bank, you will lead a team building and operating a first-class observability capability across our cloud-native microservices. Collaborating closely with product, platform, and development teams, you will influence the technical roadmap, architect, standardize, and build resilient, cost-efficient … experience designing and implementing cloud multi-region architectures in production environments Proficiency in one or more additional programming languages beyond primary experience Familiarity with observability platforms and telemetry tooling for metrics, logs, and distributed tracing at scale #ICBCareers #J-18808-Ljbffr ...

Senior Reliability Engineer

Hiring Organisation
Fitch Ratings
Location
London, UK
Employment Type
Full-time
someone who is curious about the evolving role of AI in infrastructure engineering, someone who actively explores how AI-assisted tooling, automation, and intelligent observability can raise the bar for reliability and developer experience. You will collaborate closely with global development and engineering teams to deliver reliable, resilient, and high … reliability, security, and efficiencyIdentify, contain, and mitigate risk across all cloud environments, maintaining a robust security posture for infrastructure and applicationsImplement proactive monitoring and observability practices to detect and prevent issues before they impact usersDevelop and maintain automation and tooling solutions, including AI-assisted approaches to reduce toil and accelerate ...

Senior Machine Learning Platform/Ops Engineer

Location
Greater London, England, United Kingdom
scientist creating reproducible, containerized model training environments (on-demand and scheduled), and manage compute at scale (e.g., spot/GPU autoscaling) Define and implement observability and alerting for ML systems (model drift, data quality, feature coverage, etc.) Design and scale data ingestion and feature transformation flows using batch (e.g., Spark … Kubernetes, and CI/CD workflows Understanding of ML model lifecycles: training, validation, deployment, and monitoring Strong DevOps practices: Git, IaC (Terraform), logging/observability, containerization (Docker/K8s) Ability to work independently with ML Scientists and mentor peers in reliability, testing, and delivery. Product impact driven. Exposure ...

Senior Backend Engineer

Location
Greater London, England, United Kingdom
Lead code reviews, mentor engineers, and ensure high engineering standards. Contribute to platform SDKs, internal libraries, and backend frameworks used across charters. Performance, Reliability & Observability Implement tracing, structured logging, metrics, dashboards, and alerting. Optimize services for latency, concurrency, throughput, and cost efficiency. Ensure system reliability through automated testing, load testing … Kubernetes, and cloud platforms (AWS or GCP). Strong CI/CD experience using GitHub Actions, Jenkins, Argo, or similar tools. Deep knowledge of observability practices — logs, metrics, traces, performance analysis. Proven ability to write clean, testable, well‐structured production code. Qualifications – Nice to Have Experience with streaming or queuing ...

Data Platform Engineer

Hiring Organisation
MONY Group
Location
London, UK
Employment Type
Full-time
personalisation. Stay close to the business context and apply software engineering practices to solve specific data problems. Improve monitoring, alerting, data quality checks and observability so issues are detected and understood quickly. Improve engineering workflows with AI and automationIdentify repeated or high-friction engineering tasks and turn them into reliable … improving cloud-based systems. Familiarity with infrastructure-as-code, CI/CD, version control and automated testing. Ability to reason about reliability, security, observability and operational support. Experience working with technical and non-technical stakeholders and communicating clearly. Curiosity about AI-assisted engineering and automation, with an interest in applying ...

Forward Deployed Engineer - Lead Platform Engineer

Hiring Organisation
Kyndryl
Location
London, UK
Employment Type
Full-time
guardrails, with security as a first-class concern (policy-as-code/OPA, access controls, secrets management, compliance-driven engineering) Instrument platforms for observability (Grafana, Prometheus, OpenTelemetry) Provide architectural oversight across multi-disciplinary workstreams, staying close enough to unblock the team directly Capture field learnings, codify reusable patterns and blueprints … secure network access to endpoints Hands-on experience with CI/CD pipelines, Git-based workflows, and microservices/API architectures Practical experience with observability stacks (Grafana, Prometheus, OpenTelemetry) Experience with generative AI platforms: LLM hosting, LLM gateways (e.g. LiteLLM, Portkey, Kong AI Gateway), and MLOps/LLMOps practices ...

Senior Software Engineer - Food (Distributed Systems)

Location
Greater London, England, United Kingdom
through clean, maintainable and well-tested code, promoting best practices through code reviews, pair programming, documentation and continuous improvement initiatives. Drive operational excellence and observability by designing effective monitoring and alerting, leveraging tools such as Dynatrace and participating in support activities to ensure critical supply chain and pricing data remains … uses a variety of technologies, including: Backend: Java, Spring, Spring Boot, Micronaut Frontend: React, Next.js, TypeScript, Angular Cloud & Infrastructure: Azure Cloud, Kubernetes Observability: Dynatrace Databases: SQL Server, MongoDB Caching & Performance: Ignite, Redis What's in it for you? Working at M&S means being part of something bigger - helping ...

Lead Software Engineer - AI-Native Applications

Hiring Organisation
IFS
Location
London, UK
Employment Type
Full-time
technologies such as Kubernetes, Kafka, Redpanda, PostgreSQL and MongoDB.Experience working with AWS and/or Azure. Strong understanding of CI/CD, automated testing, observability and production operations. Experience designing secure software, authentication and authorisation mechanisms and applying DevSecOps principles. Strong analytical and problem-solving skills with the ability … Experience building Enterprise SaaS or ERP products. Experience working with Model Context Protocol (MCP) or similar AI integration standards. Experience with vector databases, AI observability or AI evaluation frameworks. Experience designing and delivering enterprise-scale AI platforms or AI-powered products. QualificationsA degree in Computer Science, Software Engineering or Information ...

Production AI Engineer - Vice President

Hiring Organisation
Citigroup
Location
London, UK
Employment Type
Full-time
engineering techniques to integrate large language models (LLMs) into operational tooling, incident response pipelines, and developer productivity platforms. Leads the development of AI-native observability solutions — leveraging intelligent agents to detect anomalies, predict failures, and automate remediation before issues impact end users. Writes clean, well-tested, and well-documented code … Kanban).Operational experience of using middleware technologies (MQ, Apache Kafka, etc.) to run services at scale is desirable. Strong experience with end-to-end observability stacks (Datadog, AppDynamics, Dynatrace, etc.) is desirable. Degree in Computer Science, Mathematics, Physics, or a related technical subject is desirable. Experience of senior stakeholder management. ...

Senior AI Architect| London

Hiring Organisation
Infosys Technologies
Location
London, UK
Employment Type
Full-time
hallucination in production. LLMOps, Evaluation & Responsible AI: Experience operationalizing LLM and agentic systems at scale—evaluation harnesses and metrics for quality, groundedness, and safety; observability, tracing, and monitoring (e.g., LangSmith, LangFuse); guardrails and red-teaming; and continuous optimization of accuracy, cost, and latency. Understanding of AI governance, security, privacy, bias … standards for agentic AI—agent orchestration, MCP-based tool/data integration, shared skills and connectors, memory and state management, guardrails, human oversight, and observability—to enable safe, reliable, and scalable production deployment across teams. Solution Implementation: Collaborate with data scientists and engineers to implement Generative AI solutions, ensuring ...

Engineering Manager (Identity)

Hiring Organisation
ebury
Location
London, UK
Employment Type
Full-time
Ebury helps ambitious businesses unlock global growth, and we take the same approach with our people. We encourage innovation and movement, collaboration and problem-solving, and foster an environment where everyone can feel they belong ...