1,576 to 1,600 of 2,452 Observability Jobs in London

Intelligent Automation Engineering Manager

Location
Greater London, England, United Kingdom
deployment and operational support of AI agents and AI-powered solutions. Establish engineering standards and best practices for AI architecture, orchestration, retrieval, tool invocation, observability, governance, privacy, security and cost management. Review technical designs and architecture documentation to ensure solutions align with engineering, security and governance standards. Translate emerging … with MCP (Model Context Protocol), MCP Servers, MCP Clients or enterprise AI integration frameworks. Knowledge of LLMOps, AI evaluation frameworks, model routing and AI observability tooling. Experience with enterprise integration technologies including APIs, Middleware, ESB or iPaaS platforms. Hands‐on experience integrating internal and third-party systems. Azure cloud experience ...

SC Cleared DevOps Engineer

Hiring Organisation
Sanderson Recruitment
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
Up to £570 per day + Inside IR-35
Develop and maintain Infrastructure as Code solutions. * Create and support CI/CD pipelines to enable secure, repeatable deployments. * Implement monitoring, logging, alerting and observability capabilities. * Ensure platform resilience, availability and operational readiness. * Manage configuration, secrets and environment controls. * Embed security best practice throughout the delivery lifecycle. * Support incident response … Code, ideally Terraform, CloudFormation or AWS CDK. * Strong CI/CD pipeline experience. * Experience with containerisation technologies, ideally Docker and Kubernetes. * Strong understanding of observability, monitoring and operational support. * Knowledge of cloud security, IAM, secrets management and governance controls. * Experience delivering and supporting production cloud environments within secure or regulated ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
infrastructure Troubleshoot and resolve infrastructure issues and fix root causes of issues at the source Implement monitoring, logging and alerting solutions to ensure system observability Collaborate with engineers and architects to improve platform documentation, standards and adoption Requirements Strong hands-on experience with cloud platforms (AWS & Azure ideally) Hands …/VNets, load balancers, DNS, security groups/NSGs) Experience with secrets management and identity/access control (IAM, OIDC, Azure AD) Familiarity with observability tooling (Prometheus, Grafana, CloudWatch, Azure Monitor) Risk Benefit Statement Learn more about the LexisNexis Risk team and how we work here We know your well ...

Technical Delivery Lead

Location
Greater London, England, United Kingdom
idempotency, and dead‐letter handling. Experience implementing data quality controls, including validation, reconciliation, and exception handling. Experience establishing engineering standards for coding, testing, documentation, observability, and security. Excellent written and verbal communication skills, with the ability to explain complex technical concepts to diverse audiences and align cross‐team efforts. Strong … onboarding vendor or third‐party source systems. Experience with event‐driven or streaming services such as Pub/Sub or Dataflow. Experience with observability tooling such as Cloud Monitoring, Cloud Logging, or Datadog. Experience in banking, insurance, or other regulated financial services environments. Bits In Glass (BIG) operates ...

Senior Test Automation Engineer

Location
Greater London, England, United Kingdom
Delivery Visual Studio Git Azure DevOps Repositories Azure DevOps YAML Pipelines Event Streaming & Messaging Apache Kafka (Aiven) Schema Registry Avro JSON Messaging RabbitMQ Data & Observability Microsoft SQL Server OpenTelemetry ClickStack/HyperDX ClickHouse ELK Stack (Elasticsearch, Logstash, Kibana) Excellent analytical and problem-solving capabilities. Strong stakeholder communication and collaboration skills. … Familiarity with Azure DevOps, TFS, Jira or similar Application Lifecycle Management tools. Exposure to observability and monitoring platforms. Knowledge of event-driven and microservices architectures. Experience Proven experience in test automation engineering within Agile environments. Strong experience in C# test automation development. Experience creating and maintaining BDD frameworks and behaviour ...

Platform Engineering Manager

Location
Greater London, England, United Kingdom
suppliers and service providers, ensuring effective service delivery, adherence to agreed service levels and that they remain compliant with regulatory expectations.* Lead platform operations, observability, incident management, root cause analysis, annual disaster recovery testing and operational resilience activities.* Manage platform lifecycle activities, including technical debt reduction, end-of-support risks … Azure cloud platforms.* Knowledge of Infrastructure as Code and automation tooling.* Knowledge of CI/CD platforms and developer productivity tooling.* Knowledge of observability, monitoring and operational reporting practices.* Knowledge of secure engineering and security-by-design principles.* Knowledge of incident management, operational resilience and disaster recovery concepts.* Knowledge ...

AI Consulting –Sr. AI Architect & Client Partner

Location
Greater London, England, United Kingdom
services that underpin enterprise AI ecosystems. Guide teams on software architecture, performance, scalability, security, and maintainability. AI governance & LLMOps Architect governance frameworks for auditability, observability, explainability, and compliance. Design guardrails for hallucination, prompt injection, toxicity, and model safety. Establish LLMOps: evaluation pipelines, automated testing, CI/CD, monitoring, and production … with one of Azure, OpenAI, AWS Bedrock, Claude; Kubernetes and cloud-native deployment. LLMOps & evaluation : CI/CD for AI, automated evals, experiment tracking, observability, model lifecycle management. Responsible AI : governance frameworks, guardrails, model safety, compliance, and auditability. Preferred but not required Orchestration frameworks : LangChain/LangGraph, LlamaIndex, CrewAI, AutoGen ...

Platform Engineer

Location
Greater London, England, United Kingdom
scalable solutions to thousands of customers every day. Our mission is to make deploying and operating software effortless and safe. We focus on automation, observability, and reliability, ensuring every engineering team at Clio can move faster and with confidence. You’ll be part of a globally distributed team, collaborating closely … Code using Terraform, ensuring reproducibility and compliance. Collaborate with developers to improve CI/CD pipelines, deployment strategies, and overall developer experience. Enhance observability and reliability, refining alerting, monitoring, and incident response. Collaborate on cloud optimization projects, improving performance, cost efficiency, and security posture. Mentor and guide team members, fostering ...

Senior Site Reliability Engineer

Hiring Organisation
CISCO Systems
Location
London, United Kingdom
Salary
£ 70 K
accelerate development, validation, and optimization of configuration-as-code, improving delivery speed and reducing deployment risk.Adaptable & Problem-Solver: Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance.Ownership & Quality: Own end-to-end configuration quality, enforcing … encourage applications without these:Hands-on with Helm or KustomizeExperience with GitOps (e.g., Argo CD)Knowledge of secrets management (e.g., HashiCorp Vault)Experience with observability (metrics/logs/tracing)CollabHiringWhy Cisco? At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations ...

Staff Software Engineer-AI

Hiring Organisation
Moody's Corporation
Location
London, UK
Employment Type
Full-time
equivalent technologies in production environmentsStrong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelinesDemonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … code maintainability, system performance, reliability, security, scalability, and cost efficiencyEstablish engineering best practices through hands-on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellenceChampion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvementPartner with product managers ...

Senior Software Engineer II, Developer Experience / Operational Excellence

Location
Greater London, England, United Kingdom
confidently. Within DevEx, the Operational Excellence (OPX) team is the group that keeps production healthy at scale. We provide engineering teams the platform capabilities, observability tooling, automated safeguards, incident management tooling, and safe feature release systems they need to deliver highly available systems, ship features with confidence, and investigate … health. Reduce alert noise, surface actionable signals, and empower engineering teams to operate their services confidently with minimal operational burden Develop and evolve our observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection, to give teams real-time, actionable visibility into system health and latency Contribute to AI-driven ...

Platform Engineer

Location
Greater London, England, United Kingdom
cost‐efficient infrastructure that empowers engineering teams to deliver quickly without sacrificing reliability or compliance. You will be the driving force behind our DevOps, observability, and compliance readiness , ensuring our systems are audit‐ready, highly available, and optimized for both performance and cost. Key Responsibilities Architect, implement, and maintain cloud … incredible journey and learning a lot along the way. Requirements Technical stack : Azure (also AWS is a plus), Terraform, AKS (Kubernetes), Docker, GitHub Actions. Observability : Experience implementing logging, metrics, and tracing frameworks. Security : Familiarity with best practices, secrets management, and security scanning tools. Networking : Solid understanding of VPCs, private networking ...

Senior Full Stack Engineer

Hiring Organisation
Liberty Global
Location
London, UK
Employment Type
Full-time
challenge technical designs and code produced by delivery partners, identifying risks and opportunities for improvementChampion engineering excellence across code quality, testing, continuous delivery, security, observability and platform reliabilityCollaborate closely with AI, Machine Learning, Product, Infrastructure and DevOps teams to deliver integrated platform capabilities and share knowledge across the engineering organisationWe … into production environmentsFamiliarity with AI orchestration frameworks, intelligent workflows and emerging AI technologiesExperience with data platforms and infrastructure that support AI workloadsKnowledge of monitoring, observability and continuous delivery practicesExperience within telecommunications or other large scale customer facing digital platformsWhat's in it for you? Competitive salary and bonus, where applicable25 ...

Infrastructure Engineer

Location
Greater London, England, United Kingdom
engineering team in building reliable, scalable applications Help design and build tools to make our services scalable and highly available, automating wherever possible Develop observability and orchestration tooling to allow our customers to monitor and manage their deployments Contribute to disaster recovery, backup and redundancy tooling and strategy Assist … hands-on experience Experience developing production-ready infrastructure management tooling with either Python or Golang Familiarity with at least one of the following: Observability Tools (e.g. Prometheus, OpenTelemetry, Grafana) Databases (e.g. Postgres, DuckDB) Event Streaming platforms (e.g. Kafka) Container Orchestration (e.g. Docker, Kubernetes) Familiarity with cloud platforms such ...

Senior Data Engineer (Data Platform)

Hiring Organisation
Teya Solutions
Location
London, UK
Employment Type
Full-time
Engineering team to other parts of TeyaGuaranteeing proper data governance following best practices, while not adding 100 steps of bureaucracyImproving data reliability, quality, and observability across key datasets, while ensuring data pipelines are simple to set up. Building, maintaining, and improving data models for large and complex datasets, while ensuring … ..)Hands-on experience provisioning and managing cloud infrastructure with TerraformExperience with CI/CD pipelines, automated testing and Git-based development workflowsFamiliarity with observability practices, including logging, metrics, alerting and production troubleshooting. Strong grasp of software engineering principles and best practicesExperience contributing to or leading data warehouse architecture ...

Sr Lead Software Engineer - Java / Python

Location
Greater London, England, United Kingdom
code written by others Leads technical analysis, estimation, planning, code reviews, architecture sessions, and retrospectives to drive delivery outcomes Establishes reliability goals and implements observability, resilience patterns, and operational readiness practices Leads incident response and post-incident reviews to improve production stability and performance; identifies recurring issues and drives automation … teams Preferred qualifications, capabilities, and skills Exposure to messaging systems and market protocols (e.g., MQ/Kafka; familiarity with FIX and Solace) Experience with observability stacks and resilience engineering for low‐latency/latency‐sensitive platforms Familiarity with Python Experience operating services in regulated environments with strong auditability and controls ...

Developer Experience Engineer

Location
Greater London, England, United Kingdom
with other engineering teams, informing us of our future project work and other improvements. In this role, you will navigate novel build challenges, refine observability and alerting systems for firm-wide critical services, and build great developer workflows. Your efforts will not only advance our team's delivery capabilities … Investigate and resolve intricate systems or build issues during support rotations, building valuable relationships with other DRW development teams. Improve service reliability, performance, and observability of firm-wide critical infrastructure and tooling. Work in a hybrid infrastructure environment, cloud and on-prem, using infra-as-code patterns to reliably manage ...

Sr Lead Software Engineer - Java / Python

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
debugs code written by othersLeads technical analysis, estimation, planning, code reviews, architecture sessions, and retrospectives to drive delivery outcomesEstablishes reliability goals and implements observability, resilience patterns, and operational readiness practicesLeads incident response and post-incident reviews to improve production stability and performance; identifies recurring issues and drives automation/remediationUpholds … globally distributed teamsPreferred qualifications, capabilities, and skillsExposure to messaging systems and market protocols (e.g., MQ/Kafka; familiarity with FIX and Solace)Experience with observability stacks and resilience engineering for low-latency/latency-sensitive platformsFamiliarity with PythonExperience operating services in regulated environments with strong auditability and controlsJ.P. Morgan ...

Senior Python Developer

Hiring Organisation
King Digital Entertainment
Location
London, UK
Employment Type
Full-time
internal and external systemsDeploy, monitor, and operate containerized applications in cloud-native environmentsContribute to technical designs, architectural discussions, and engineering best practicesImprove system reliability, observability, performance, and operational excellenceDeliver high-quality, well-tested, maintainable codeParticipate in code reviews and knowledge-sharing initiatives across teamsSupport CI/CD pipelines and automation … Teams, SharePoint, or related enterprise platformsExperience building internal developer platforms, business applications, or workflow automation solutionsExperience working with large-scale data processing systemsFamiliarity with observability and monitoring tools such as Grafana or SentryExperience improving application performance and scalability in production environmentsAbout KingWith a mission of Making the World Playful, King ...

Platform Chapter Lead - Engineering

Location
Greater London, England, United Kingdom
paved paths, less friction — using metrics (e.g. DORA) and real feedback to keep improving it. Keep it reliable and compliant. Oversee performance, resilience and observability for revenue‐critical services through peak traffic, and maintain security and compliance (e.g. PCI‐DSS, GDPR). Bring the business with you. Align platform strategy … balance cost, speed and risk. Strong cloud‐native and modern DevOps background — cloud (ideally AWS) and Kubernetes, CI/CD, infrastructure as code and observability — with enough engineering depth (e.g. Java, .NET, Python) to be credible with strong engineers. Excellent communication and the ability to influence at executive level, plus ...

Lead Java Developer

Location
Greater London, England, United Kingdom
adoption and ensure successful rollout of new capabilities. Lead root cause analysis on production issues, drive long‐term stability improvements, and strengthen monitoring and observability across the platform. Recommended Experience Strong experience in Core Java, J2EE, Spring Framework Exposure to Python scripting and data analysis Experience in fast moving Capital … such as Kafka, JMS, gRPC etc Proficient in latency measurement and performance optimization of Java based platforms with focus on JVM tuning Experience with observability stacks like ELK, Prometheus, Grafana, Kiali, Jaeger etc. Sound knowledge for persistence technologies such as relational databases, NoSQL databases, off heap storages and distributed caches ...

SRE Engineer

Hiring Organisation
Ricoh
Location
London, UK
Employment Type
Full-time
practices, tooling, and engineering standardsDriving infrastructure‐as‐code and automation across Azure and on‐premImproving image bakery pipelines for secure, repeatable server buildsEmbedding observability using metrics, logs, traces, and effective alertingEnsuring all practices align with ISO 27001 and internal security frameworksManaging automated patching, vulnerability remediation and configuration complianceBuilding dashboards … code (Terraform, ARM/Bicep), configuration management (Ansible, PowerShell DSC), and CI/CD tooling (Azure DevOps, GitHub Actions)Experience with monitoring and observability stacksSolid understanding of OS fundamentals (Windows/Linux), security, networkingBackground in scripting or software development (PowerShell, Python, Go)Experience with containers and orchestration (Docker, Kubernetes ...

Senior DevOps Engineer

Location
Greater London, England, United Kingdom
based teams on anything related to the technical aspects of our platform. What you'll do Help develop and maintain a robust monitoring and observability framework to ensure system performance, reliability, and early issue detection. Leverage your experience in DevOps to design, maintain, and improve cloud infrastructure. Own and evolve … roles. You’re an IT generalist who’s comfortable shaping and defining the role with us. You have a strong understanding of monitoring and observability tools (e.g., Datadog or Grafana). You have experience applying SRE principles such as SLOs, and error budgets. You have worked with incident management processes ...

Senior AI Engineer

Location
Greater London, England, United Kingdom
passthrough and logic centre for enterprise datasets, consuming other MCP servers and presenting them through one governed interface; and the orchestration, evaluation, and observability services beneath Libros and PRISM, two of the projects we are delivering with Percepta. As a Senior Applied AI Engineer, reporting to the Principal AI Engineer … own. Platforms built jointly with Percepta transfer into our ownership with maintainable designs and a team that can extend them without external help. Evaluation, observability, and control are built into what you ship rather than bolted on before release. Your responsibilities Build and run our AI products and platform Design ...

Lead Java Developer

Hiring Organisation
Citigroup
Location
London, UK
Employment Type
Full-time
adoption and ensure successful rollout of new capabilities. Lead root cause analysis on production issues, drive long‐term stability improvements, and strengthen monitoring and observability across the platform. Recommended Experience: Strong experience in Core Java, J2EE, Spring FrameworkExposure to Python scripting and data analysisExperience in fast moving Capital Markets Front … messaging technologies such as Kafka, JMS, gRPC etcProficient in latency measurement and performance optimization of Java based platforms with focus on JVM tuningExperience with observability stacks like ELK, Prometheus, Grafana, Kiali, Jaeger etc. Sound knowledge for persistence technologies such as relational databases, NoSQL databases, off heap storages and distributed cachesHands ...