1,601 to 1,625 of 2,452 Observability Jobs in London

Cloud Network Engineer - Active UK Government Security Clearance Required

Location
Greater London, England, United Kingdom
Support routing, switching, firewalling, VPN, and load-balancing technologies Contribute to network upgrades, migrations, optimisation, and continuous improvement Monitor network health and performance using observability and monitoring tooling Participate in root-cause analysis and problem management Maintain technical documentation, configuration records, and operational procedures Collaborate with engineering, service management … networking BGP/OSPF VLANs/VRFs MPLS Firewalls such as Palo Alto or Check Point VPN/IPSec Load balancing Network monitoring and observability Enterprise networking in production environments Nice to have: Cisco ACI/Nexus Cisco SD-WAN/SDA AWS networking - VPC, Transit Gateway, Route 53 Azure ...

Lead Java Developer

Location
Greater London, England, United Kingdom
adoption and ensure successful rollout of new capabilities.* Lead root cause analysis on production issues, drive long‐term stability improvements, and strengthen monitoring and observability across the platform.**Recommended Experience:*** Strong experience in Core Java, J2EE, Spring Framework* Exposure to Python scripting and data analysis* Experience in fast moving Capital … such as Kafka, JMS, gRPC etc* Proficient in latency measurement and performance optimization of Java based platforms with focus on JVM tuning* Experience with observability stacks like ELK, Prometheus, Grafana, Kiali, Jaeger etc.* Sound knowledge for persistence technologies such as relational databases, NoSQL databases, off heap storages and distributed caches ...

Lead DevOps Engineer

Location
Greater London, England, United Kingdom
Terragrunt).* Ensure deployments are secure, fast and auditable.* Raise the bar on automation so routine operational tasks are eliminated, not managed.* Own the observability strategy across the platform using Datadog.* Define what good looks like: SLO/SLA dashboards, alerting thresholds, runbooks and the feedback loops that let teams … within a compliant environment, and comfortable in security audits and risk conversations.* Comfortable with security tooling such as CrowdStrike and Rapid7, SIEM/SOC.* Observability ownership with Datadog (or equivalent), having defined SLOs, built the dashboards and set the alerting culture.* Experience leading or mentoring a DevOps team, with ...

Principal Software Engineer

Location
Greater London, England, United Kingdom
needs of the wider vertical.* Champion AWS-native design and operational practices, including secure and resilient use of serverless, messaging, storage, compute, identity, observability and infrastructure-as-code capabilities.* Guide teams in building secure APIs, backend services and integrations that meet the needs of business-critical financial services and regulated … make appropriate design decisions in partnership with data specialists.* Experience designing reusable APIs, services and integrations, with a strong focus on secure development, observability, resilience and maintainability.* Strong understanding of cloud security, application security and risk-aware design within a regulated environment.* Excellent analytical and problem-solving skills, with ...

DevSecOps Architect - Director

Location
Greater London, England, United Kingdom
regulation, covering explainability, human oversight, drift monitoring and incident handling. Define the reference architecture for AI platform components: model gateway, retrieval layer, guardrails, observability and audit trail. Qualifications, Skills and Experience Significant experience architecting CI/CD and DevSecOps platforms at enterprise scale in a regulated environment. Demonstrable delivery … Veracode, Trivy, secrets management (HashiCorp Vault, cloud-native equivalents), OPA/policy-as-code, SBOM and SLSA concepts. Platform: Kubernetes, containers, Terraform, service mesh, observability stacks AI/ML: MLOps and LLMOps tooling, model registries, vector stores and retrieval architectures, evaluation frameworks, agent frameworks, AI gateway patterns Experience operating under ...

Full Stack Engineer

Hiring Organisation
Liberty Global
Location
London, UK
Employment Type
Full-time
features and platform capabilities end to end, from design through build, test and deployment into production. Contribute to production readiness: documentation, runbooks, monitoring and observability – ensuring nothing ships until it's genuinely ready. Maintain and improve existing platform components, addressing technical debt and performance issues with clear prioritisation. Standards & QualityChampion … with data pipelines, data quality processes and the infrastructure that supports AI/ML at scale. Experience with DevOps, CI/CD pipelines and observability tooling. Exposure to telecom or large-scale consumer-facing platforms. SKILLS & BEHAVIOURAL COMPETENCIESTechnical Depth & JudgementConsistently makes sound technical decisions, balancing short-term delivery with long ...

Senior AI Platform Engineer

Hiring Organisation
9fin
Location
London, UK
Employment Type
Full-time
developer tooling that enable self-service AI development across engineering teams. Design secure, scalable deployment pipelines for AI models and applications. Build AI observability capabilities including monitoring, tracing, evaluation, cost optimisation, and production quality measurement. Collaborate closely with AI Engineers, Backend Engineers and Engineering Leadership to define platform architecture … monitor, and operate AI services in production. AI Operations & ObservabilityHave experience implementing monitoring, tracing, evaluation, and cost optimisation for AI systems. Have experience with observability solutions such as Arize Phoenix, Langfuse, or LangsmithUnderstand the operational challenges of deploying LLM-powered applications, including latency, reliability, hallucination monitoring, and model quality evaluation. ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
enablement work Established objective measures of "done" and documented assumptions clearly, so both infrastructure and tooling initiatives had transparent success criteria Contributed to observability as code initiatives, both inside the team and across engineering Infrastructure & reliability Provided a highly available and performant platform, including our AWS, Kubernetes (EKS) and MongoDB … genuine interest in developer experience: you get satisfaction from removing friction for other engineers, not just from the infrastructure itself Worked with monitoring and observability tools such as Datadog General Proficiency in at least one of: Go, Python, Typescript, or strong shell scripting skills Strong communication skills and comfort influencing ...

Principal - Head of Compliance Architecture

Hiring Organisation
Northern Trust
Location
London, UK
Employment Type
Full-time
Drives enterprise resiliency, recoverability, and operational excellence initiatives including active-active architectures, hot-hot disaster recovery models, blue-green deployment strategies, high availability design, observability, fault tolerance, and business continuity readiness.* Establishes and improves engineering discipline and software delivery practices, including CI/CD pipelines, DevSecOps operating models, automated testing … warehouses, master data management, metadata management, AI enablement, and advanced analytics platforms. Enterprise resiliency, disaster recovery, active-active architectures, high availability, blue-green deployments, observability, and operational readiness frameworks. Modern DevSecOps and software engineering disciplines including CI/CD, automated testing, infrastructure-as-code, secure software delivery, and software lifecycle ...

Senior Site Reliability Engineer

Location
City Of London, England, United Kingdom
initiatives, drive automation efforts to reduce operational toil, and help build resilient systems that deliver exceptional customer experiences. You will leverage your expertise in observability, incident response, and distributed systems to proactively identify and resolve reliability challenges. Working closely with engineering teams, you will design and implement solutions that improve … Networking & Security: Proficiency in VPCs, networking, ALBs, Route53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, and cloud security best practices. Incident Response & Observability: Skilled in troubleshooting using logs, metrics, alarms, deployment history, root cause analysis, rollback decisions, and operational runbooks. Linux & Automation: Strong Linux and Git fundamentals with Bash ...

Staff Software Engineer, Liquidity Management (C#/.NET) New London, UK

Location
Greater London, England, United Kingdom
work that distributes effectively across the team Establish coding standards, review practices, and testing strategies that improve overall code quality Champion engineering methodologies including observability, incident response, and production excellence Share your enthusiasm for tech trends, explore and learn new technologies, engage with tech communities, mentor fellow engineers, and lead … cloud platform (preferably Azure) Experience with containerization (Docker) and orchestration (Kubernetes) Understanding of infrastructure as code and CI/CD pipeline design Knowledge of observability: logging, metrics, tracing, and alerting strategies Experience with microservices deployment patterns and service mesh concepts Core Technologies: Deep expertise with C#/.NET Expert ...

Full Stack Engineer

Location
Greater London, England, United Kingdom
features and platform capabilities end to end, from design through build, test and deployment into production. Contribute to production readiness: documentation, runbooks, monitoring and observability – ensuring nothing ships until it’s genuinely ready. Maintain and improve existing platform components, addressing technical debt and performance issues with clear prioritisation. Standards & Quality … with data pipelines, data quality processes and the infrastructure that supports AI/ML at scale. Experience with DevOps, CI/CD pipelines and observability tooling. Exposure to telecom or large-scale consumer-facing platforms. SKILLS & BEHAVIOURAL COMPETENCIES Technical Depth & Judgement Consistently makes sound technical decisions, balancing short-term delivery ...

Senior Software Engineer - Customer & Claims

Location
Greater London, England, United Kingdom
Create an environment where less experienced engineers can grow and do their best work. Continuous Improvement Drive improvements to engineering practices, testing, deployment and observability within the domain. Continuously grow your own capability and that of the engineers around you. Contribute to Homeprotect's wider engineering standards as they mature. … oriented language), and familiarity with a modern frontend framework. Deep, hands‐on understanding of modern software engineering practices: continuous integration and deployment, automated testing, observability and cloud‐native development. Experience operating production systems, in a public cloud environment such as Azure or GCP, that need to be reliable, secure ...

AI and Cloud Director - Technology Consulting (AI, data and analytics)

Hiring Organisation
Business Integration Partners
Location
London, UK
Employment Type
Full-time
AI and Cloud DirectorLocation: London/HybridBusiness: BIP UKPractice: xTech - AI, Data and CloudReporting to: UK xTech LeadershipEmployment type: Full timeAbout BIP UKFounded in 2003, BIP is an international consulting firm with more than 6 ...

Software Engineer (Backend)

Location
Greater London, England, United Kingdom
Software Engineer (backend) Location: London or New York - in office 4 days a week About Lorum Global payments are not broken. Incentives are. Clearing has been deprioritized inside balance sheet driven institutions whose models rely ...

AI & ML Engineer

Location
Greater London, England, United Kingdom
About the role The AI & ML Engineering team accelerates the adoption of AI across the business, championing innovation while ensuring our machine learning solutions are robust, scalable, and cost-efficient. We enable teams to solve ...

Senior Data Engineer

Hiring Organisation
McKesson
Location
London, UK
Employment Type
Full-time
Role OverviewThe Senior Data Engineer is the technical owner of the ClarusONE data platform. This is a hands-on engineering role with real ownership: alongside building and delivering data pipelines, you will raise the engineering ...

Senior Infrastructure Engineer

Location
Greater London, England, United Kingdom
mission is to abstract away multi-cloud complexity and reduce cognitive load for other developers. We treat our platform infrastructure which spans, orchestration, observability, and data. As a first-class product, enabling our internal teams and global clients to operate securely and reliably at massive scale. Duties Scale the Core … engineers and bank SREs. Design for Resiliency: Engineer robust, fault-tolerant architectures and Disaster Recovery/Business Continuity Planning (DR/BCP) systems. Advance Observability: Enhance our comprehensive observability suites that are shared across the entire Thought Machine product ecosystem, providing clients with deep, out-of-the-box monitoring ...

Senior Data Platform Engineer

Location
Greater London, England, United Kingdom
policy-as-code, and action logging. Build platform-, tool-, and agent-level kill switches, alongside dry-run/safe-testing modes for agent workflows. Observability & FinOps Implement AI system observability, including prompt logging, output monitoring, quality scoring, and drift detection. Establish operational monitoring for gateway usage, latency, error rates … plain English. And to really stand out from the crowd Experience with AI gateways, LLM proxies, or model-routing patterns. Exposure to AI observability tools, vector databases (Databricks AI Search), RAG pipelines, or non-deterministic system monitoring. Hands-on experience implementing policy-as-code, action logging, or AI kill switches. ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
initiatives, drive automation efforts to reduce operational toil, and help build resilient systems that deliver exceptional customer experiences. You will leverage your expertise in observability, incident response, and distributed systems to proactively identify and resolve reliability challenges. Working closely with engineering teams, you will design and implement solutions that improve … call rota and provide on-call support for other SRE engineers.* Can write advanced automation scripts for incident response, including failovers and rollbacks.**Observability** * Has a deep technical understanding of observability techniques across the full stack and can bring clarity to complex incidents or performance issues.* Able to create templated ...

Lead Cloud Engineer

Location
Greater London, England, United Kingdom
data orchestration toolsets (e.g., dbt, Apache Airflow), ETL/ELT methodologies, real‐time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost‐control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated ...

IBM Netcool / Observability Technical Lead

Hiring Organisation
Deerfoot Recruitment Solutions
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
£780 - £830 per day
Netcool/Observability Technical Lead Inside IR35 Contract -up to £827pd London Hybrid - 4 Days Onsite/1 Day WFH per Week Banking Are you the person who knows exactly why an ObjectServer failover didn't behave as expected, and how to stop a flood of duplicate events before anyone … shape how thousands of infrastructure and application events are detected, correlated and actioned across EMEA, and you'll have genuine scope to modernise observability capability rather than simply keep the lights on. This is a hands-on technical leadership role with no direct reports, so your influence comes from your ...

Enterprise AI Deployment Architect

Location
Greater London, England, United Kingdom
commercial execution across product, engineering, and sales teams. You will translate business needs into deployment playbooks, drive measurable outcomes, and ensure governance, security, and observability across customer environments. #J-18808-Ljbffr ...

Site Reliability Engineer — Cloud & Live Ops (Hybrid)

Location
Uxbridge, England, United Kingdom
resilient, secure, and highly available platforms underpinning live, broadcast‐adjacent services. Based at Stockley Park in Uxbridge with hybrid options, the role involves improving observability, incident response, automation, disaster recovery, and collaborating with engineering, operations and project stakeholders. #J-18808-Ljbffr ...

SRE Engineer – Hybrid Cloud Reliability & Automation

Location
Greater London, England, United Kingdom
ensure ISO 27001 security alignment. This hands‐on role has significant influence across engineering, security and operations teams, with a focus on automation, observability and end‐to‐end service reliability. #J-18808-Ljbffr ...