1 to 25 of 80 Observability Jobs in Central London

Senior Observability Engineer

Location
City Of London, England, United Kingdom
Title Senior Observability Engineer Job Description Senior Observability Engineer Location London Employment type Permanent, Full Time Reporting into Senior Engineering Manager - SRE and Observability About IG Group IG Group (LSE: IGG) is a leading global fintech company, established in 1974 and headquartered in London. As a constituent of the FTSE … role IG Group’s systems move billions of dollars every day - and our clients expect them to be fast, reliable, and transparent. As an Observability Engineer, you will own the platforms and practices that give IG's engineering teams deep, real-time insight into how those systems behave. This ...

Production Engineer

Location
City Of London, England, United Kingdom
identify patterns, and drive intelligent automation solutions Hands-on experience with containerization technologies such as Docker and Kubernetes, including cluster management, deployment, scaling, and observability Deep practical knowledge of algorithmic trading workflows, including the behaviour, lifecycle, and risk controls of execution algos used across the EMEA markets Experience designing ...

Cloud Engineer

Hiring Organisation
Anson Mccade
Location
Central London, London, United Kingdom
Employment Type
Permanent, Work From Home
DevOps Engineer. Microservices architecture and API design. Docker and container platforms including Kubernetes, Amazon EKS or Amazon ECS. AWS operational concepts, including monitoring, observability and FinOps. CI/CD tooling such as Jenkins, Bamboo, TeamCity or Bitbucket. Automated testing frameworks and continuous testing practices. Linux. Networking, routing and firewalls. Cloud ...

Principal Software Engineer - Platform Engineering - Accelerator Business

Location
Westminster, West End, United Kingdom
Preferred qualifications, capabilities, and skills Advanced knowledge ofCI/CD, application resiliency, and secure delivery (e.g., SLSA framework and GitOps). Deep experience with Observability and Monitoring tools (e.g., Prometheus, Grafana, OTEL). Expertise in performance optimisation of distributed systems (e.g., caching, network latency). Practical experience with Service Mesh ...

Lead Site Reliability Engineer

Location
City of Westminster, England, United Kingdom
Fluency in Python & deep knowledge of software applications and technical processes with emerging depth in one or more technical disciplines Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, etc. Proficiency in continuous ...

Senior Software Engineer

Hiring Organisation
Method-Resourcing
Location
City of London, London, United Kingdom
Employment Type
Permanent
distributed systems where availability and performance are critical Identifying and solving complex performance and scalability challenges Contributing to architectural and technical decisions Improving resilience, observability and overall system reliability Working closely with engineers and technical stakeholders to deliver high-quality software Supporting infrastructure, deployment and continuous improvement The Technology ...

Senior AI Engineer

Location
City Of London, England, United Kingdom
agentic AI solutions to production standard within a defined technical approach. Implement and test tool use, retrieval pipelines, and agent workflows. Contribute to evaluation, observability and guardrails for agentic systems, identifying operational requirements for logging, instrumentation and alerting. Instrument solutions to capture outcome data against established baselines. Integrate AI capabilities ...

Machine Learning Engineering Lead

Hiring Organisation
Hackajob Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent
maintainability, reliability, and production support. Strong understanding of ML engineering and MLOps practices, including model lifecycle management, CI/CD, testing, monitoring, release management, observability, and operational support. Practical experience with LLM-based capabilities, including retrieval-augmented generation, semantic search, embeddings, prompt design, evaluation, guardrails, and observability. Experience with agentic ...

Security Platform Engineer, UK Security Operations

Location
Westminster, West End, United Kingdom
Python, Go, or Bash. Experience with Kubernetes security, including workload isolation, Role-Based Access Control (RBAC), and network policies, containerisation, orchestration, and Kubernetes observability tools (e.g., Falco, Prometheus, Grafana). Experience with infrastructure-as-code and configuration management tools (e.g., Terraform, Helm, ArgoCD). Active, or the ability to obtain ...

Expert Associate Partner, Architecture (Cybersecurity)

Location
City of Westminster, England, United Kingdom
security incident through architecture choices, not just tooling. Data security: experience with DLP, data classification, and cryptographic controls across hybrid environments. Logging and observability: clear understanding of what good telemetry looks like from a security operations perspective, and able to design architectures to produce it. Able to speak ...

Lead Agentic AI Engineer

Location
City Of London, England, United Kingdom
implement evaluation frameworks for quality, grounding, task success, safety, latency, and cost Experience deploying AI solutions into production with focus on cost optimisation, scalability, observability, security, and governance Experienced in using Microsoft GenAI ecosystem, including M365, Copilot Studio, Azure AI Foundry, and Microsoft Agent Framework Proficient in Python and production ...

Site Reliability Engineer, Infrastructure - ThousandEyes

Location
City Of London, England, United Kingdom
user experiences. ThousandEyes is deeply integrated across the Cisco technology portfolio, delivering AI-powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios. Our distributed Site Reliability Engineering team of approximately nine engineers owns the availability, latency, performance, efficiency, monitoring, emergency response, and capacity planning of the platform ...

Senior Front-End Engineer React

Hiring Organisation
Halian Technology Limited
Location
Central London, London, United Kingdom
Employment Type
Permanent
Salary
£90,000
validation. Take ownership of functionality from concept through to production. Define and influence frontend architecture, documenting and presenting technical decisions. Ensure security, reliability and observability are considered from the outset. Build highly accessible, responsive and user-friendly interfaces. Integrate with third-party platforms including identity verification, document management ...

Associate Director Lead AI Architect

Hiring Organisation
Anson Mccade
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
services . Work with self-hosted inference and model-serving technologies where appropriate, including tools such as vLLM, SGLang and Ollama . Design evaluation, observability and monitoring frameworks to assess model performance, system health, reliability and business impact. Optimise AI solutions for performance, scalability, security and cost-efficiency. Identify where ...

Senior C# Developer .NET 9/10, AWS Lambda, SQS, SNS Outside IR35

Hiring Organisation
Smart Sourcer Limited
Location
City of London, London, United Kingdom
Employment Type
Contract
Contract Rate
From £700 to £800 per day
versioned, well-documented APIs that other engineers love to consume. Championing a high-quality engineering culture: testing discipline, peer review, CI/CD excellence, observability, performance and secure coding. Developing hands-on skills with Claude Code , including AI-assisted development patterns, smart refactoring, automated context management and productivity-boosting workflows. ...

Lead Full Stack Engineer

Location
Westminster, West End, United Kingdom
while keeping strong human accountability for high risk decisions, verification, and final quality Build software with a reliability first mindset, incorporating secure coding practices, observability, testing, and disciplined rollout/rollback strategies Ensure production systems meet required levels of resiliency, performance, scalability, and operability To be successful in this role ...

Head of Engineering (Convo AI)

Location
City Of London, England, United Kingdom
clear quality standards and performance expectations across teams. As a senior engineering leader, you will drive engineering excellence across software engineering, AI engineering, testing, observability, reliability and operational excellence. You will lead the delivery of Conversational AI products, AI-powered servicing journeys, intelligent assistants and agentic customer experiences, ensuring solutions ...

Lead DevOps Engineer

Location
City Of London, England, United Kingdom
partnership with engineering, infrastructure, security, and operations teams. It is an opportunity to improve reliability, scalability, security, and delivery while advancing DevOps, platform engineering, observability, and AI-enabled infrastructure tooling. Responsibilities Lead the design and evolution of scalable, secure, and highly available trading infrastructure. Provide technical direction, mentorship, and guidance … assisted engineering tools to accelerate development, automation, troubleshooting, documentation, and operational workflows. Identify, design, and help implement AI-enabled operational capabilities for infrastructure automation, observability, incident response, and platform engineering. Contribute to the strategy for safe, practical, and secure adoption of AI across infrastructure and DevOps practices. Lead and participate ...

Lead Identity and Security Engineer - 12 Month FTC

Location
Westminster, West End, United Kingdom
across identity and security. Work collaboratively with architecture, platform, application engineering and security teams to solve complex identity challenges. Contribute to the reliability, scalability, observability and operational maturity of identity services. Produce, maintain and review technical documentation, architecture decisions, standards and security guidance. Identify and address security risks and continuously ...

Senior Data Platform Engineer (Data Lake and Catalog)

Location
City Of London, England, United Kingdom
example in engineering quality through clean, well-tested code and strong coding practices. Promote an SLO-driven culture by contributing to reliability goals, observability, and incident learnings. Partner closely with engineers and other stakeholders to clarify requirements and deliver effectively. Break down complex problems into pragmatic technical solutions, estimates ...

Software Engineering Manager

Location
Westminster, West End, United Kingdom
gateway, agentic runtime, auth, data retrieval, eval tooling) Experience operating production distributed systems on AWS/Azure, with a strong grasp of reliability, observability, and incident response at scale Excellent communication skills, with the ability to build consensus and translate technical topics for engineers and executives alike Strong sense ...

Senior Forward Deployed Engineer, Google Cloud Public Sector

Location
City of Westminster, England, United Kingdom
Google’s AI products and customer's live infrastructure, including APIs, legacy data silos, and security perimeters. Build high-performance evaluation (Eval) pipelines and observability frameworks to ensure agentic systems meet requirements for accuracy, safety, and latency. Identify repeatable field patterns and technical friction points in Google’s AI stack ...

Senior Frontend Engineer

Location
City Of London, England, United Kingdom
software works Make architectural calls on your domain, record down, present and defend them in front of the tribe. Build in security, reliability and observability from the start. Use AI tooling well. We use Claude Code daily and we expect judgement about where it earns its place and where ...

Vice President, Site Reliability Engineering

Location
Westminster, West End, United Kingdom
implement, and continuously improve Service Level Indicators, Service Level Objectives, and service health measures aligned to operational and business priorities. Build and optimize monitoring, observability, and alerting capabilities using tools such as Prometheus, Grafana, AppDynamics, and Splunk. Apply AIOps capabilities to improve event correlation, anomaly detection, root cause analysis, predictive … enterprise or production environments. Demonstrated ability to define and operationalize SLIs, SLOs, dashboards, alerts, and health indicators. Hands-on experience with enterprise monitoring and observability platforms including Prometheus, Grafana, AppDynamics, and Splunk. Strong troubleshooting, analytical, and problem-solving skills in complex distributed or production environments. Strong verbal and written communication ...

Production Engineer

Location
City Of London, England, United Kingdom
Group OverviewThe TP ICAP Group is a world leading provider of market infrastructure.Our purpose is to provide clients with access to global financial and commodities markets, improving price discovery, liquidity, and distribution of data, through ...