1,776 to 1,800 of 2,250 Observability Jobs in London

Backend Software Engineer

Location
Greater London, England, United Kingdom
engineering, performance, and reliability problems that sit outside of their scope. Contribute to the shared foundations other engineers depend on: deployment pipelines, service templates, observability, and the environments their work runs in. You should apply if You have built and worked on production backend systems. You will spend your time … whether that means learning a new skill, reaching out to stakeholders to clarify requirements, or suggesting an alternate approach. You treat quality, security, and observability as engineering fundamentals rather than optional extras, and you flag technical debt rather than letting it accumulate silently. You actively seek feedback ...

Agentic Commerce Architect

Location
City Of London, England, United Kingdom
intelligence — including embedding pipelines, vector retrieval, semantic search, and structured product data — ensuring agent outputs are accurate, grounded, and commercially reliable Embedding AI governance, observability, and responsible AI principles into architecture design — including audit logging, human-in-the-loop escalation points, and performance monitoring — as first-class concerns rather than … deployment, including familiarity with cloud-native services and CI/CD-based delivery practices Understanding of AI governance and responsible AI principles — including observability, tracing, auditability, and how to build guardrails into production AI systems Preferred Experience Experience contributing to technical proposals, reference architectures, or delivery accelerators in a consulting ...

Application Support Engineer (Third Party Systems), London

Hiring Organisation
Isomorphic Labs
Location
London, UK
Employment Type
Full-time
Isomorphic Labs is applying frontier AI to help unlock deeper scientific insights, faster breakthroughs, and life-changing medicines with an ambition to solve all disease. The future is coming. A future enabled and enriched by ...

Data Architect

Location
Greater London, England, United Kingdom
We believe in the power of ingenuity to build a positive human future.We challenge where it matters and own the outcome.As strategies, technologies, and innovation collide, we create opportunity from complexity. Our teams of interdisciplinary ...

Data Platform Engineer

Hiring Organisation
ed Resourcing Ltd
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP 70,000 - 80,000 Annual
Data Platform Engineer £70,000 - £75,000 + 10% bonus + excellent benefits We're looking for a hands-on Data Platform Engineer who enjoys getting stuck into infrastructure, automation and platform engineering. This isn ...

Azure Platform Engineer - DevOps, IaC & Observability

Location
Greater London, England, United Kingdom
technical delivery. You will be designing, deploying and supporting secure, scalable Azure environments, building CI/CD pipelines, IaC and automation, and ensuring observability, resilience and secure governance across Azure. #J-18808-Ljbffr ...

Senior SRE Technical Lead — Reliability & Observability

Location
Greater London, England, United Kingdom
seeking a Technical Lead SRE in Greater London. In this role, you will enhance the reliability engineering capabilities, collaborating with various teams to establish observability standards and ensure operational excellence. The ideal candidate will have over 10 years of experience in SRE or related fields, strong AWS and Kubernetes skills ...

Senior Platform SRE: Reliable, Scalable Observability

Location
Greater London, England, United Kingdom
team to own availability and performance of mission-critical services, and to lead on-call and incident response. You will build tooling, improve observability, and drive platform scalability in collaboration with product teams, while lowering costs and improving developer productivity. The ideal candidate has around 8 years of distributed systems ...

Kubernetes SRE: GitOps, Canary Deployments & Observability

Location
Greater London, England, United Kingdom
secure deployments across dev, staging, and production. Embedded in a hybrid team, you will apply AI-assisted tooling to accelerate delivery and collaborate on observability, policy, and data services for #J-18808-Ljbffr ...

Observability Engineer (Dynatrace) — Telemetry & Performance

Location
Greater London, England, United Kingdom
Computacenter is seeking a Monitoring & Observability Engineer (Dynatrace) to design, implement and manage observability across customer IT estates in the UK. You will collect telemetry, diagnose issues and drive proactive improvements across teams. The role requires strong Dynatrace/Grafana/Splunk experience, scripting skills, cloud familiarity (Azure/ ...

Cloud Advisory Architecture Associate Manager

Hiring Organisation
Accenture
Location
London, UK
Employment Type
Full-time
where GenAI and Agentic play a role. Champion system performance, resilience, and efficiency: Proactively identifying and addressing consumption and scalability challenges. Champion full stack observability using modern full stack observability, SRE and AIOps. Manage & Mentor: Lead teams of architects and engineers, providing technical coaching, career counselling, performance management, and coaching ...

Service Reliability Support Manager

Location
Greater London, England, United Kingdom
batch monitoring and major incident command. It works in partnership with business-aligned Application Support teams and the Technology Centre of Excellence to embed observability engineering, telemetry, and automation — including AI-assisted triage and self-healing — across the estate. Role Summary The Service Reliability Lead owns Marex's central, cross … reactive, manual, ticket-driven model into an engineering-first Service Reliability capability, in which AI-assisted triage and automation absorb routine work and observability is owned as an engineering discipline connected to the Technology Centre of Excellence. The role provides oversight of the existing team, ensuring all current responsibilities continue ...

Head of Digital Platform Enablement

Hiring Organisation
S Merrick LTD
Location
Central London, London, United Kingdom
Employment Type
Permanent
large global enterprise. The role is effectively the Product Owner for Digital Platform Enablement and spans three core pillars: ServiceNow/service management platforms, observability/monitoring, and developer/engineering platforms. The client is moving from a traditional project-led, design-build-run model towards a product-centric, agile … Skills/Experience for the Head of Digital Platform Enablement role: Meaningful ownership of ServiceNow/service management platforms, workflows, automation and self-service Observability strategy across user experience, applications, cloud, infrastructure and networks Developer platforms and tooling including Azure DevOps and/or GitHub CI/CD, Infrastructure ...

Principal Software Engineer-AI

Location
Greater London, England, United Kingdom
production or in platforming (LLM or MCP gateway, agentic runtime, auth, data retrieval, eval tooling) Experience running AI systems in production at scale, including observability, cost and capacity planning, regression detection, and incident response for AI-powered applications Experience operating production distributed systems on AWS/Azure, with a strong … grasp of reliability, observability, and incident response at scale Deep knowledge of cloud-native technologies, serverless applications, event-driven architectures, data and inference pipelines, relational, NoSQL, and vector databases, and modern software architecture patterns Proven track record of owning multi-year technical strategy and architectural roadmaps, guiding teams from ...

Lead Network Operations Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
organisation's network and security infrastructure across datacentre and office environments. This is a hands-on technical leadership role focused on operational excellence, automation, observability, and incident response — ensuring high availability, resilience, and a strong security posture. Key responsibilities: Own the day-to-day performance, stability, availability, and security … campus environmentsLead response to critical incidents, driving rapid diagnosis, containment, and long-term remediationChampion an automation-first approach, reducing operational toil through tooling, observability, and emerging technologies (including agentic AI)Develop and refine operational tooling, runbooks, and incident response frameworks to ensure consistent, high-quality service deliverySupport an event-driven ...

Front End developer

Location
Greater London, England, United Kingdom
applications. You’ll work across Python services and React/Vite frontends , collaborating closely with engineering teams to improve application reliability, CI/CD, observability, security and developer tooling. Key experience: Docker and CI/CD pipelines Azure cloud experience is essential; AWS is a bonus Good knowledge of Linux … networking, security and observability Experience supporting scalable application platforms and cloud deployments Able to work independently and establish reusable engineering patterns and best practices A great opportunity for someone who enjoys working across development, cloud and platform engineering. Additional Information #TalanUK #J-18808-Ljbffr ...

Data Reliability Engineer

Hiring Organisation
Ashdown Group
Location
London, UK
Employment Type
Full-time
work from home 2 days per week. This is a high-impact role focused on improving data quality, reducing incidents, and building scalable observability across a modern enterprise data platform. You'll help ensure data across the organisation is accurate, reliable, and trusted for critical business decision-making. … style roles, with strong SQL and Python skills and experience working in modern cloud-based data environments. Hands-on experience with data observability tools such as Grafana, Monte Carlo, or Acceldata, and data governance/quality platforms like Informatica, Collibra or Microsoft Purview is highly desirable. Experience within the Azure ...

Senior Software Engineer – Agentic Development Enablement

Location
Greater London, England, United Kingdom
such as Claude Code and GitHub Copilot Design and implement guardrails, controls, and engineering patterns for AI-assisted development Contribute to endpoint and platform observability, telemetry, and policy enforcement Define how controls work consistently across local development environments and CI/CD pipelines Explore changes to development environments, including containerised … production Requirements Strong hands-on background as a software engineer Broad technical understanding across developer tooling, cloud platforms, operating systems, desktop environments, security controls, observability, telemetry, and CI/CD Experience working on developer workflows and engineering ways of working, not only end-user application delivery Ability to work ...

Agent Engineer

Location
Greater London, England, United Kingdom
users to create user-friendly solutions for complex business processes Participate in technical design reviews, planning sessions, and code reviews Contribute to infrastructure and observability practices with the Engineering team Continuously improve quality, reliability, and usability across internal platforms Support Elliptic's mission to make crypto markets safer, more transparent … framework experience such as NestJS or Express is nice to have Terraform or infrastructure-as-code experience is nice to have Datadog or similar observability platform experience is nice to have DynamoDB or other NoSQL database experience at scale is nice to have Distributed or event-driven architecture experience, including ...

Principal Software Engineer-AI

Hiring Organisation
Moody's Corporation
Location
London, UK
Employment Type
Full-time
production or in platforming (LLM or MCP gateway, agentic runtime, auth, data retrieval, eval tooling)Experience running AI systems in production at scale, including observability, cost and capacity planning, regression detection, and incident response for AI-powered applicationsExperience operating production distributed systems on AWS/Azure, with a strong grasp … reliability, observability, and incident response at scaleDeep knowledge of cloud-native technologies, serverless applications, event-driven architectures, data and inference pipelines, relational, NoSQL, and vector databases, and modern software architecture patternsProven track record of owning multi-year technical strategy and architectural roadmaps, guiding teams from AI prototype through production deployment ...

Azure Systems Engineer

Location
Greater London, England, United Kingdom
integrations and production environments Implementing Infrastructure as Code (IaC) and automation solutions Supporting identity, access management, security and governance across Azure Improving platform monitoring, observability, resilience and performance Supporting SQL Server and Azure SQL environments, including access, backups and troubleshooting Contributing to vulnerability management, incident response and operational security Strong … Entra ID, IAM, RBAC and privileged access Understanding of secure application delivery, configuration and secrets management Experience with Azure Monitor, Application Insights or equivalent observability tooling Knowledge of SQL Server/Azure SQL , including access management, backups and troubleshooting Strong understanding of TCP/IP, DNS, routing, firewalls and private ...

Lead Site Reliability Engineer London

Location
Greater London, England, United Kingdom
build the future of insurance, we're hiring. Overview of the Role You will build the Site Reliability Engineering function at Zego, embedding reliability, observability and operational excellence as core engineering concerns - AI is the primary lever for doing that at scale, not a bolt on. You will … with adoption measured rather than assumed. Encode standards into automation rather than enforcing them by hand. Guardrails in CI, agents that check reliability and observability posture on pull requests, and safe defaults in shared infrastructure so the reliable path is the easy path. Build SRE capability on Zego ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
Drive automation across infrastructure, application delivery, operational processes, and platform management Establish platform standards, engineering patterns, and best practices Improve platform reliability, scalability, performance, observability, and operational efficiency Reduce engineering friction and accelerate software delivery Establish engineering principles and guardrails for security, reliability, and governance Lead complex platform initiatives from … automated software delivery Experience automating operational processes and platform lifecycle management Experience establishing repeatable, standardised engineering workflows Experience designing for resilience, fault tolerance, observability, and operational excellence Experience applying SRE principles and practices Experience in performance analysis, capacity planning, scalability engineering, and proactive reliability improvement Experience establishing service-level objectives ...

Senior Software Engineer, GenAI Platform

Location
Greater London, England, United Kingdom
technical direction across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability, and mentor engineers as you go. This role is ideal for a senior engineer who enjoys owning ambiguous, high-impact systems and pushing the cost/performance … hours and cutting inference cost by multiples — while giving product teams a clean choice across open-weight and closed-source models with reliability, fallback, observability, and cost controls built in. Build platforms that support rapid experimentation while meeting production standards for latency, scale, monitoring, SLOs, playbooks, and operational excellence. Partner ...

Software Engineer, GPU Infrastructure- ChatGPT Engineering

Location
Greater London, England, United Kingdom
large-scale GPU infrastructure supporting ChatGPT inference. Build internal platforms, tooling, and AI-powered agents that automate fleet operations and reduce operational overhead. Improve observability, reliability, and operational efficiency across thousands of GPUs. Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response. Identify infrastructure bottlenecks and implement … software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure. Understand infrastructure observability, monitoring, capacity planning, and incident management. Enjoy identifying cross-team pain points and building reusable platforms that improve developer productivity. Are comfortable working across software ...