1,801 to 1,825 of 2,250 Observability Jobs in London

Senior Platform Software Engineer - SRE

Location
Greater London, England, United Kingdom
doing: The Senior Platform Software Engineer role in our SRE team combines software engineering practices with cloud infrastructure, distributed systems patterns, storage systems and observability to deliver on a wide range of projects - ranging from tooling to core Platform services which serve production traffic. Own the availability and performance … mission-critical services and build automation to prevent problems recurrence. Improve the system’s scalability, observability, and alerting. Build tooling to improve our platform and accelerate the overall software development. Practice sustainable incident response and blameless postmortems. Collaborate with product teams to help them tackle technical issues and design ...

Senior AI Engineer (AI Platform)

Hiring Organisation
ASOS
Location
London, UK
Employment Type
Full-time
will also contribute to the production foundations needed to operate AI capabilities reliably, including LLMOps, model access patterns, prompt and agent lifecycle practices, evaluation, observability and secure enterprise integration. This is a hands-on engineering role where the capabilities you build will be used by other teams across ASOS, helping … systems, including latency monitoring, alerting, scaling considerations and operational readinessApplying CI/CD and software engineering best practices to AI platform and agentic componentsEmbedding observability by default, ensuring AI systems are measurable, debuggable and auditable through logs, metrics and tracesWorking with Cloud Infrastructure and Security teams to design secure, scalable ...

Principal Network Engineer

Location
Greater London, England, United Kingdom
across firewalls, NAT, VPN, security policies, and multi-tenant segmentation. Design highly available and scalable security architectures appropriate for mission-critical AI infrastructure. Reliability, Observability & Operations Lead complex technical escalations and root-cause analysis for network performance, reliability, and stability issues. Establish measurable SLOs and operational standards for network services. … technical direction for network observability, telemetry, monitoring, and alerting. Ensure clear visibility into fabric health, traffic patterns, performance, and capacity. Develop runbooks, automation, and engineering improvements that systematically reduce operational toil. Act as a senior 3rd/4th line escalation point for complex networking issues. Network Data & Configuration Management Ensure ...

Senior DevOps Engineer

Location
Greater London, England, United Kingdom
Senior DevOps Engineer to support the team that keeps ThreatAware's platform running and shipping. You'll own AWS infrastructure, CI/CD pipelines, observability, and platform reliability — and define the developer workflow end-to-end, pioneering how AI can streamline every step. You report to the CTO. … cost optimisation. You know this space and you care about getting it right. Define the developer workflow and keep improving it. Build the tooling, observability, and automation that let engineers focus on code, not friction. Pioneer AI-assisted DevOps — explore and implement ways AI can streamline builds, deployments, monitoring ...

Senior Microsoft Power Platform and AI Developer

Hiring Organisation
Methods Consulting
Location
London, UK
Employment Type
Full-time
agents use trusted knowledge, approved tools, Model Context Protocol (MCP) services, approvals, human hand-offs and safe failure paths. You will ensure that evaluation, observability, identity, permissions and data boundaries are built into delivery rather than added later. You will review code and designs, mentor developers, improve engineering practices … C# or Python, with experience extending low-code services appropriately. Strong software engineering practice, including testing, version control, code review, documentation, CI/CD, observability and supporting live services. Ability to lead technical decisions for medium-to-high complexity work, communicate trade-offs and escalate architecture or security risks appropriately. ...

Senior Microsoft Power Platform and AI Developer

Location
Greater London, England, United Kingdom
agents use trusted knowledge, approved tools, Model Context Protocol (MCP) services, approvals, human hand-offsand safe failure paths. You will ensure that evaluation, observability, identity,permissionsand data boundaries are built into delivery rather than added later. You will review code and designs, mentor developers, improve engineeringpracticesand support live services. … JavaScript,C#or Python, with experience extending low-code services appropriately. Strong software engineering practice, including testing, version control, code review, documentation, CI/CD, observability and supporting live services. Ability to lead technical decisions for medium-to-high complexity work, communicate trade-offs and escalated architecture or security risks appropriately. ...

Sales Specialist (UK/I) - DevOps & DevEx

Hiring Organisation
Adaptavist
Location
London, UK
Employment Type
Full-time
secured, deployed and operated across UK/I.This role acts as a specialist advisor focused on Developer Experience, Platform Engineering, DevSecOps, Cloud Native Engineering, Observability and AI-enabled software delivery. This role requires a deep understanding of modern software engineering practices, developer productivity, platform operating models and software delivery transformation. … aligned to DevOps, Developer Experience and software delivery transformation initiatives. Identify, qualify and pursue opportunities across Developer Experience, Platform Engineering, DevSecOps, Cloud Native Engineering, Observability, and AI-assisted Software Delivery and Software Supply Chain Security. Develop customer-specific hypotheses, transformation opportunities and executive points of view that align engineering challenges ...

AI Engineer IRC302970

Location
Greater London, England, United Kingdom
products. As a Senior/Lead AI Engineer, you will own and drive the development of our core agentic frameworks, evaluation pipelines, and observability tooling to ensure they operate with safety, trust, and intelligence at scale. You won’t just be integrating AI into a product, you will lead … CrewAI is a plus. Production Python engineering: Clean, modular, testable, maintainable code — you care about system reliability as much as model output. Evaluation & observability: Practical experience instrumenting tracing (LangSmith, Arize) and building CI/CD pipelines built specifically for LLMs. Engineering foundations: Distributed systems, scalable data pipelines (Kafka ...

Lead Quality Engineer

Location
Greater London, England, United Kingdom
Lead QA at Waracle, you will play a pivotal leadership role, shaping test strategies across multiple squads, establishing automation architectures, and championing observability and non-functional requirements. Operating at a strategic level, you’ll mentor engineers, build strong relationships with senior stakeholders, and influence the adoption of forward-thinking testing … platforms. Quality & Automation Architecture: Defining squad-level automation approaches and embedding robust quality gates within CI/CD pipelines to elevate software standards. NFR & Observability Championing: Defining comprehensive strategies for performance, security, and accessibility, ensuring systems are inherently testable and debuggable in production. Delivery & Team Wellbeing: Leading QA workstreams with ...

Staff Backend Engineer - Data Platform

Location
Greater London, England, United Kingdom
drive the technical vision and implementation for our foundational data platform — from experimentation, event ingestion pipelines to our data lake, governance frameworks, and data observability and real-time analytics capabilities. You will work closely with data scientists, machine learning engineers, backend teams, and product leaders, providing deep technical expertise … analytics. Champion engineering excellence, setting high technical standards and advocating for best practices in system design, maintainability, performance, and privacy. Lead efforts in data observability, governance, and privacy-by-design principles, ensuring their robust implementation across the organization. Mentor and coach engineers, elevating the technical capabilities of the team ...

Engineering Manager - Customer and Claims · London Office ·

Location
Greater London, England, United Kingdom
every day. Identify opportunities to simplify systems, reduce technical debt and remove fragile, hard-to-change configuration. Help shape Homeprotect's engineering standards, tooling, observability and platform capabilities as the organisation continues its product engineering transformation. Promote modern software engineering practices including CI/CD, Infrastructure as Code, automated testing … observability and secure software development. KNOWLEDGE, SKILLS & EXPERIENCE #J-18808-Ljbffr ...

Software Reliability Engineer

Location
Greater London, England, United Kingdom
secrets management, authentication and authorisation Simplifying and automating application deployment processes, including automated database changes with Liquibase and installed software with Ansible. Introducing standard observability patterns. Overhauling exception handling and logging. Required Qualifications : Several years of experience with Test-Driven Development (TDD) using multiple test frameworks. Really excellent understanding … languages and/or TypeScript. Highest quality coding skills An excellent understanding of safe practices for critical systems, including deployment architecture and observability At home with multiple continuous integration and deployment systems Excellent understanding of secure coding practices Competent with Docker and Openshift Actively embracing AI coding Comfortable with build ...

Technical Lead

Location
Greater London, England, United Kingdom
Support engineers through complex technical challenges without becoming the decision-maker or implementation owner for every issue. Promote strong engineering practices including testing, automation, observability, security, documentation and sustainable software development. Encourage constructive technical challenge, knowledge sharing and continuous learning across the team. Work with the Engineering Manager to identify … sustainability of the Marketing & Commercial Data capabilities throughout their lifecycle. Ensure solutions are designed and engineered with appropriate consideration for security, resilience, scalability, observability, maintainability and supportability. Work with the Engineering Manager and Commercial Platform team to ensure new and changed capabilities are operationally ready and can be effectively supported ...

Engineering Lead

Location
Greater London, England, United Kingdom
share ownership of technical direction. Shape how the team works, not just what it builds Drive improvements to developer experience, CI/CD, observability and release practices that make the whole team faster and more confident. Make pragmatic trade-offs that balance reliability, performance and cost across AWS services. Decide … modern serverless (Lambda, API Gateway, SQS, EventBridge, DynamoDB, S3, CloudWatch) or other cloud platforms. A pragmatic approach to technical decisions, balancing reliability, observability, performance and cost, and bringing engineers along on the reasoning. Track record of raising engineering standards as a force-multiplier, through code reviews, pairing, design discussions ...

Platform Principal Engineer

Location
City Of London, England, United Kingdom
self-service capabilities. Upskill and Mentor: Transition the in-house engineering team into a high-performing internal platform team throughout the platform build process. Observability: Design and implement enterprise-grade logging, metrics, and tracing for Kubernetes at scale. IaC Leadership: Implement and manage Infrastructure as Code to a senior standard … Terraform/Open Tofu module design. (MUST) Kubernetes Engineering: GitOps (Argo CD/Flux), secrets management, ingress/mesh, and OPA/Gatekeeper. (MUST) Observability: OpenTelemetry (MUST) Tooling: Spacelift, Atlantis, or Terraform Cloud (Desired) Governance: EPAC (Enterprise Policy as Code) (Desired) What You'll Bring To Us: Recent, hands ...

AI Engineer

Location
Greater London, England, United Kingdom
full-stack AI engineering role, with applied AI product delivery at its core. You will build AI workflows, tools and integrations, evaluation and observability capabilities, together with the APIs, services and user interfaces needed to ship them reliably. This is not a research-only role. You will apply agreed enterprise … feedback loops. Manage prompts, model configuration, tool schemas and routing as tested product assets, including fallbacks and cost/latency trade‐offs. Implement AI observability through traces, logs, metrics and evaluation results, so quality, reliability, failure modes, latency and cost are visible and can be improved. Build clear APIs ...

Backend Engineers (Ruby)

Location
Greater London, England, United Kingdom
technical insight into upcoming work and helping pull the team together to ship it Deliver your work using agile methodologies and tools like tests, observability, A/B tests, and feature flags Mentor colleagues to help them grow as engineers, and actively support their development Contribute to cross-cutting concerns … Rails, building data models, APIs, and business logic services in a production monolith Comfortable delivering with agile methodologies and practices like automated testing, observability, A/B testing, and feature flags A collaborative approach — you work well with product, design, and analytics partners, and enjoy shaping work beyond just ...

Agent Engineer

Location
Greater London, England, United Kingdom
concept through to production. Take part in technical design reviews, planning sessions, and code reviews to continuously improve system quality. Contribute to infrastructure and observability practices alongside the Engineering team — you won't own this alone, but you'll be expected to care about how your services run in production. … design. Nice to Have Experience with Node.js frameworks like NestJS or Express. Hands-on experience with Terraform, or infrastructure-as-code tooling. Experience with observability platforms like Datadog (metrics, tracing, alerting). Exposure to DynamoDB or other NoSQL databases at scale. Experience with distributed or event-driven architectures ...

Embedded Software Engineer

Location
Greater London, England, United Kingdom
vulnerability management and secure provisioning. Edge & Cloud Integration Integrate devices with cloud IoT platforms and backend services. Define and improve telemetry, health monitoring, and observability for deployed devices. Support reliable operation of connected devices in the field, including debugging fleet issues and improving resilience. Compliance & Testing Support testing and validation …/CD workflows for embedded software. Practical debugging experience using lab and software tools such as logic analysers, protocol analysers, network sniffers, or observability platforms. Strong problem-solving skills and the ability to work across hardware and software boundaries. Nice to Have Experience with low-power wireless technologies such ...

Head of Data Platform

Hiring Organisation
Arch Capital Group
Location
London, UK
Employment Type
Full-time
Lead the design, build and operation of a scalable, secure and resilient modern data platform. Drive modern engineering practices across DataOps, CI/CD, observability, service management, data quality and production support. Create AI-ready data foundations by ensuring data is governed, well modelled, permission-aware and suitable for advanced … platforms, ideally including technologies such as Snowflake, orchestration tools, ELT, BI and analytics platforms. A strong understanding of platform engineering, DataOps, CI/CD, observability, reliability, resilience, data quality, security and production support. Practical knowledge of AI, machine learning and Generative AI workloads, including the data, governance, retrieval, MLOps ...

GenAI Engineer

Location
Greater London, England, United Kingdom
indexing, retrieval policies, grounding, guardrails) and agent frameworks. Take basic infra ownership on GCP (or AWS/Azure): networking, autoscaling, CI/CD, IaC, observability, and cost tuning. Participate in on‐call for your area and drive root‐cause analysis with crisp follow‐ups. 15% Collaborate Pair with back … series analysis (forecasting, change‐point, drift). Cloud & ops: Basic infra ownership on GCP (or AWS/Azure): networking, autoscaling, CI/CD, IaC, observability, and cost control. Communication: You explain results clearly, align stakeholders, and write crisp docs. Bonus points DevOps wizardry; GPU/accelerator experience. Multimodal pipelines (text ...

Managed Service Operations - Head of Practice

Location
Greater London, England, United Kingdom
service outcomes. Key responsibilities You will lead the development and maturity of Made Tech’s operational capabilities incident, problem, and change management; monitoring and observability; automation and AIOps; governance; operational playbooks; runbooks; service health metrics; and 24/7/365 support patterns. You will ensure our teams have … ITIL practices blended with modern DevOps, SRE, Agile and platform‐engineering approaches. Broad technical awareness across cloud platforms, application architectures, data platforms, networks, observability tooling, security‐by‐design, and automation. Ability to create and evolve operational standards, playbooks, governance models, templates, and frameworks that drive consistency, stability, and efficiency. Skilled ...

Senior DevOps Developer - Remote/Hybrid (Canada)

Hiring Organisation
InfoTech Research Group
Location
London, UK
Employment Type
Full-time
data store environments, including MySQL 8, PostgreSQL, MS SQL Server, and SQLite, with consideration for performance, reliability, scalability, and availability. Establish and improve reliability, observability, monitoring, and disaster recovery practices, ensuring appropriate visibility into system health, performance, and potential risks. Identify opportunities to improve automation, infrastructure efficiency, system reliability, security … MySQL 8, PostgreSQL, MS SQL Server, and SQLite. Strong experience working across Windows and Mac environments. Strong understanding and practical experience with reliability, observability, monitoring, and disaster recovery. Proven ability to troubleshoot and resolve complex issues spanning application, infrastructure, deployment, and production environments. Strong understanding of the software development lifecycle ...

Software Engineer III - Fullstack (Java, React, Python and AI) Engineer

Location
Greater London, England, United Kingdom
with product, marketing, and partners to translate requirements into well-designed technical solutions Improve engineering excellence through code reviews, test automation, CI/CD, observability, performance tuning, and operational best practices Ensure solutions meet security, privacy, and compliance expectations including consent management, data minimization, and access controls Required Qualifications, Capabilities … with campaign management and attribution systems Background in AI-enabled content generation, experimentation, or workflow automation Skills in test automation, CI/CD, and observability tools Experience optimizing performance and scalability Familiarity with consent management and data minimization practices Ability to drive innovation in personalization and measurement Experience working ...

Lead Mobile Engineer

Location
Greater London, England, United Kingdom
deployment processes, including CI/CD pipelines and app store submissions for Google Play and Apple App Store.* Champion automated testing, app stability, observability, and the responsible use of AI to improve developer productivity and software quality.**Knowledge, Skills and Behaviours**Essential* 7+ years of professional software engineering experience, with … experience using mobile testing frameworks and automated testing approaches.* Experience owning mobile CI/CD pipelines, build automation and app store release processes.* An observability mindset, with experience using crash reporting, performance monitoring and analytics tools to improve app reliability.* Comfortable working with Git, Jira, Confluence and modern agile engineering ...