1,951 to 1,975 of 2,452 Observability Jobs in London

Agent Engineer

Location
Greater London, England, United Kingdom
users to create user-friendly solutions for complex business processes Participate in technical design reviews, planning sessions, and code reviews Contribute to infrastructure and observability practices with the Engineering team Continuously improve quality, reliability, and usability across internal platforms Support Elliptic's mission to make crypto markets safer, more transparent … framework experience such as NestJS or Express is nice to have Terraform or infrastructure-as-code experience is nice to have Datadog or similar observability platform experience is nice to have DynamoDB or other NoSQL database experience at scale is nice to have Distributed or event-driven architecture experience, including ...

Principal Software Engineer-AI

Hiring Organisation
Moody's Corporation
Location
London, UK
Employment Type
Full-time
production or in platforming (LLM or MCP gateway, agentic runtime, auth, data retrieval, eval tooling)Experience running AI systems in production at scale, including observability, cost and capacity planning, regression detection, and incident response for AI-powered applicationsExperience operating production distributed systems on AWS/Azure, with a strong grasp … reliability, observability, and incident response at scaleDeep knowledge of cloud-native technologies, serverless applications, event-driven architectures, data and inference pipelines, relational, NoSQL, and vector databases, and modern software architecture patternsProven track record of owning multi-year technical strategy and architectural roadmaps, guiding teams from AI prototype through production deployment ...

Azure Systems Engineer

Location
Greater London, England, United Kingdom
integrations and production environments Implementing Infrastructure as Code (IaC) and automation solutions Supporting identity, access management, security and governance across Azure Improving platform monitoring, observability, resilience and performance Supporting SQL Server and Azure SQL environments, including access, backups and troubleshooting Contributing to vulnerability management, incident response and operational security Strong … Entra ID, IAM, RBAC and privileged access Understanding of secure application delivery, configuration and secrets management Experience with Azure Monitor, Application Insights or equivalent observability tooling Knowledge of SQL Server/Azure SQL , including access management, backups and troubleshooting Strong understanding of TCP/IP, DNS, routing, firewalls and private ...

Lead Site Reliability Engineer London

Location
Greater London, England, United Kingdom
build the future of insurance, we're hiring. Overview of the Role You will build the Site Reliability Engineering function at Zego, embedding reliability, observability and operational excellence as core engineering concerns - AI is the primary lever for doing that at scale, not a bolt on. You will … with adoption measured rather than assumed. Encode standards into automation rather than enforcing them by hand. Guardrails in CI, agents that check reliability and observability posture on pull requests, and safe defaults in shared infrastructure so the reliable path is the easy path. Build SRE capability on Zego ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
Drive automation across infrastructure, application delivery, operational processes, and platform management Establish platform standards, engineering patterns, and best practices Improve platform reliability, scalability, performance, observability, and operational efficiency Reduce engineering friction and accelerate software delivery Establish engineering principles and guardrails for security, reliability, and governance Lead complex platform initiatives from … automated software delivery Experience automating operational processes and platform lifecycle management Experience establishing repeatable, standardised engineering workflows Experience designing for resilience, fault tolerance, observability, and operational excellence Experience applying SRE principles and practices Experience in performance analysis, capacity planning, scalability engineering, and proactive reliability improvement Experience establishing service-level objectives ...

Senior Software Engineer, GenAI Platform

Location
Greater London, England, United Kingdom
technical direction across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability, and mentor engineers as you go. This role is ideal for a senior engineer who enjoys owning ambiguous, high-impact systems and pushing the cost/performance … hours and cutting inference cost by multiples — while giving product teams a clean choice across open-weight and closed-source models with reliability, fallback, observability, and cost controls built in. Build platforms that support rapid experimentation while meeting production standards for latency, scale, monitoring, SLOs, playbooks, and operational excellence. Partner ...

Oracle BSS Stack Lead

Location
Greater London, England, United Kingdom
Management controls. Drive automation across deployment, monitoring and operational processes. Implement continuous integration and continuous delivery capabilities using Jenkins, GitHub and Azure DevOps. Define observability standards using Dynatrace, Splunk and other enterprise monitoring tools. Reduce repetitive operational activity through sustainable automation. Build trusted working relationships with customer stakeholders, business teams … clearly to technical teams, business stakeholders and senior decision-makers. Experience with cloud technologies, containers and Kubernetes would be beneficial. Knowledge of monitoring and observability platforms would be beneficial. Familiarity with Oracle OSS components would be advantageous but is not essential. Not a perfect fit? Concerned you may not meet ...

Software Engineer, GPU Infrastructure- ChatGPT Engineering

Location
Greater London, England, United Kingdom
large-scale GPU infrastructure supporting ChatGPT inference. Build internal platforms, tooling, and AI-powered agents that automate fleet operations and reduce operational overhead. Improve observability, reliability, and operational efficiency across thousands of GPUs. Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response. Identify infrastructure bottlenecks and implement … software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure. Understand infrastructure observability, monitoring, capacity planning, and incident management. Enjoy identifying cross-team pain points and building reusable platforms that improve developer productivity. Are comfortable working across software ...

Go Back-End Developer

Location
Greater London, England, United Kingdom
ClickHouse analytics data warehouse Participate in design discussions: contribute to architectural decisions, write technical design docs, conduct code reviews Ensure service reliability and observability: metrics, logs, alerts, and participation in incident post-mortems Work with financial data security: authentication, anti-fraud mechanisms, and compliance with regulatory requirements Requirements 3+ years … ClickHouse Familiarity with AWS (EKS, RDS, DynamoDB), GitLab CI, ArgoCD Experience integrating with financial APIs (Bloomberg/Refinitiv, brokerage APIs, Fireblocks, etc.) Experience with observability systems (Prometheus, Grafana, VictoriaMetrics/VictoriaLogs) Working conditions Fully remote work Product development with no outstaffing — real influence on architecture and roadmap Small team, minimal ...

Senior AI Engineer (AI Platform)

Location
Greater London, England, United Kingdom
will also contribute to the production foundations needed to operate AI capabilities reliably, including LLMOps, model access patterns, prompt and agent lifecycle practices, evaluation, observability and secure enterprise integration. This is a hands‐on engineering role where the capabilities you build will be used by other teams across ASOS, helping … latency monitoring, alerting, scaling considerations and operational readiness Applying CI/CD and software engineering best practices to AI platform and agentic components Embedding observability by default, ensuring AI systems are measurable, debuggable and auditable through logs, metrics and traces Working with Cloud Infrastructure and Security teams to design secure ...

Senior Platform Software Engineer - SRE

Location
Greater London, England, United Kingdom
doing: The Senior Platform Software Engineer role in our SRE team combines software engineering practices with cloud infrastructure, distributed systems patterns, storage systems and observability to deliver on a wide range of projects - ranging from tooling to core Platform services which serve production traffic. Own the availability and performance … mission-critical services and build automation to prevent problems recurrence. Improve the system’s scalability, observability, and alerting. Build tooling to improve our platform and accelerate the overall software development. Practice sustainable incident response and blameless postmortems. Collaborate with product teams to help them tackle technical issues and design ...

Senior AI Engineer (AI Platform)

Hiring Organisation
ASOS
Location
London, UK
Employment Type
Full-time
will also contribute to the production foundations needed to operate AI capabilities reliably, including LLMOps, model access patterns, prompt and agent lifecycle practices, evaluation, observability and secure enterprise integration. This is a hands-on engineering role where the capabilities you build will be used by other teams across ASOS, helping … systems, including latency monitoring, alerting, scaling considerations and operational readinessApplying CI/CD and software engineering best practices to AI platform and agentic componentsEmbedding observability by default, ensuring AI systems are measurable, debuggable and auditable through logs, metrics and tracesWorking with Cloud Infrastructure and Security teams to design secure, scalable ...

Principal Network Engineer

Location
Greater London, England, United Kingdom
across firewalls, NAT, VPN, security policies, and multi-tenant segmentation. Design highly available and scalable security architectures appropriate for mission-critical AI infrastructure. Reliability, Observability & Operations Lead complex technical escalations and root-cause analysis for network performance, reliability, and stability issues. Establish measurable SLOs and operational standards for network services. … technical direction for network observability, telemetry, monitoring, and alerting. Ensure clear visibility into fabric health, traffic patterns, performance, and capacity. Develop runbooks, automation, and engineering improvements that systematically reduce operational toil. Act as a senior 3rd/4th line escalation point for complex networking issues. Network Data & Configuration Management Ensure ...

Senior DevOps Engineer

Location
Greater London, England, United Kingdom
Senior DevOps Engineer to support the team that keeps ThreatAware's platform running and shipping. You'll own AWS infrastructure, CI/CD pipelines, observability, and platform reliability — and define the developer workflow end-to-end, pioneering how AI can streamline every step. You report to the CTO. … cost optimisation. You know this space and you care about getting it right. Define the developer workflow and keep improving it. Build the tooling, observability, and automation that let engineers focus on code, not friction. Pioneer AI-assisted DevOps — explore and implement ways AI can streamline builds, deployments, monitoring ...

Senior Microsoft Power Platform and AI Developer

Hiring Organisation
Methods Consulting
Location
London, UK
Employment Type
Full-time
agents use trusted knowledge, approved tools, Model Context Protocol (MCP) services, approvals, human hand-offs and safe failure paths. You will ensure that evaluation, observability, identity, permissions and data boundaries are built into delivery rather than added later. You will review code and designs, mentor developers, improve engineering practices … C# or Python, with experience extending low-code services appropriately. Strong software engineering practice, including testing, version control, code review, documentation, CI/CD, observability and supporting live services. Ability to lead technical decisions for medium-to-high complexity work, communicate trade-offs and escalate architecture or security risks appropriately. ...

Senior Microsoft Power Platform and AI Developer

Location
Greater London, England, United Kingdom
agents use trusted knowledge, approved tools, Model Context Protocol (MCP) services, approvals, human hand-offsand safe failure paths. You will ensure that evaluation, observability, identity,permissionsand data boundaries are built into delivery rather than added later. You will review code and designs, mentor developers, improve engineeringpracticesand support live services. … JavaScript,C#or Python, with experience extending low-code services appropriately. Strong software engineering practice, including testing, version control, code review, documentation, CI/CD, observability and supporting live services. Ability to lead technical decisions for medium-to-high complexity work, communicate trade-offs and escalated architecture or security risks appropriately. ...

Sales Specialist (UK/I) - DevOps & DevEx

Hiring Organisation
Adaptavist
Location
London, UK
Employment Type
Full-time
secured, deployed and operated across UK/I.This role acts as a specialist advisor focused on Developer Experience, Platform Engineering, DevSecOps, Cloud Native Engineering, Observability and AI-enabled software delivery. This role requires a deep understanding of modern software engineering practices, developer productivity, platform operating models and software delivery transformation. … aligned to DevOps, Developer Experience and software delivery transformation initiatives. Identify, qualify and pursue opportunities across Developer Experience, Platform Engineering, DevSecOps, Cloud Native Engineering, Observability, and AI-assisted Software Delivery and Software Supply Chain Security. Develop customer-specific hypotheses, transformation opportunities and executive points of view that align engineering challenges ...

Software Engineer - Calibration

Hiring Organisation
Woven by Toyota
Location
London, UK
Employment Type
Full-time
data correctness and coverage, (4) Managing the teams calibration and validation services, ensuring robust and predictable operation, and (5) Improving the maintainability, reliability, and observability of these systems, (6) Maintaining and improving the team repositories, with an emphasis on simplification and long-term code health, (7) Evolving the teams Vision … scalability. Experience managing and maintaining software repositories, including code reviews, dependency management, versioning, and enforcing contribution standards across multiple teams. Experience with monitoring and observability tooling to maintain and improve production service health. The ability to thrive in a fast-paced environment and collaborate effectively across teams and disciplines, supported ...

Data Architect / Engineering Lead

Location
Greater London, England, United Kingdom
deployment predictable through Git-based development, CI/CD for the warehouse and semantic layer, automated provisioning and consistent environments. Set standards for testing, observability, backfills, quarantine and the day-to-day operation of analytics as software. Lead the Data Engineering discipline through standards, technical direction, mentoring, hiring input … attribute-level security, PII classification and GDPR-sensitive data. Experience treating analytics as software, including Git, CI/CD, automated testing, environment parity, observability, SLOs and cost awareness. Functional or discipline leadership experience, setting standards, raising the technical bar and mentoring without relying on formal line management. The ability ...

AI Engineer IRC302970

Location
Greater London, England, United Kingdom
products. As a Senior/Lead AI Engineer, you will own and drive the development of our core agentic frameworks, evaluation pipelines, and observability tooling to ensure they operate with safety, trust, and intelligence at scale. You won’t just be integrating AI into a product, you will lead … CrewAI is a plus. Production Python engineering: Clean, modular, testable, maintainable code — you care about system reliability as much as model output. Evaluation & observability: Practical experience instrumenting tracing (LangSmith, Arize) and building CI/CD pipelines built specifically for LLMs. Engineering foundations: Distributed systems, scalable data pipelines (Kafka ...

Lead Quality Engineer

Location
Greater London, England, United Kingdom
Lead QA at Waracle, you will play a pivotal leadership role, shaping test strategies across multiple squads, establishing automation architectures, and championing observability and non-functional requirements. Operating at a strategic level, you’ll mentor engineers, build strong relationships with senior stakeholders, and influence the adoption of forward-thinking testing … platforms. Quality & Automation Architecture: Defining squad-level automation approaches and embedding robust quality gates within CI/CD pipelines to elevate software standards. NFR & Observability Championing: Defining comprehensive strategies for performance, security, and accessibility, ensuring systems are inherently testable and debuggable in production. Delivery & Team Wellbeing: Leading QA workstreams with ...

Staff Backend Engineer - Data Platform

Location
Greater London, England, United Kingdom
drive the technical vision and implementation for our foundational data platform — from experimentation, event ingestion pipelines to our data lake, governance frameworks, and data observability and real-time analytics capabilities. You will work closely with data scientists, machine learning engineers, backend teams, and product leaders, providing deep technical expertise … analytics. Champion engineering excellence, setting high technical standards and advocating for best practices in system design, maintainability, performance, and privacy. Lead efforts in data observability, governance, and privacy-by-design principles, ensuring their robust implementation across the organization. Mentor and coach engineers, elevating the technical capabilities of the team ...

Engineering Manager - Customer and Claims · London Office ·

Location
Greater London, England, United Kingdom
every day. Identify opportunities to simplify systems, reduce technical debt and remove fragile, hard-to-change configuration. Help shape Homeprotect's engineering standards, tooling, observability and platform capabilities as the organisation continues its product engineering transformation. Promote modern software engineering practices including CI/CD, Infrastructure as Code, automated testing … observability and secure software development. KNOWLEDGE, SKILLS & EXPERIENCE #J-18808-Ljbffr ...

Software Reliability Engineer

Location
Greater London, England, United Kingdom
secrets management, authentication and authorisation Simplifying and automating application deployment processes, including automated database changes with Liquibase and installed software with Ansible. Introducing standard observability patterns. Overhauling exception handling and logging. Required Qualifications : Several years of experience with Test-Driven Development (TDD) using multiple test frameworks. Really excellent understanding … languages and/or TypeScript. Highest quality coding skills An excellent understanding of safe practices for critical systems, including deployment architecture and observability At home with multiple continuous integration and deployment systems Excellent understanding of secure coding practices Competent with Docker and Openshift Actively embracing AI coding Comfortable with build ...

Technical Lead

Location
Greater London, England, United Kingdom
Support engineers through complex technical challenges without becoming the decision-maker or implementation owner for every issue. Promote strong engineering practices including testing, automation, observability, security, documentation and sustainable software development. Encourage constructive technical challenge, knowledge sharing and continuous learning across the team. Work with the Engineering Manager to identify … sustainability of the Marketing & Commercial Data capabilities throughout their lifecycle. Ensure solutions are designed and engineered with appropriate consideration for security, resilience, scalability, observability, maintainability and supportability. Work with the Engineering Manager and Commercial Platform team to ensure new and changed capabilities are operationally ready and can be effectively supported ...