126 to 150 of 156 Observability Jobs in Central London

Senior Software Engineer

Hiring Organisation
The Portfolio Group
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£90000/annum
Establish automated testing across backend and frontend applications, including unit, contract and end-to-end testing. Work with the platform engineering team on deployment, observability, logging, tracing and operational readiness. Act as the technical owner for the application and integration layer, making and documenting key architectural decisions. Provide technical guidance … such as Lambda, ECS, API Gateway, S3, CloudFront, Cognito and IAM. Experience designing and operating distributed or event-driven systems. A strong understanding of observability, testing and CI/CD practices. Experience working with data platforms or stores such as MongoDB, OpenSearch or Databricks. Experience integrating internal systems and third ...

Staff Python Engineer (ML)

Location
City Of London, England, United Kingdom
apps in a service architecture. Furthering Developer Experience (DevEx) by mentoring others in writing code that is intuitive, clear, and easy to test Developing observability for new and existing ML applications and GenAI/LLM integrations , making use of the Grafana Stack (Prometheus, Loki, Tempo) Develop integrations and services that … Backend-Engineering Experience owning projects from start to finish, including speccing, architecture, development, testing, deployment, release and monitoring Strong skills in building maintainable tests, observability and tracing systems. Knowledge of best practices for performance optimisation, memory management. Familiarity with Kubernetes , Docker and other cloud infrastructure, ops and containerised tools. Strong ...

Principal Software Engineer-AI

Location
Westminster, West End, United Kingdom
production or in platforming (LLM or MCP gateway, agentic runtime, auth, data retrieval, eval tooling) Experience running AI systems in production at scale, including observability, cost and capacity planning, regression detection, and incident response for AI-powered applications Experience operating production distributed systems on AWS/Azure, with a strong … grasp of reliability, observability, and incident response at scale Deep knowledge of cloud-native technologies, serverless applications, event-driven architectures, data and inference pipelines, relational, NoSQL, and vector databases, and modern software architecture patterns Proven track record of owning multi-year technical strategy and architectural roadmaps, guiding teams from ...

Product Associate - SRE Team - Chase UK

Location
Westminster, West End, United Kingdom
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … services are designed, delivered, and operated with reliability in mind. Job responsibilities Support the product strategy and delivery of reliability capabilities, including standards, observability, incident practices, automation, and developer experience improvements. Partner with engineers, site reliability engineers, and cross-functional teams to understand problems, gather requirements, and translate ideas into ...

Platform Principal Engineer

Location
City Of London, England, United Kingdom
self-service capabilities. Upskill and Mentor: Transition the in-house engineering team into a high-performing internal platform team throughout the platform build process. Observability: Design and implement enterprise-grade logging, metrics, and tracing for Kubernetes at scale. IaC Leadership: Implement and manage Infrastructure as Code to a senior standard … Terraform/Open Tofu module design. (MUST) Kubernetes Engineering: GitOps (Argo CD/Flux), secrets management, ingress/mesh, and OPA/Gatekeeper. (MUST) Observability: OpenTelemetry (MUST) Tooling: Spacelift, Atlantis, or Terraform Cloud (Desired) Governance: EPAC (Enterprise Policy as Code) (Desired) What You'll Bring To Us: Recent, hands ...

Agentic AI Forward Deployed Engineer

Hiring Organisation
HCLTech
Location
City of London, London, United Kingdom
harnesses and test suites that measure agent correctness, safety and regression before anything ships. Own AgentOps/DevSecOps: CI/CD for agents, versioning, observability and telemetry, shift-left security, and Responsible AI governance baked in from day one. Run a continuous, adaptable feedback loop: feed production telemetry, evals … Eval-driven development: designing evaluation harnesses and measuring agent quality, safety and reliability. Standards-based integration and DevSecOps: APIs, secure auth, CI/CD, observability and AgentOps. Ability to conceptualize a business problem as an agent quickly, and operate effectively in ambiguous, customer-embedded settings. Client-facing maturity: translates fluidly ...

Senior UI Platform Engineer, London

Location
City Of London, England, United Kingdom
developer workflows Help define standards for application structure, testing, deployment, and supportability Partner with other engineering teams on shared concerns such as authentication, authorization, observability, and frontend/backend integration patterns Contribute to internal platform applications such as onboarding, access administration, and service discovery experiences Evaluate and maintain third-party … internal tooling Experience with Vite, Next.js, or both Experience with UI/component libraries such as Ant Design, AG Grid, or similar Experience with observability tooling such as Datadog, OpenTelemetry, or Grafana Experience with Microsoft Entra ID/Azure AD or similar identity platforms Experience publishing and maintaining internal ...

Senior Pre-Sales Consultant, Enterprise Payments

Location
City Of London, England, United Kingdom
workshops, both virtually and on-site. Lead and contribute to RFI and RFP responses, providing detailed technical input across areas including: Integration Security Resiliency Observability Compliance Identify risks, gaps, constraints, and assumptions while clearly articulating solution trade-offs. Plan, design, and support Proof of Concept (POC) activities, enabling customers … identity frameworks, including OpenID Connect and OAuth Cloud platforms such as: AWS Microsoft Azure Google Cloud Platform (GCP) IBM Cloud Oracle Cloud Monitoring and observability tools, including Splunk and SIEM solutions Pre-Sales and Customer-Facing Experience Proven experience in a customer-facing technical role, solution architecture role ...

Dataiku Solution Architect

Hiring Organisation
Everforth Quinnox
Location
City of London, London, United Kingdom
Employment Type
Permanent
dashboards use controlled, reconciled, and traceable data from the governed platform. Define access, refresh, performance, lineage, and reconciliation standards for reporting solutions. Support curve observability and the monitoring of data quality, source availability, processing status, and workflow completion. Nonfunctional Architecture Define nonfunctional requirements for performance, scalability, security, availability, resiliency, recoverability … observability, maintainability, and supportability. Design monitoring and alerting across Dataiku, Power Automate, PostgreSQL or Amazon RDS, integrations, WebApps, and reporting components. Establish recovery patterns for failed source deliveries, workflow errors, data-quality issues, integration failures, and interrupted processing. Define capacity and performance considerations for regional processing, historical replay, concurrent users ...

Vice President, Production Services Application Support

Location
Westminster, West End, United Kingdom
risk while improving platform stability. Drive initiatives through to completion with strong ownership, accountability, urgency, and quality. Champion an automation-first mindset, leveraging AI, observability, and tooling to reduce manual effort and improve service quality. Identify systemic issues and drive sustainable remediation through process simplification, platform improvements, and close partnership … incident, problem, and change management with measurable improvements in stability and service recovery. Demonstrated automation-first and AI-enabled mindset, with experience driving tooling, observability, and process automation. Strong ownership mentality and execution focus, with the ability to take initiatives from concept through delivery and embed sustainable outcomes. Deep technical ...

Backend Java Developer – Data Fabric / Platform Engineering

Location
City Of London, England, United Kingdom
platforms/query engines (e.g., Starburst or similar) Own API contracts with living documentation in CI/CD Build production-grade, testable pipelines Drive observability, reliability, and performance Contribute to architecture decisions (modularity, DI, extensibility) What You Bring (Must-Have) Strong hands-on experience in Java (17/21) + … backend performance Production-grade testing using JUnit 5, Mockito Experience with clean architecture, DI, modular design Comfortable owning CI/CD, code quality, observability Familiarity with Docker, Maven, Jenkins ⭐ Nice to Have Apache Calcite Starburst or federated query engines JVM performance tuning High-throughput service interfaces (REST/gRPC) Data ...

Staff Security Engineer

Location
City Of London, England, United Kingdom
engineer who enjoys solving complex security challenges at scale. You’ll work across engineering, data, AI and digital workplace teams to build security observability, automate control assurance and influence how security is embedded into products, platforms and processes. If you're passionate about turning security data into actionable insight, building … raising the security maturity of a fast-moving technology organisation, we'd love to hear from you. About the role Designing and building security observability capabilities that provide meaningful visibility across systems, infrastructure and applications Developing automated control monitoring, evidence collection and continuous testing solutions that strengthen security governance Partnering ...

Staff Software Engineer - Customer Data Platform

Location
City of Westminster, England, United Kingdom
outcomes. Work with adjacent teams when needed to align on shared components and dependencies. Actively participate in on-call support, contributing to operational stability, observability, and performance. Coach and support other engineers through pairing, reviews, and mentoring, helping raise the team’s overall capability. Contribute to team OKRs and actively … that balance speed, maintainability, and long-term scalability. You actively identify technical debt, risks, or inefficiencies and take action to address them. You use observability and metrics to validate behavior, debug issues, and improve system health. You regularly support and unblock teammates, helping them deliver more effectively. You contribute ...

Senior Security Engineer: Observability, Automation & Risk

Location
City Of London, England, United Kingdom
seeking a Staff Security Engineer to advance security observability, automate controls and strengthen governance across its global tech estate. This senior individual contributor role partners with engineering, data, AI and digital workplace teams to raise security maturity in a fast-moving environment. You will design scalable security observability, automate assurance ...

Principal Data Engineer

Location
City Of London, England, United Kingdom
follow the identical pattern so they are handover-ready by design. Drive data quality as a first-class, firm-wide concern: establish data contracts, observability, SLA/SLO monitoring, and automated alerting and remediation across ingestion and transformation layers, and hold squads to those standards. Act as the senior technical … with the ability to set standards, conduct code and design reviews, and grow engineers’ capabilities Strong grasp of data quality practices: data contracts, pipeline observability, SLA/SLO definition, and automated alerting and remediation Solid understanding of SQL transformation patterns and modern tooling such as dbt, alongside experience managing ingestion ...

Senior Software Engineer

Location
City Of London, England, United Kingdom
Impact and Responsibilities We are seeking a Senior Software Engineer who thrives on untangling complex systems and modernising core infrastructure without breaking production. This is an exciting opportunity to modernise core C#/SQL systems ...

Lead Backend Engineer

Location
City Of London, England, United Kingdom
About us At Zego, we know that traditional motor insurance holds good drivers back. It’s too complicated, too expensive, and it doesn't take into account how well you actually drive. That’s why ...

Lead Python Backend Engineer - Processing & Observability

Location
City Of London, England, United Kingdom
seeking a skilled software developer to enhance their data orchestration pipeline using Python. Join the Data Intelligence Team to work on AI features, develop observability for ML applications, and mentor peers. The role offers remote flexibility, a supportive work environment, and generous compensation along with various benefits. Ideal candidates have ...

Senior DevOps Engineer - ELK

Hiring Organisation
ECS
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
£500 - £700 per day
become available to join one of the world's leading technology organisations as a Senior DevOps Engineer, helping to deliver and scale critical observability, automation, and platform engineering solutions across a complex enterprise environment. As a Senior DevOps Engineer, you will be responsible for: Designing, deploying, and optimising large-scale … Elasticsearch environments. Leading improvements to monitoring, logging, and observability platforms. Driving automation and infrastructure-as-code best practices. Supporting and enhancing Kubernetes and GitOps deployments. Troubleshooting complex performance and reliability issues. Collaborating with technical teams to improve platform scalability, resilience, and security. Promoting DevOps best practices across engineering teams. Requirements ...

Staff ML Engineer | Agentic AI & Applied ML | London (Hybrid) | Contract | Inside IR35

Location
City Of London, England, United Kingdom
implementation patterns Designing and evolving production RAG and retrieval architectures Establishing effective LangGraph/LangChain patterns for agentic applications Improving AI evaluation, testing, observability and production monitoring Developing guardrails, controls and approaches to hallucination and model risk Supporting the move towards increasingly high-risk and high-complexity AI/… based applications Retrieval Augmented Generation (RAG) LangChain and/or LangGraph Vector databases and retrieval MLOps and production deployment AI evaluation, testing and observability AI governance, model risk and engineering controls ML frameworks such as PyTorch, TensorFlow or Scikit-learn Experience operating in complex, regulated or high-risk environments would ...

Senior Software Development Engineer

Location
City Of London, England, United Kingdom
expectations and can be reused across multiple brands and platforms. Drive engineering excellence for the services you own by championing code quality, automated testing, observability, performance optimization, and simplification, taking technical responsibility for service health, scalability, resilience, and the ongoing reduction of technical debt and operational overhead. Provide technical mentorship … integrations, and communicating trade-offs to both technical and non-technical stakeholders. Track record of improving operational excellence at the team level through enhanced observability, automation, performance tuning, and data-driven analysis of incidents and customer impact. Hands-on experience integrating or consuming AI/ML-enabled services or platforms ...

Senior Principal Software Engineer Dev O

Location
City Of London, England, United Kingdom
partners, the SPSE will combine hands‐on technical leadership with strategic influence. They will shape and deliver cross‐cutting improvements spanning operational maturity, resilience, observability, DevSecOps and application security, including the adoption of Application Security Posture Management (ASPM) capabilities, while exploring Agentic Dev Operations and AI‐enabled automation to improve … that standardise approaches, simplify the estate, drive efficiency and close capability gaps. Define and embed pragmatic standards, guardrails and engineering practices for reliability, resilience, observability, service performance and operational readiness, balancing consistency with the needs of autonomous engineering teams. Provide technical leadership for the ongoing DevSecOps programme, integrating security into ...

Principal AI Engineer

Hiring Organisation
Intellias
Location
City of London, London, United Kingdom
Our client is a leading global investment management firm headquartered in London, managing over $228B in assets. Technology, data science, machine learning, and AI are at the heart of its investment and research ecosystem. The ...

Databricks Champion Architect

Location
City Of London, England, United Kingdom
We’re hiring a Databricks Champion Architect to define and govern our modern lakehouse architecture. You’ll blend solution design, hands‐on technical leadership, and platform enablement; setting patterns, accelerating delivery teams, and ensuring production ...

IBM Netcool / Observability Technical Lead

Hiring Organisation
Deerfoot Recruitment Solutions
Location
City, London, United Kingdom
Employment Type
Contract
Contract Rate
GBP 780 - 830 Daily
Netcool/Observability Technical Lead Inside IR35 Contract -up to £827pd London Hybrid - 4 Days Onsite/1 Day WFH per Week Banking Are you the person who knows exactly why an ObjectServer failover didn't behave as expected, and how to stop a flood of duplicate events before anyone ...