101 to 118 of 118 Observability Jobs in the City of London

Staff Security Engineer

Location
City Of London, England, United Kingdom
engineer who enjoys solving complex security challenges at scale. You’ll work across engineering, data, AI and digital workplace teams to build security observability, automate control assurance and influence how security is embedded into products, platforms and processes. If you're passionate about turning security data into actionable insight, building … raising the security maturity of a fast-moving technology organisation, we'd love to hear from you. About the role Designing and building security observability capabilities that provide meaningful visibility across systems, infrastructure and applications Developing automated control monitoring, evidence collection and continuous testing solutions that strengthen security governance Partnering ...

Senior Security Engineer: Observability, Automation & Risk

Location
City Of London, England, United Kingdom
seeking a Staff Security Engineer to advance security observability, automate controls and strengthen governance across its global tech estate. This senior individual contributor role partners with engineering, data, AI and digital workplace teams to raise security maturity in a fast-moving environment. You will design scalable security observability, automate assurance ...

Principal Data Engineer

Location
City Of London, England, United Kingdom
follow the identical pattern so they are handover-ready by design. Drive data quality as a first-class, firm-wide concern: establish data contracts, observability, SLA/SLO monitoring, and automated alerting and remediation across ingestion and transformation layers, and hold squads to those standards. Act as the senior technical … with the ability to set standards, conduct code and design reviews, and grow engineers’ capabilities Strong grasp of data quality practices: data contracts, pipeline observability, SLA/SLO definition, and automated alerting and remediation Solid understanding of SQL transformation patterns and modern tooling such as dbt, alongside experience managing ingestion ...

Senior Software Engineer

Location
City Of London, England, United Kingdom
Impact and Responsibilities We are seeking a Senior Software Engineer who thrives on untangling complex systems and modernising core infrastructure without breaking production. This is an exciting opportunity to modernise core C#/SQL systems ...

Lead Backend Engineer

Location
City Of London, England, United Kingdom
About us At Zego, we know that traditional motor insurance holds good drivers back. It’s too complicated, too expensive, and it doesn't take into account how well you actually drive. That’s why ...

Lead Python Backend Engineer - Processing & Observability

Location
City Of London, England, United Kingdom
seeking a skilled software developer to enhance their data orchestration pipeline using Python. Join the Data Intelligence Team to work on AI features, develop observability for ML applications, and mentor peers. The role offers remote flexibility, a supportive work environment, and generous compensation along with various benefits. Ideal candidates have ...

Senior DevOps Engineer - ELK

Hiring Organisation
ECS
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
£500 - £700 per day
become available to join one of the world's leading technology organisations as a Senior DevOps Engineer, helping to deliver and scale critical observability, automation, and platform engineering solutions across a complex enterprise environment. As a Senior DevOps Engineer, you will be responsible for: Designing, deploying, and optimising large-scale … Elasticsearch environments. Leading improvements to monitoring, logging, and observability platforms. Driving automation and infrastructure-as-code best practices. Supporting and enhancing Kubernetes and GitOps deployments. Troubleshooting complex performance and reliability issues. Collaborating with technical teams to improve platform scalability, resilience, and security. Promoting DevOps best practices across engineering teams. Requirements ...

Staff ML Engineer | Agentic AI & Applied ML | London (Hybrid) | Contract | Inside IR35

Location
City Of London, England, United Kingdom
implementation patterns Designing and evolving production RAG and retrieval architectures Establishing effective LangGraph/LangChain patterns for agentic applications Improving AI evaluation, testing, observability and production monitoring Developing guardrails, controls and approaches to hallucination and model risk Supporting the move towards increasingly high-risk and high-complexity AI/… based applications Retrieval Augmented Generation (RAG) LangChain and/or LangGraph Vector databases and retrieval MLOps and production deployment AI evaluation, testing and observability AI governance, model risk and engineering controls ML frameworks such as PyTorch, TensorFlow or Scikit-learn Experience operating in complex, regulated or high-risk environments would ...

Senior Software Development Engineer

Location
City Of London, England, United Kingdom
expectations and can be reused across multiple brands and platforms. Drive engineering excellence for the services you own by championing code quality, automated testing, observability, performance optimization, and simplification, taking technical responsibility for service health, scalability, resilience, and the ongoing reduction of technical debt and operational overhead. Provide technical mentorship … integrations, and communicating trade-offs to both technical and non-technical stakeholders. Track record of improving operational excellence at the team level through enhanced observability, automation, performance tuning, and data-driven analysis of incidents and customer impact. Hands-on experience integrating or consuming AI/ML-enabled services or platforms ...

Senior Principal Software Engineer Dev O

Location
City Of London, England, United Kingdom
partners, the SPSE will combine hands‐on technical leadership with strategic influence. They will shape and deliver cross‐cutting improvements spanning operational maturity, resilience, observability, DevSecOps and application security, including the adoption of Application Security Posture Management (ASPM) capabilities, while exploring Agentic Dev Operations and AI‐enabled automation to improve … that standardise approaches, simplify the estate, drive efficiency and close capability gaps. Define and embed pragmatic standards, guardrails and engineering practices for reliability, resilience, observability, service performance and operational readiness, balancing consistency with the needs of autonomous engineering teams. Provide technical leadership for the ongoing DevSecOps programme, integrating security into ...

Principal AI Engineer

Hiring Organisation
Intellias
Location
City of London, London, United Kingdom
Our client is a leading global investment management firm headquartered in London, managing over $228B in assets. Technology, data science, machine learning, and AI are at the heart of its investment and research ecosystem. The ...

Databricks Champion Architect

Location
City Of London, England, United Kingdom
We’re hiring a Databricks Champion Architect to define and govern our modern lakehouse architecture. You’ll blend solution design, hands‐on technical leadership, and platform enablement; setting patterns, accelerating delivery teams, and ensuring production ...

IBM Netcool / Observability Technical Lead

Hiring Organisation
Deerfoot Recruitment Solutions
Location
City, London, United Kingdom
Employment Type
Contract
Contract Rate
GBP 780 - 830 Daily
Netcool/Observability Technical Lead Inside IR35 Contract -up to £827pd London Hybrid - 4 Days Onsite/1 Day WFH per Week Banking Are you the person who knows exactly why an ObjectServer failover didn't behave as expected, and how to stop a flood of duplicate events before anyone ...

Cloud FinOps Analyst

Hiring Organisation
Manufacturing Recruitment Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£60,000
across Azure and Snowflake environments. A key focus of the role is leading the FinOps optimisation activities, embedding governance frameworks, and overseeing AKS cost observability using tooling such as Power BI, Kubecost etc. The FinOps Analyst partners closely with Engineering, Data, Cloud Operations, and Finance teams to enable a cost … optimisation, and waste elimination. Develop, maintain, and enforce cloud and data platform cost governance frameworks including tagging, budgeting, guardrails, and accountability processes. Oversee cost observability tooling (Kubecost, Snowflake dashboards, cloud cost portals) to ensure visibility of usage, forecasts, and budget performance. Manage budgeting, forecasting, cost allocation, and financial reporting ...

Sr Director, Platform Engineering Data Platform & Agentic Platform

Hiring Organisation
Hackajob Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent
operate agent workflow platform capabilities aligned to product-defined standards and interfaces, including traceability, state handling, and convergence patterns Implement production-grade evaluation, observability, auditability, and guardrail mechanisms required for safe AI workflows Implement security controls, access governance, encryption, and audit requirements in partnership with InfoSec while ensuring enterprise SDLC … large-scale SaaS systems with production operations accountability Demonstrated success building and operating platforms adopted by multiple product teams, including reliability discipline (SLOs), observability, and incident management Strong hands-on technical leadership background in distributed systems and platform engineering Deep experience with data platform engineering at scale, including ingestion ...

Clickhouse Solutions Architect

Location
City Of London, England, United Kingdom
Role We are looking for a ClickHouse Solutions Architect to join our team supporting the design and implementation of a greenfield, enterprise-scale ClickHouse observability platform for a global banking client. This is a genuine greenfield build at significant scale — there is no incumbent platform to inherit or work around. … where benchmarks disprove the design Establish infrastructure-as-code, CI/CD and environment promotion for schema and configuration changes Productionisation Define and implement observability — system table monitoring, metrics, alerting thresholds, capacity headroom tracking Establish backup, restore and disaster recovery, and validate them by test Implement security and governance — RBAC ...

Principal Product Engineer

Hiring Organisation
SR2 | Socially Responsible Recruitment | Certified B Corporation™
Location
City of London, London, United Kingdom
Docker MySQL/relational databases High-volume data processing and distributed workflows Typed APIs and integration contracts Event-driven architecture and reliable processing Observability, resilience and operational excellence What you’ll be doing Own the technical direction and architecture of the backend platform Design scalable systems capable of handling millions … events Establish clear boundaries between application, data and delivery layers Build robust, well-typed APIs and processing workflows Improve reliability through idempotency, retries, observability and failure isolation Lead technical design reviews and raise engineering standards Mentor engineers through code reviews and technical pairing Drive improvements across scalability, security and maintainability ...

Software Engineer, Platform Systems

Location
City Of London, England, United Kingdom
systems that provide visibility into large‐scale training workloads and help operate them reliably at scale. You’ll work on failure detection, tracing, and observability systems that identify slow or faulty nodes, surface performance bottlenecks, and help engineers understand and optimize massive distributed training jobs. This infrastructure is critical … systems for large‐scale AI training jobs Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable visibility into system behavior Improve observability, reliability, and performance across OpenAI’s training platform Debug and resolve issues in complex, high‐throughput distributed systems Collaborate with systems, infrastructure, and research teams ...