2,126 to 2,150 of 2,280 Observability Jobs in London

Sr Director, Platform Engineering – Data Platform & Agentic Platform

Location
Greater London, England, United Kingdom
operate agent workflow platform capabilities aligned to product‐defined standards and interfaces, including traceability, state handling, and convergence patterns Implement production‐grade evaluation, observability, auditability, and guardrail mechanisms required for safe AI workflows Implement security controls, access governance, encryption, and audit requirements in partnership with InfoSec while ensuring enterprise SDLC … large‐scale SaaS systems with production operations accountability Demonstrated success building and operating platforms adopted by multiple product teams, including reliability discipline (SLOs), observability, and incident management Strong hands‐on technical leadership background in distributed systems and platform engineering Deep experience with data platform engineering at scale, including ingestion ...

Cloud FinOps Analyst

Hiring Organisation
Manufacturing Recruitment Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£60,000
across Azure and Snowflake environments. A key focus of the role is leading the FinOps optimisation activities, embedding governance frameworks, and overseeing AKS cost observability using tooling such as Power BI, Kubecost etc. The FinOps Analyst partners closely with Engineering, Data, Cloud Operations, and Finance teams to enable a cost … optimisation, and waste elimination. Develop, maintain, and enforce cloud and data platform cost governance frameworks including tagging, budgeting, guardrails, and accountability processes. Oversee cost observability tooling (Kubecost, Snowflake dashboards, cloud cost portals) to ensure visibility of usage, forecasts, and budget performance. Manage budgeting, forecasting, cost allocation, and financial reporting ...

Operational and Technology Platforms Lead - ITSM - CMDB - ITIL

Hiring Organisation
Tria Recruitment
Location
London, UK
Employment Type
Full-time
ResponsibilitiesDefine and drive platform strategy, governance and roadmaps. Standardise and rationalise platforms across multiple business divisions. Lead the development of ITSM, ITAM, CMDB and observability capabilities. Deliver greater visibility of applications, assets, licences and technology risks. Manage technology vendors, budgets and platform performance. Build and lead the Operational & Technology Platforms … platforms, technical teams or technology functions. Experience working within complex, federated or multi-business organisations. Strong knowledge of ITSM, ITAM, CMDB, endpoint management and observability tooling. A track record of platform transformation, standardisation and continuous improvement. Experience managing technology platforms within engineering, industrial, manufacturing, energy, defence or other asset-intensive ...

Operational and Technology Platforms Lead - ITSM - CMDB - ITIL

Hiring Organisation
Tria
Location
London, United Kingdom
Employment Type
Permanent
Define and drive platform strategy, governance and roadmaps. Standardise and rationalise platforms across multiple business divisions. Lead the development of ITSM, ITAM, CMDB and observability capabilities. Deliver greater visibility of applications, assets, licences and technology risks. Manage technology vendors, budgets and platform performance. Build and lead the Operational & Technology Platforms … platforms, technical teams or technology functions. Experience working within complex, federated or multi-business organisations. Strong knowledge of ITSM, ITAM, CMDB, endpoint management and observability tooling. A track record of platform transformation, standardisation and continuous improvement. Experience managing technology platforms within engineering, industrial, manufacturing, energy, defence or other asset-intensive ...

Tech Lead

Location
Greater London, England, United Kingdom
Product, Sales, Marketing, Support, Operations, Legal, and Compliance Mentor and coach engineers, supporting their technical growth and confidence Set standards for code quality, testing, observability, and operational excellence Collaborate with Product and Design to shape solutions, challenge assumptions, and manage trade-offs Lead technical discussions, reviews, and incident investigations Contribute … Modelling Testing Strategies System Design React PostgreSQL Soft Skills Mentoring Communication Collaboration Influence Coaching Industry Keywords Production Software Regulated Environment Operational Excellence Code Quality Observability Tools & Technologies GCP React Native Temporal AI-Assisted Engineering Tools #J-18808-Ljbffr ...

Applied AI Engineer New

Location
Greater London, England, United Kingdom
built quickly and as separate systems. The next challenge is to bring these approaches together: build reusable agentic infrastructure, establish a robust evaluation and observability layer, and create systems that allow us to automate new workflows quickly and reliably as Dwelly scales. This is not an AI research role. … loops. Move us from one-off AI solutions toward reusable infrastructure where new workflows can be introduced quickly and with predictable reliability. 2. Evaluation & observability Build the evaluation framework that allows us to understand how our agents perform and why they succeed or fail. Make testing, tracing, debugging, and evaluating ...

Solutions Engineer

Location
Greater London, England, United Kingdom
Riverbed. Empower the Experience Riverbed, the leader in AIOps for observability, helps organizations optimize their users' experiences by leveraging AI automation for the prevention, identification, and resolution of IT issues. With over 20 years of experience in data collection and AI and machine learning, Riverbed's open and AI-powered … observability platform and solutions optimize digital experiences and greatly improve IT efficiency. Riverbed also offers industry-leading Acceleration solutions that provide fast, agile, secure acceleration of any app, over any network, to users anywhere. Together with our thousands of market-leading customers globally – including 95% of the FORTUNE ...

Systems Engineer SRE Golang - FinTech

Hiring Organisation
Client Server
Location
London, UK
Employment Type
Full-time
level. Reliability will be central to everything you do. You'll design systems with that are resilient, adaptable and highly available, working across performance, observability and security. You'll also work with Infrastructure as Code such as Terraform and Kubernetes. You'll be collaborating with and learning from a hugely … application and the infrastructure needed to support itYou have a strong knowledge of network, operating system and application level securityYou have experience with systems observability/SRE e.g. logs, metricsYou have experience with IaC, Terraform preferredYou have a good understanding of Kubernetes and how it worksYou have a good knowledge ...

Systems Engineer SRE Golang - FinTech

Hiring Organisation
Client Server
Location
Central London, London, United Kingdom
Employment Type
Permanent, Work From Home
level. Reliability will be central to everything you do. You'll design systems with that are resilient, adaptable and highly available, working across performance, observability and security. You'll also work with Infrastructure as Code such as Terraform and Kubernetes. You'll be collaborating with and learning from a hugely … infrastructure needed to support it You have a strong knowledge of network, operating system and application level security You have experience with systems observability/SRE e.g. logs, metrics You have experience with IaC, Terraform preferred You have a good understanding of Kubernetes and how it works You have ...

Technical Account Manager

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
team rewriting the rules of observability. A platform at the bleeding edge, empowering businesses to understand and act on their data through better observability, all in real time. Innovation removes the need for unnecessary indexing, cutting costs and complexity. Truly end to end, the platform encompasses everything; from logs … This is a high-impact role that demands serious technical firepower. You'll need deep experience with Cloud native tooling, hands-on knowledge of observability tools e.g. Grafana, DataDog or Splunk, and the ability to troubleshoot containerised environments like a pro. You have the technical knowledge and the confidence ...

Senior/Staff Software Engineer (Nova Core)

Location
Greater London, England, United Kingdom
operations across Nova Cloud deployments. This role focuses on the Nova Core “inner loop”: service architecture, APIs, data models, persistence, authn/authz, observability, and developer experience that other Nova modules and product teams depend on. What you’ll do Own and ship critical Nova Core backend services (e.g., common … engineering and product teams. What success looks like Core Nova services are delivered, adopted, and operated reliably with clear SLIs/SLOs and runbooks. Observability is strong enough that incidents are detected quickly and resolved faster over time (improving MTTD/MTTR). API versioning and compatibility practices reduce integration ...

Senior/Staff Software Engineer

Location
Greater London, England, United Kingdom
tooling, optimizing for reliability, latency, cost, and debuggability in production. Build and maintain the surrounding infrastructure: data pipelines, evaluation harnesses, prompt and model management, observability, and safety/guardrails. Work across the stack—from backend integrations and APIs to simple UI hooks—to deliver complete AI features, not just model … workflows quickly, then harden what works. Own, downscope, ship, iterate: one clear owner per feature, from prototype to production. Fundamentals done well: evaluation, observability, and safety are part of the first version, not an afterthought. Competitive salary and meaningful equity. Health, dental, and vision coverage. Flexible time off and support ...

Operations and SRE Manager

Location
Greater London, England, United Kingdom
internal and external customers. You will be responsible for driving reliability improvements, advancing automation and AI-Ops capabilities, and leading a team focused on observability, incident response, operational excellence, and continuous improvement. Responsibilities Lead the transformation of the Operations function towards an AI-Ops operating model, driving the adoption … improvement actions are owned, tracked and completed. Strengthen operational process adherence, ensuring responsibilities are clear and delegation is effective. Drive SRE practices across observability, automation, disaster recovery, design for reliability, on-call readiness and production support. Protect service levels by ensuring engineering effort is balanced across InfoSec commitments, operational tickets ...

Software Engineer, ChatGPT Infrastructure

Location
Greater London, England, United Kingdom
diagnosing performance, scalability, or reliability issues in production environments. Understanding of distributed systems, data storage, concurrency, asynchronous processing, or networking. Familiarity with modern deployment, observability, and cloud infrastructure practices. Ability to lead complex technical work and collaborate effectively across product and infrastructure teams. Experience with cloud infrastructure, containerized environments … observability tools is useful but not required. We welcome candidates from backend engineering, distributed systems, platform engineering, and reliability backgrounds. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general‐purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities ...

Software Engineer (Simulation, Evaluation, Validation)

Location
Greater London, England, United Kingdom
differences between on-road and simulated execution, identifying issues across data, inference and simulated components Improve simulation reproducibility, reliability and debuggability through automated testing, observability and better developer tooling Profile and improve simulator performance, helping us run increasingly large evaluation workloads efficiently Work with internal users and adjacent engineering teams … sensor data such as camera, radar, lidar or GNSS, including modelling uncertainty or noiseExperience integrating machine-learning inference into production systemsExperience with performance profiling, observability or debugging distributed systemsFamiliarity with large-scale batch processing, cloud infrastructure or GPU-based workloads #J-18808-Ljbffr ...

Platform Software Engineer

Location
Greater London, England, United Kingdom
Modelling infrastructure state and making changes idempotent, attributable and auditable Defining stable contracts between storage platform and developer platform services Building tests, fixtures and observability that make failures safe to diagnose Working with users and partner teams to turn operational problems into bounded engineering work Participating in incident response … reconciliation systems Understanding of retries, partial failure, concurrency, idempotency and eventual consistency Strong Linux and production troubleshooting skills Experience with CI/CD, observability and safe deployment practices A practical approach to security, permissions, audit and change control A calm, methodical approach to incidents Desirable experience includes orchestration, configuration management ...

Senior Cloud Data Engineer - KSP

Location
Greater London, England, United Kingdom
optimise streaming data pipelines (Kafka/Flink or equivalent) to enable near real‐time data availability. Ensure data quality and reliability through validation frameworks, observability, and robust handling of late‐arriving or inconsistent data. Design data contracts and schemas that enable reliable integration between upstream event producers and downstream consumers. … medallion architecture or similar data layering approaches. Experience working with streaming technologies (Kafka, Flink, or similar). Strong understanding of data quality, testing, and observability practices. Experience designing schemas and handling data consistency challenges in distributed systems. Ability to work closely with stakeholders to translate business needs into scalable data ...

Senior Software Engineer - Runtime Platform, Robot Software

Location
Greater London, England, United Kingdom
support new product features, which is critical to the success of Wayve’s mission. The Runtime Platform team equips all Wayve teams with the observability, profiling tools, and infrastructure needed to understand and optimise software performance across our development fleet. We work closely with teams to investigate issues, reduce bottlenecks … caches, context switches, ...), and thread synchronisation Desirable Familiarity with Nvidia performance tools such as NV NSight, NV Lumos and tegrastat Familiarity with observability tools such as Grafana (logs, metrics, traces), Databricks, Datadog Familiarity with QNX and Momentics is a plus This is a full-time role based in London. ...

Engineering Manager - People & Procurement New London

Location
Greater London, England, United Kingdom
owns the integrations and data flows connecting People & Procurement platforms with the wider FT technology estate. Own operational performance and establish effective approaches to observability, incident and problem management, change management, resilience, disaster recovery and business continuity. Ensure relevant security, privacy, data governance, compliance and technology standards are met, with … system dependencies. Experience owning technology across its lifecycle, from change and delivery through production operation and continuous improvement. Strong understanding of reliability, security, privacy, observability, operational risk and technology lifecycle management. Experience translating business priorities into engineering plans and balancing feature delivery with platform health and technical investment. Experience leading ...

Staff Software Engineer, AI Reliability Engineering

Location
Greater London, England, United Kingdom
Responsibilities Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity. Design and implement monitoring and observability systems across the token path. Assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud providers Lead incident response … more ML hardware accelerators (GPUs, TPUs, Trainium). Understand ML-specific networking optimizations like RDMA and InfiniBand. Have expertise in AI-specific observability tools and frameworks. Have experience with chaos engineering and systematic resilience testing. Have contributed to open-source infrastructure or ML tooling. The annual compensation range for this ...

AI Productivity Engineer

Hiring Organisation
Aircall
Location
London, UK
Employment Type
Full-time
productionize AI-driven solutions with strong autonomy on how problems are solvedAutomate and streamline workflows across GitLab, Jira, CI/CD, Slack, and observability toolsDesign and operate internal AI services and orchestration layers (e.g. MCP servers)Own solutions end-to-end: discovery design build measure iterateWork hands-on with engineering … injectionAI-powered tooling or internal platformsSolid backend engineering skills (APIs, services, integrations)Experience working with developer tools (CI/CD, GitHub/GitLab, Jira, observability)Strong product mindset and comfort operating in ambiguous problem spacesNice to HaveParticularly interesting profiles are engineers who have built developer tools and are now evolving ...

IT Infrastructure Solutions Architect

Location
Greater London, England, United Kingdom
roadmap.**Key responsibilities*** Define and maintain reference architectures and target-state designs for VMware VCF 9.0 platform architecture and lifecycle patterns.* Define Aria Operations observability strategy (telemetry standards, alert philosophy, capacity/performance governance, service reporting) and ensure operational adoption.* Define VCF Automation platform approach (catalog/service design, templates … iSCSI), VSAN, NAS, and software-defined storage concepts.* Experience or exposure to infrastructure-as-code* Proven capability to architect and operationalize enterprise monitoring/observability standards (Logic Monitor and Aria Operations).* Proven capability to architect, govern, and troubleshoot provisioning automation (VCF Automation).* Proven backup/recovery architecture ...

Engineering Manager Sotheby's · London / Remote

Location
Greater London, England, United Kingdom
priorities and deliver meaningful outcomes Track record of hiring well, onboarding effectively, and building teams that retain good people A focus on engineering quality: observability, analytics and reporting, testing practices, and ensuring the team ships work they can stand behind Useful: Familiarity with our core stack: Go, Scala, React/… feature delivery, technical investment and operational health Set and maintain the quality bar for your team’s area of the platform, including engineering practices, observability, reliability and how work gets reviewed Run the delivery cadence: planning, prioritisation, retrospectives, and the day-to-day work of keeping things moving and unblocked ...

Engineering Lead – ERP Services

Location
Greater London, England, United Kingdom
impacts. Operational Excellence & Risk Owning the operational performance of ERP Services and ensuring critical services are appropriately monitored and supported. Establishing effective approaches to observability, incident management, problem management, change management and operational readiness. Ensuring recurring incidents and operational problems are investigated to their underlying causes and addressed sustainably. Maintaining … developing engineering capability and strengthening technical ownership within teams. Experience working with Finance or other complex enterprise business domains. Strong understanding of reliability, security, observability, operational risk and technology lifecycle management. Experience translating business priorities into engineering plans and balancing feature delivery with platform health and technical investment. Experience leading ...

Lead Technical Program Manager - Low Latency Market Trading

Location
City Of London, England, United Kingdom
Leverage your deep technical expertise and leadership to guide cutting-edge projects, fostering growth and innovation in a dynamic environment. As a Lead Technical Program Manager in Corporate Technology, you will lead the delivery of ...