151 to 156 of 156 Observability Jobs in Central London

Head of Technology Operations & Service Delivery

Location
City of Westminster, England, United Kingdom
Report transparently on cloud utilisation, demand, optimisation opportunities and cost drivers. Work with engineering, architecture, finance and suppliers to improve efficiency without weakening control. Observability & operational intelligence Develop proactive monitoring and end-to-end observability across critical services. Use operational data to identify service degradation, capacity constraints and control weaknesses … regulated bank or financial institution. Experience with Microsoft Azure, Microsoft 365 and cloud-first operating environments. Experience building service catalogues, service ownership models and observability capabilities. ITIL v4 or advanced service management qualification. What we can offer you: 8% company pension contribution and 3% individual contribution (which ...

Cloud FinOps Analyst

Hiring Organisation
Manufacturing Recruitment Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£60,000
across Azure and Snowflake environments. A key focus of the role is leading the FinOps optimisation activities, embedding governance frameworks, and overseeing AKS cost observability using tooling such as Power BI, Kubecost etc. The FinOps Analyst partners closely with Engineering, Data, Cloud Operations, and Finance teams to enable a cost … optimisation, and waste elimination. Develop, maintain, and enforce cloud and data platform cost governance frameworks including tagging, budgeting, guardrails, and accountability processes. Oversee cost observability tooling (Kubecost, Snowflake dashboards, cloud cost portals) to ensure visibility of usage, forecasts, and budget performance. Manage budgeting, forecasting, cost allocation, and financial reporting ...

Sr Director, Platform Engineering Data Platform & Agentic Platform

Hiring Organisation
Hackajob Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent
operate agent workflow platform capabilities aligned to product-defined standards and interfaces, including traceability, state handling, and convergence patterns Implement production-grade evaluation, observability, auditability, and guardrail mechanisms required for safe AI workflows Implement security controls, access governance, encryption, and audit requirements in partnership with InfoSec while ensuring enterprise SDLC … large-scale SaaS systems with production operations accountability Demonstrated success building and operating platforms adopted by multiple product teams, including reliability discipline (SLOs), observability, and incident management Strong hands-on technical leadership background in distributed systems and platform engineering Deep experience with data platform engineering at scale, including ingestion ...

Clickhouse Solutions Architect

Location
City Of London, England, United Kingdom
Role We are looking for a ClickHouse Solutions Architect to join our team supporting the design and implementation of a greenfield, enterprise-scale ClickHouse observability platform for a global banking client. This is a genuine greenfield build at significant scale — there is no incumbent platform to inherit or work around. … where benchmarks disprove the design Establish infrastructure-as-code, CI/CD and environment promotion for schema and configuration changes Productionisation Define and implement observability — system table monitoring, metrics, alerting thresholds, capacity headroom tracking Establish backup, restore and disaster recovery, and validate them by test Implement security and governance — RBAC ...

Principal Product Engineer

Hiring Organisation
SR2 | Socially Responsible Recruitment | Certified B Corporation™
Location
City of London, London, United Kingdom
Docker MySQL/relational databases High-volume data processing and distributed workflows Typed APIs and integration contracts Event-driven architecture and reliable processing Observability, resilience and operational excellence What you’ll be doing Own the technical direction and architecture of the backend platform Design scalable systems capable of handling millions … events Establish clear boundaries between application, data and delivery layers Build robust, well-typed APIs and processing workflows Improve reliability through idempotency, retries, observability and failure isolation Lead technical design reviews and raise engineering standards Mentor engineers through code reviews and technical pairing Drive improvements across scalability, security and maintainability ...

Software Engineer, Platform Systems

Location
City Of London, England, United Kingdom
systems that provide visibility into large‐scale training workloads and help operate them reliably at scale. You’ll work on failure detection, tracing, and observability systems that identify slow or faulty nodes, surface performance bottlenecks, and help engineers understand and optimize massive distributed training jobs. This infrastructure is critical … systems for large‐scale AI training jobs Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable visibility into system behavior Improve observability, reliability, and performance across OpenAI’s training platform Debug and resolve issues in complex, high‐throughput distributed systems Collaborate with systems, infrastructure, and research teams ...