10 of 10 Observability Jobs in Westminster

Lead Data Engineer

Hiring Organisation
Hackajob Ltd
Location
Westminster, Greater London, UK
processing and data quality frameworks using Python, PySpark, and dbt Build and optimize batch and streaming data pipelines with strong performance, fault tolerance, and observability Develop and operate workflow orchestration (e.g., Apache Airflow) to schedule, monitor, and manage data movement and transformations Model and transform data for analytics using ...

Principal Software Engineer - Platform Engineering - Accelerator Business

Hiring Organisation
Hackajob Ltd
Location
Westminster, Greater London, UK
Preferred qualifications, capabilities, and skills Advanced knowledge ofCI/CD, application resiliency, and secure delivery (e.g., SLSA framework and GitOps). Deep experience with Observability and Monitoring tools (e.g., Prometheus, Grafana, OTEL). Expertise in performance optimisation of distributed systems (e.g., caching, network latency). Practical experience with Service Mesh ...

Senior Machine Learning Scientist

Location
Westminster, West End, United Kingdom
enable repeatable releases Develop reusable engineering assets (libraries, templates, reference architectures, infrastructure-as-code patterns) to reduce technical debt and accelerate delivery Implement observability for AI services (logging/metrics/tracing), model performance monitoring, and quality/drift checks with actionable alerting Partner with data scientists, data engineers, platform ...

Software Engineer III - Data Analytics Platform

Hiring Organisation
Hackajob Ltd
Location
Westminster, Greater London, UK
degradation, autoscaling behaviors, incident follow-ups, and runbooks. Contribute to system design by breaking down ambiguous problems, proposing approaches, and making pragmatic tradeoffs. Add observability with metrics, tracing, logging, dashboards, and actionable alerts tied to SLOs. Support safe deployments through CI/CD improvements, canarying, feature flags, backward compatibility ...

Head of Production Management- Personal Investing

Hiring Organisation
Hackajob Ltd
Location
Westminster, Greater London, UK
ability to influence across technical and non-technical audiences in complex, regulated environments. Strong technical background in modern production environments including cloud-native infrastructure, observability, CI/CD pipelines, and automation frameworks. Proven track record defining and governing production management standards — change, incident, capacity, and automation — across multiple engineering teams ...

Senior Engineering Lead

Hiring Organisation
Hackajob Ltd
Location
Westminster, Greater London, UK
. Build the tooling, controls, and change-management practices that make SOX compliance a property of how we Engineer. Operational maturity. Drive deployment frequency, observability, incident response, and change control to a standard appropriate for systems the CFO signs off on. Co-own the roadmap with Finance and Product . ...

Lead SRE

Hiring Organisation
Hackajob Ltd
Location
Westminster, Greater London, UK
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … knowledge of microservice infrastructure components, including service discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Westminster, Greater London, UK
undergoing a multi-year convergence and modernization journey. You will play a pivotal role in shaping our next-generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast-paced front-office environments, enjoys … Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with ...

Product Associate - SRE Team

Hiring Organisation
Hackajob Ltd
Location
Westminster, Greater London, UK
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … services are designed, delivered, and operated with reliability in mind. Job responsibilities Support the product strategy and delivery of reliability capabilities, including standards, observability, incident practices, automation, and developer experience improvements. Partner with engineers, site reliability engineers, and cross-functional teams to understand problems, gather requirements, and translate ideas into ...

Lead Software Engineer - Proxy/SSE Network Security

Hiring Organisation
Hackajob Ltd
Location
Westminster, Greater London, UK
resilience outcomes. Drive operational excellence at scale for perimeter, proxy, and SSE services in the US, including incident, change, and problem management rigor, observability and resiliency validation practices, automation to improve repeatability and evidence quality, reduction of client and partner impact, and execution of Technology Lifecycle Management (TLM) and modernization … design, exception frameworks, audit-ready traceability, and measurable risk reduction reporting. Experience with large-scale operations for externally facing or security enforcement services, including observability strategy, resilience testing, incident response alignment, and reduction of repeat incidents and client-impacting events. Experience designing and operating hybrid edge architectures and cloud interconnect ...