1,201 to 1,225 of 1,896 Observability Jobs

Principal Java Engineer

Hiring Organisation
Jobleads-UK
Location
Wallingford, England, United Kingdom
members to drive technical excellence and continuous improvement. Production systems Lead root cause analysis and resolution of complex production issues. Drive improvements in system observability, monitoring and operational performance. Ensure applications are designed and operated to meet reliability, availability and performance targets. Partner with Operations, DevOps and QA teams … Claude, Codex, GitLab Duo, etc.) REST APIs, OpenAPI, Microservices, Event‐driven architecture (RabbitMQ) Containers, Docker, AWS, Linux CI/CD with GitLab Pipelines & Jenkins Observability: logging, metrics and monitoring MySQL, Apache Solr Front‐end UI (e.g. Angular) Person Specification Strategic and systems‐thinking mindset Excellent communication and stakeholder management skills ...

Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Azure-hosted services. Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Proactively monitor platform health and performance using observability tooling. Perform root cause analysis and implement permanent fixes for recurring incidents. Participate in incident management and on-call support rotations where required. Lead blameless … SDLC methods Strong understanding of Site Reliability Engineering principles, including SLIs, SLOs, SLAs, error budgets, reliability targets, and service health measurement. Experience designing observability strategies across metrics, logs, traces, synthetic monitoring, alerting, dashboards, and operational telemetry. Ability to define actionable alerts that identify customer-impacting symptoms, reduce noise, and support ...

Senior Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Welwyn Garden City, England, United Kingdom
GitOps workflows to enable safe, fast, and repeatable delivery. Championing DevSecOps principles, embedding security and compliance into the software delivery lifecycle. Establishing and improving observability, monitoring, and incident response practices, including vulnerability management and remediation. Mentoring engineers and contributing to a strong engineering culture through knowledge sharing, documentation, and technical … would be great if you have the following Experience with Helm, Kustomize, and Kubernetes ecosystem tooling. Familiarity with Azure and Azure DevOps. Experience with observability platforms and Kubernetes policy enforcement tools. Proficiency in scripting or programming (e.g. Bash, Python, PowerShell, C#). Experience designing multi‐region or highly available systems. ...

Senior Software Engineer (Infrastructure)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
engineering team in building reliable, scalable applications Help design and build tools to make our services scalable and highly available, automating wherever possible Develop observability and orchestration tooling to allow our customers to monitor and manage their deployments Contribute to disaster recovery, backup and redundancy tooling and strategy Assist … hands‐on experience Experience developing production‐ready infrastructure management tooling with either Python or Golang Familiarity with at least one of the following: Observability Tools (e.g. Prometheus, OpenTelemetry, Grafana) Databases (e.g. Postgres, DuckDB) Event Streaming platforms (e.g. Kafka) Container Orchestration (e.g. Docker, Kubernetes) Familiarity with cloud platforms such ...

Lead Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
these practices within an application or platform Fluency in at least one programming language such as Java, Python, Go, etc. Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency … high-availability services Deep understanding of distributed system design principles, networking (TCP/IP, DNS, load balancing), Linux internals. Contributions to open-source observability or telemetry projects. Experience working with agent control planes and management protocols. Hands‐on knowledge of OpAMP is highly desirable. We recognize that our people ...

Senior Data Engineer

Hiring Organisation
CGI
Location
City and Borough of Leeds, United Kingdom
Employment Type
Full Time
join multi-disciplinary teams where pragmatic ownership, creative solutioning and strong technical craft translate directly into business outcomes-building reusable patterns, improving observability and accelerating client roadmaps for modern data adoption. CGI was recognised in the Sunday Times Best Places to Work List 2025 and has been named … will work across Foundry and cloud ecosystems (or quickly gain Foundry expertise), translating stakeholder needs into maintainable, production-grade solutions while championing quality, observability and reusable engineering patterns. You will collaborate closely with architects, product owners and analysts to shape roadmaps and deliver measurable outcomes. Required qualifications to be successful ...

Safety Engineer United Kingdom +11 more

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
models into production systems. Architect robust APIs, data pipelines, and service architectures supporting real‐time and batch moderation workflows. Implement comprehensive monitoring, alerting, and observability systems; establish SLIs, SLOs, and performance benchmarks. Partner with ML engineers to translate research models into production‐ready systems and integrate them across our product … Python expertise (asynchronous Python, backend frameworks). Infrastructure & DevOps proficiency: cloud platforms (AWS/GCP), containerization (Docker/K8s), CI/CD pipelines. Observability mindset with experience in monitoring tools (Prometheus, Grafana) and building observable systems. Track record of taking products or systems from 0→1 with measurable impact, including ...

Senior Manager- Software Engineering

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
microservices, including REST and event‐driven patterns, and enforce best practices for versioning, contracts, and backward compatibility. Advance operational excellence by defining SLOs, improving observability and alerting, hardening on‐call procedures and runbooks, and leading incident response and post‐mortems. Solve complex distributed system challenges (such as throughput, latency, consistency … with microservices, API design, and event‐driven systems, as well as experience with containerization and orchestration (Docker, Kubernetes). Strong operational mindset: expert in observability, monitoring, incident response, performance engineering, and adherence to security best practices. Familiarity with CI/CD pipelines, automated testing strategies (unit, integration, e2e), and modern ...

Staff Cloud SRE - AI/ML Platform & GPU Compute

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
escalation, communications, and root cause analysis. Translate post‐incident learning into durable architectural or automation improvements. Continuously reduce alert noise and recurring operational burden. Observability & Operational Excellence Design and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. Build dashboards that reflect real user‐centric platform … Python, Go, C++) with a bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g. Datadog, Prometheus, Grafana, OpenTelemetry). Clear communication skills, including leading incidents, writing postmortems, and influencing teams to prioritise reliability improvements. Desirable skills Familiarity ...

Principal Software Development Engineer

Hiring Organisation
Jobleads-UK
Location
Reigate, England, United Kingdom
pipelines, Infrastructure as Code, automation frameworks, and database-as-code practices using Redgate Flyway. Take ownership of critical customer systems, ensuring operational resilience, observability, performance optimisation, and rapid incident response. Collaborate closely with Product, Delivery, Operations, and Commercial teams to shape technical solutions, delivery plans, and strategic outcomes. Promote secure … Connect or Genesys Cloud. Proven ability to design and deliver secure, scalable, and resilient cloud-native solutions within complex enterprise environments. Strong understanding of observability, operational support, reliability engineering, and end-to-end ownership practices. Knowledge of regulated financial services environments, including UK GDPR and FCA Consumer Duty requirements. Excellent ...

Production AI Engineer - Vice President

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
engineering techniques to integrate large language models (LLMs) into operational tooling, incident response pipelines, and developer productivity platforms. Lead the development of AI‐native observability solutions—leveraging intelligent agents to detect anomalies, predict failures, and automate remediation before issues impact end users. Write clean, well‐tested, and well‐documented code … . Operational experience of using middleware technologies (MQ, Apache Kafka, etc.) to run services at scale is desirable. Strong experience with end‐to‐end observability stacks (Datadog, AppDynamics, Dynatrace, etc.) is desirable. Degree in Computer Science, Mathematics, Physics, or a related technical subject is desirable. Experience of senior stakeholder management. ...

Financial Risk Analytics – Senior Product Analyst

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
test scenarios, release notes, and operational documentation. The role will support the design and continuous improvement of data pipelines, APIs, integration services, controls, monitoring, observability, exception management, and automation across Market Data and Integration workflows. The Senior Product Analyst will investigate complex data and workflow issues using SQL, Python, logs … market data vendors, data mastering, golden-source design, curve construction, historical market data, pricing services, scenario generation, or analytics input validation. Experience with observability, production support, automated controls, regression testing, reconciliation, model input validation, machine learning, NLP, or responsible AI applications in financial analytics. CFA, FRM, CQF, or other relevant ...

Staff Engineer - Data

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
create unnecessary complexity, risk, duplicated capability or long‐term support burden. Raise the quality bar for data products through clear ownership, robust testing, reconciliation, observability, lineage, documentation, performance and supportability. Collaborate with cross‐functional teams to address security, GDPR, PII handling, role‐based access, auditability and data governance are designed … services across batch, streaming and event‐driven patterns. Deep understanding of engineering practice: clean design, testing strategy, CI/CD, infrastructure as code, observability, performance, security, incident response and DevSecOps. Experience with cloud data services and modern data stacks. Relevant technologies may include Snowflake, Azure/AWS/GCP data ...

Enterprise Data Integration Engineer | Enterprise Technology

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
maintain high-performance data pipelines capable of processing large-scale, distributed workloads while meeting operational and regulatory requirements. Implement data quality controls, monitoring, observability, and operational metrics to improve the reliability and resilience of enterprise data products. Define and implement automated testing practices for data and integration solutions, including unit … proficiency in Python and SQL for enterprise data engineering and integration workloads. Deep understanding of how integration architectures and engineering practices influence data quality, observability, lineage, traceability, and auditability. Experience delivering production-grade data solutions that balance scalability, reliability, security, and operational supportability. Ability to operate effectively as both ...

Financial Risk Analytics – Senior Product Analyst

Hiring Organisation
S&P Global
Location
Greater London, United Kingdom
Employment Type
Full Time
test scenarios, release notes, and operational documentation. The role will support the design and continuous improvement of data pipelines, APIs, integration services, controls, monitoring, observability, exception management, and automation across Market Data and Integration workflows. The Senior Product Analyst will investigate complex data and workflow issues using SQL, Python, logs … market data vendors, data mastering, golden-source design, curve construction, historical market data, pricing services, scenario generation, or analytics input validation. Experience with observability, production support, automated controls, regression testing, reconciliation, model input validation, machine learning, NLP, or responsible AI applications in financial analytics. CFA, FRM, CQF, or other relevant ...

Senior Backend Engineer - Java

Hiring Organisation
Capco
Location
Borough of Tameside, United Kingdom
Employment Type
Full Time
This job is with Capco, an inclusive employer and a member of myGwork – the largest global platform for the LGBTQ+ business community. Please do not contact the recruiter directly. Senior Backend Engineer – Java Location: London ...

AVP Site Reliability Engineer - SRE/Infrastructure/Python/Powershell/AWS/Observability/ITIL - PERM

Hiring Organisation
Scope AT Limited
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP Annual
Site Reliability Engineer - SRE/Infrastructure/Python/Powershell/AWS/Observability/ITIL - PERM - Financial Services Job Purpose: The role is primarily responsible for developing SRE methodologies and ensuring they are applied to the Cloud hosted environment. In addition, the role will act as a central point … methodologies, collaborating closely with other infrastructure teams to optimize infrastructure and deployment processes, focusing on automation and operational excellence. Drives continuous improvement in system observability, alerting, and capacity planning through the definition and implementation of SLA, SLOs & SLIs Define and enhance frameworks for Toil identification, analysis & remediation to identify opportunities ...

Platform Engineer: Scalable AWS, CI/CD & Observability

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
availability and performance of our platform Evolve and optimise our CI/CD pipelines to enable fast, safe and consistent deployments Implement and enhance observability (monitoring, alerting, logging) using tools like DataDog to ensure high-availability APM Apply best practices in security, scalability and cost optimisation across our infrastructure Support ...

SRE - Site Reliability Engineer - Observability & Performance

Hiring Organisation
Sanderson Recruitment
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Contract
Contract Rate
£550 - £600 per day
Observability and Performance Up to £600 per day outside IR35 6 month initial contract Bristol - Largely remote I'm currently working with a client who is looking for an SRE to implement and enhance observability across Java applications, middleware and Linux infrastructure using Grafana. The role is focused on monitoring … monitoring, alerting and instrumentation. The environment is currently hosted on traditional infrastructure, with an AWS migration planned, offering the opportunity to develop cloud-ready observability, automation and operational capabilities as the platform evolves. Essential Skills: Strong hands-on experience in DevOps, SRE, Platform Engineering or Systems Engineering environments. Expertise ...

Technical Architecture Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
where GenAI and Agentic play a role Champion system performance, resilience, and efficiency: Proactively identifying and addressing consumption and scalability challenges. Champion full stack observability using modern full stack observability, SRE and AIOps Manage & Mentor: Lead teams of architects and engineers, providing technical coaching, career counselling, performance management, and coaching ...

Lead Windows Server Platform Engineer - Automation & Security

Hiring Organisation
Jobleads-UK
Location
Knutsford, England, United Kingdom
Windows Server Engineering Lead to define and deliver the Windows Server platform roadmap within a large-scale enterprise. You will drive standardisation, automation, security, observability, and continuous service improvement across the estate. You will lead multi-disciplinary engineering teams, champion infrastructure-as-code, and work with cloud, security, and platform ...

Principal Software Engineer, Distributed Identity Workflows

Hiring Organisation
Jobleads-UK
Location
United Kingdom
design across provisioning workflows, pipelines, and distributed execution, mentoring engineers and advancing distributed systems expertise. This hands-on role emphasizes correctness, performance, and observability, with emphasis on asynchronous processing, queues, and retries. You will drive modernization and scalable architectures across teams. #J-18808-Ljbffr ...

Senior Software Engineer - Platform & Deployment Lead

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
will influence architecture, mentor peers, and drive high-quality software with a production-first mindset. You will collaborate with management and partners, champion observability, and help evolve the technology roadmap while maintaining scalable production systems. #J-18808-Ljbffr ...

Senior Software Engineer, Order Management Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Collaborating with Product Managers and Fulfilment stakeholders, you will shape the technical direction of the Order Management platform, driving improvements in architecture, automation, and observability while ensuring #J-18808-Ljbffr ...

Principal Software Engineer, Distributed Identity Workflows

Hiring Organisation
Jobleads-UK
Location
United Kingdom
role covers system design, debugging complex cross‐service workflows, and mentoring engineers in distributed systems across provisioning workflows and pipelines. You will drive reliability, observability, and correctness in asynchronous execution, shaping legacy systems into modern, resilient architectures while collaborating across teams. #J-18808-Ljbffr ...