1,551 to 1,575 of 1,767 Permanent Observability Jobs

Engineering Manager – Payments, BACS, CHAPS

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Responsibilities Deliver and continuously improve payment services across UK payment rails (BACS, CHAPS, SWIFT). Ensure systems meet high standards for availability, resilience, security and regulatory compliance. Own delivery outcomes in partnership with Product, Technology ...

Staff Engineer (Tech Lead)

Hiring Organisation
Talent Locker
Location
Warminster, Wiltshire, South West, United Kingdom
Employment Type
Permanent
Staff Engineer (Tech Lead) - Warminster, Hybrid - Security Cleared - Up to £120,000 We're looking for an experienced Staff Engineer/Technical Lead who thrives on solving complex engineering challenges and wants to play a ...

Senior Reliability Engineer – Platform & Observability

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Identity’s OneLogin team is seeking Senior Software Engineers who own production systems and drive reliability, observability, and operability across the stack. You’ll tackle complex issues, design for resilience, and apply AI-assisted development approaches to accelerate debugging and root-cause analysis. The role emphasizes end-to-end ownership ...

OpenSearch & Observability SRE – Hybrid (London)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
sector. The role is hybrid and initially a 6-month contract with strong prospects to extend. You will focus on OpenSearch deployment, monitoring, and observability across critical systems. You will design and operate OpenSearch environments, develop dashboards in Grafana, and support Geneos monitoring, automation, and SRE practices in collaboration with ...

Senior Lead SRE: Reliability, Observability & Resiliency

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
JPMorgan Chase & Co. is seeking a Senior Lead Site Reliability Engineer in Glasgow, Scotland. This role is pivotal for enhancing the reliability and observability of critical platforms. You will lead technical initiatives and contribute significantly to business impact through your expertise. The ideal candidate will have advanced proficiency in software ...

AWS SRE: Reliability, Observability & Cost Optimisation

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Site Reliability Engineer in London for a hybrid role, requiring 3+ years of SRE experience, especially in Kubernetes. Responsibilities include improving system reliability, observability, and cost efficiency. The ideal candidate will work closely with development and platform teams, and should be familiar with operational and incident workflows. This full-time ...

Senior Software Engineer II — Reliability & Observability

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
self-healing systems at scale, delivering platform tooling that engineers across the company adopt for their services. You will own incident management tooling, evolve observability infrastructure with SLOs and real-time signals, and contribute to AI-driven automation that reduces toil and speeds delivery. #J-18808-Ljbffr ...

Software Engineer, Observability

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
software, trusted by companies like OpenAI, PayPal, Ramp, Supreme, and millions of developers worldwide. We are looking for a Software Engineer to join our Observability team in London, UK. Vercel users rely on Observability to monitor and understand their applications’ health and behavior. In this hybrid role, you will design … implement, and maintain Observability products that meet high-quality standards and integrate seamlessly with the Vercel Platform. This role is based in London, UK, and includes in-office anchor days on Monday, Tuesday, and Friday for those within commuting distance. If located beyond, the role is fully remote.What You Will ...

Unix Engineer

Hiring Organisation
Wolviston Management Services
Location
Glasgow, City of Glasgow, United Kingdom
Employment Type
Permanent
server estates. Automation of infrastructure operations and maintenance activities. Development of orchestration workflows using Apache Airflow. Building automation tooling using Python and Ansible. Enhancing observability, monitoring and operational controls. Supporting the introduction of live-patching technologies. The role requires a strong infrastructure engineering mindset, with particular emphasis on UNIX/… orchestration workflows within Apache Airflow. Create automated pre-change validation and post-change verification processes. Improve efficiency and reduce manual intervention through automation. Platform Observability Implement monitoring and observability capabilities across automated infrastructure workflows. Develop dashboards, alerting and operational reporting. Work with tooling including: Prometheus Grafana Loki Ensure automation ...

Senior Dynatrace Observability Lead - Remote

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
Experis - ManpowerGroup is seeking a Senior Observability Engineer to lead the Dynatrace observability practice for a UK client. You will own architecture, design, deployment, and ongoing optimization, while mentoring junior consultants and coordinating with customer stakeholders. The role emphasizes hands-on delivery, design documentation, and proactive operations across cloud ...

Principal Machine Learning Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
data engineering teams to implement scalable data lakehouse oriented feature architectures and enterprise‐grade ML governance. Champion engineering standards for model quality, documentation, observability, and platform resilience. Feature Engineering & Data Architecture Architect highly scalable, production‐ready feature pipelines within Lakehouse environments. Set the technical direction for fallback and resilience strategies … including scoring metrics, latency, error analytics, and SLOs. Partner with platform teams to optimise cost, scale, and reliability of inference endpoints. Monitoring, Drift Detection & Observability Define observability standards for feature drift, concept drift, performance degradation, and data integrity. Lead the creation of dashboards, benchmarks, and automated alerting across ...

Senior Automation Engineer

Hiring Organisation
Raytheon
Location
Glenrothes, Fife, Scotland, United Kingdom
Employment Type
Permanent, Work From Home
commissioning of robotic cells and assembly systems; perform First Article Inspection (FAI) and ensure compliance with safety standards i.e. ISO 9001 or AS9100. Observability & Support: Maintain platform observability and respond to incidents through Root Cause Analysis (RCA) to improve service efficiency. System Integration: Designing and implementing interfaces between MES (e.g. ...

Staff Data Engineer – Data Quality & Governance

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
adjustments@depop.com. For any other non-disability related questions, please reach out to our Talent Partners. Role We’re building a Data Quality, Observability & Governance Team to improve the reliability, trust, and compliance of Depop’s data ecosystem. As a Staff Data Engineer in this team, you’ll lead … reduce the mean time to detection and resolution of data incidents, by establishing data contracts between producers and consumers, developing robust data observability systems, and embedding governance and GDPR compliance principles across the data lifecycle. You’ll collaborate with product engineering, data platform, analytics, and legal teams to build confidence ...

Software Engineer, Observability

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
shaping our story, you’ll help define what comes next. About the Role: We are looking for a Software Engineer to join our Observability team. Vercel users rely on Observability to monitor and understand their applications’ health and behavior. In this role, you will design, implement, and maintain Observability products … large-scale data ingestion, storage, and processing from distributed systems. Develop cutting-edge visualization tools to provide insights into application behavior and performance. Integrate observability features with popular frontend tools, frameworks, and build systems to enhance developer experience. Write clean, efficient, and well-documented code, ensuring platform reliability through thorough ...

Automation Engineer

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
Title: Automation and Observability Engineer Location: Glasgow (hybrid) Contract: 12-months (possible extensions) Pay rate: up to £500 p/d PAYE Are you an experienced Automation and Observability Engineer looking to make an impact in a global financial services environment? We are seeking a talented engineer to join … Enterprise Technology Services team, delivering high-quality automation and observability solutions to enhance data protection and system reliability. What you'll do: Work closely with internal teams to automate processes and improve alerting/observability solutions. Focus on enhancing the reliability of the data protection environment through automation or improved ...

Principal Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Wales, United Kingdom
working on Mondays and Fridays Take technical ownership of the platforms that support Centerprise services and customer operations. You will lead improvements in reliability, observability and automation while remaining closely involved in complex engineering, major incidents and service recovery. Role Summary As Principal Platform Engineer, you will … hours escalation when required. Identify and address technical debt, operational risk and platform weaknesses. Ensure services remain supportable, recoverable and operationally efficient. Observability and service health Own monitoring and observability tooling, standards and operational dashboards. Develop service health metrics that provide clear and useful operational insight. Improve the quality ...

Senior Software Engineer – Order Management Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
technical direction of the Order Management platform, driving improvements in architecture, integration patterns, automation, reliability, and security. Drive operational excellence by designing systems with observability, monitoring, resilience, and supportability at their core. Use tools such as Dynatrace to improve platform visibility, monitoring, and alerting capabilities. Collaborate closely with Product Managers … fulfilment operations. Strong Java engineer with experience designing and developing enterprise services, APIs, and integrations across distributed systems. Experienced in CI/CD practices, observability, production support, and operational excellence. Comfortable designing systems with resilience, recoverability, monitoring, and supportability in mind. A collaborative engineer who enjoys working across disciplines, raising ...

Staff Backend Engineer - Grafana Second Horizon | UK | Remote

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … remote opportunity, and we would be interested in applicants located in Spain, Sweden, UK, Ireland or Germany. The Opportunity: At Grafana Labs, we build observability tools that help users understand, respond to, and improve their systems – regardless of scale, complexity, or tech stack. We recently started a skunkworks initiative with ...

Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
robust security, access controls, secrets management, encryption, and audit trails directly into the infrastructure foundations to meet strict regulatory standards from day one. Data & Observability: Own the underlying data and event backbone required for transactional correctness, alongside the observability stack needed for rapid, real‐time issue detection and resolution. Qualifications ...

Staff Software Engineer (Reliability & Platform)

Hiring Organisation
Jobleads-UK
Location
United Kingdom
rather than handing work off between silos. In this role, you will focus on improving production operability and reliability by debugging complex issues, strengthening observability, and eliminating recurring problems at the source. Responsibilities Own production operability by debugging complex issues, improving system visibility, and eliminating recurring problems at the source … resolve (MTTR), and recurrence rates for issues Identify systemic issues and eliminate recurring problems through code fixes, architecture improvements, and better operational tooling Improve observability across services — logs, metrics, and alerting — for faster diagnosis and resolution Design and improve debugging workflows, runbooks, and internal tooling for engineers Reduce operational burden ...

SRE Engineer

Hiring Organisation
Pinnacle Technical Resources
Location
Jersey City, New Jersey, United States
Employment Type
Permanent
Salary
USD 65 Annual
business processes; i.e., the ability to understand 'the business' and the impact of technology solutions Experience working with thirdparty applications and integrations Familiarity with observability practices such as white and black box monitoring, service level objective alerting, and telemetry collection using tools like Grafana, Dynatrace, Prometheus, Datadog and Splunk. Recognize … eliminate toil through systems engineering or automation, and implement observability patterns to improve service level indicators, objectives monitoring, and alerting solutions. Note: Pay Range: $60 - $65 The specific compensation for this position will be determined by a number of factors, including the scope, complexity and location of the role ...

SRE Managing Consultant

Hiring Organisation
Akkodis
Location
City, London, United Kingdom
Employment Type
Permanent
Salary
GBP 90,000 - 100,000 Annual
include: Define and embed SRE engagement models aligned to modern engineering and traditional ITSM/ITIL practices Establish SLIs, SLOs, and Error Budgets Shape observability strategies using metrics, logs, and traces Design incident response models and post-incident learning loops Reduce toil through automation and engineering excellence Deliver SRE capability … Looking For Extensive experience in SRE, cloud operations, or DevOps Proven consulting or advisory background Experience with AWS, Azure, or GCP Strong observability and incident management expertise Ability to obtain UK SC clearance Modis International Ltd acts as an employment agency for permanent recruitment and an employment business ...

Senior Director Technology

Hiring Organisation
Jobleads-UK
Location
Langley, England, United Kingdom
buildsourproductionAWSplatform, and onboards development teams and products to this platform.You’llbe responsible fordriving the move toinfrastructure as code, account management, automated pipelines, observability, resilience, operating coverage and cost control. You’llalso help define how the platform supports AI, data engineering and high-scale product demand, including where AWS Bedrock … standards for infrastructure as code, CI/CD, AWS account management, platform guardrails and developer enablement. Improve the operational model for the platform, including observability, incident response, reliability and 24/7 support coverage. Partner with Product and Commercial teams to get ahead of major demand changes, customer commitments ...

Staff Security Engineer

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
engineer who enjoys solving complex security challenges at scale. You’ll work across engineering, data, AI and digital workplace teams to build security observability, automate control assurance and influence how security is embedded into products, platforms and processes. If you're passionate about turning security data into actionable insight, building … raising the security maturity of a fast-moving technology organisation, we'd love to hear from you. About the role Designing and building security observability capabilities that provide meaningful visibility across systems, infrastructure and applications Developing automated control monitoring, evidence collection and continuous testing solutions that strengthen security governance Partnering ...

Operations Engineer

Hiring Organisation
ASCENT PROFESSIONAL SERVICES LTD
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£60,000
continuity. Key Responsibilities Provide operational support for enterprise platforms, applications, integrations, and associated technologies. Monitor system health, availability, and performance using monitoring, alerting, and observability tools. Analyse, troubleshoot, and resolve incidents affecting services and platforms. Perform root cause analysis and contribute to implementing permanent solutions to prevent recurring issues. Coordinate … within IT operations, support engineering, or service management environments. Experience supporting business-critical production services and operational platforms. Knowledge of monitoring, logging, alerting, and observability practices. Experience working with incident, problem, change, and release management processes. Excellent communication skills with the ability to collaborate effectively across multiple technical and business ...