776 to 800 of 4,719 Observability Jobs

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Cardiff, South Glamorgan, Wales, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Darlington, County Durham, North East, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Belfast, County Antrim, Northern Ireland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Salford, Greater Manchester, North West, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Value Stream Engineering Lead

Location
Greater London, England, United Kingdom
fostering a culture of collaboration and continuous improvement Own and continuously improve engineering standards and practices across code quality, testing, CI/CD, observability, performance, DevOps, automation and secure, resilient system design Partner closely with Product and Architecture to optimise the flow of value from idea to delivery, identifying blockers … coaching engineers through technical challenges in real time Strong ownership of modern engineering standards and practices, including code quality, testing strategy, CI/CD, observability, performance optimisation, DevOps and secure, resilient system design Experience designing and delivering within modern cloud environments such as AWS, Azure or GCP, with a strong ...

Software Engineer (Infrastructure)

Location
Greater London, England, United Kingdom
engineering team in building reliable, scalable applications Help design and build tools to make our services scalable and highly available, automating wherever possible Develop observability and orchestration tooling to allow our customers to monitor and manage their deployments Contribute to disaster recovery, backup and redundancy tooling and strategy Assist … hands‐on experience Experience developing production‐ready infrastructure management tooling with either Python or Golang Familiarity with at least one of the following: Observability Tools (e.g. Prometheus, OpenTelemetry, Grafana) Databases (e.g. Postgres, DuckDB) Event Streaming platforms (e.g. Kafka) Container Orchestration (e.g. Docker, Kubernetes) Familiarity with cloud platforms such ...

Sr Lead Software Engineer - Java / Python

Location
Greater London, England, United Kingdom
code written by others Leads technical analysis, estimation, planning, code reviews, architecture sessions, and retrospectives to drive delivery outcomes Establishes reliability goals and implements observability, resilience patterns, and operational readiness practices Leads incident response and post-incident reviews to improve production stability and performance; identifies recurring issues and drives automation … teams Preferred qualifications, capabilities, and skills Exposure to messaging systems and market protocols (e.g., MQ/Kafka; familiarity with FIX and Solace) Experience with observability stacks and resilience engineering for low‐latency/latency‐sensitive platforms Familiarity with Python Experience operating services in regulated environments with strong auditability and controls ...

Developer Experience Engineer

Location
Greater London, England, United Kingdom
with other engineering teams, informing us of our future project work and other improvements. In this role, you will navigate novel build challenges, refine observability and alerting systems for firm-wide critical services, and build great developer workflows. Your efforts will not only advance our team's delivery capabilities … Investigate and resolve intricate systems or build issues during support rotations, building valuable relationships with other DRW development teams. Improve service reliability, performance, and observability of firm-wide critical infrastructure and tooling. Work in a hybrid infrastructure environment, cloud and on-prem, using infra-as-code patterns to reliably manage ...

Sr Lead Software Engineer - Java / Python

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
debugs code written by othersLeads technical analysis, estimation, planning, code reviews, architecture sessions, and retrospectives to drive delivery outcomesEstablishes reliability goals and implements observability, resilience patterns, and operational readiness practicesLeads incident response and post-incident reviews to improve production stability and performance; identifies recurring issues and drives automation/remediationUpholds … globally distributed teamsPreferred qualifications, capabilities, and skillsExposure to messaging systems and market protocols (e.g., MQ/Kafka; familiarity with FIX and Solace)Experience with observability stacks and resilience engineering for low-latency/latency-sensitive platformsFamiliarity with PythonExperience operating services in regulated environments with strong auditability and controlsJ.P. Morgan ...

Software Engineer Lead - Site Reliability

Location
Telford, England, United Kingdom
practical improvements across the Customer Digital Platform. You’ll strengthen our DevOps foundations across Azure API platform integration, cloud infrastructure, CI/CD automation, observability, deployment safety and operational readiness, working closely with internal squads, strategic partners and third parties. Job Type: Permanent Location: Birmingham, London, Telford, Edinburgh (Hybrid working … operational readiness. Good awareness of Site Reliability Engineering principles, with the ability to apply them pragmatically to DevOps and platform improvement work. Experience improving observability through monitoring, logging, tracing, alerting, service-health dashboards and actionable telemetry. Experience improving resilience through failover, graceful degradation, recovery testing and practical reliability improvements. Understanding ...

Lead SRE - AWS,Python

Hiring Organisation
JP Morgan Chase
Location
United Kingdom
Salary
£ 70 K
accelerate delivery velocityCollaborate cross-functionally with software engineering, architecture, and security teams to embed reliability and resiliency principles early in the design processChampion observability practices by building and maintaining monitoring, alerting, and dashboarding solutions that provide actionable insights into system healthMentor and guide junior engineers, fostering a culture of continuous … Java, Bash) for automation and tooling developmentExperience defining and managing service level indicators, service level objectives, and error budgets in production environmentsStrong background in observability tooling, including metrics, logging, and distributed tracing platformsDemonstrated experience leading incident response processes, conducting blameless post-mortems, and driving systemic reliability improvementsExperience with container orchestration ...

Lead SRE - AWS,Python

Location
Glasgow, Scotland, United Kingdom
delivery velocity Collaborate cross-functionally with software engineering, architecture, and security teams to embed reliability and resiliency principles early in the design process Champion observability practices by building and maintaining monitoring, alerting, and dashboarding solutions that provide actionable insights into system health Mentor and guide junior engineers, fostering a culture … automation and tooling development Experience defining and managing service level indicators, service level objectives, and error budgets in production environments Strong background in observability tooling, including metrics, logging, and distributed tracing platforms Demonstrated experience leading incident response processes, conducting blameless post-mortems, and driving systemic reliability improvements Experience with container ...

Platform Chapter Lead - Engineering

Location
Greater London, England, United Kingdom
paved paths, less friction — using metrics (e.g. DORA) and real feedback to keep improving it. Keep it reliable and compliant. Oversee performance, resilience and observability for revenue‐critical services through peak traffic, and maintain security and compliance (e.g. PCI‐DSS, GDPR). Bring the business with you. Align platform strategy … balance cost, speed and risk. Strong cloud‐native and modern DevOps background — cloud (ideally AWS) and Kubernetes, CI/CD, infrastructure as code and observability — with enough engineering depth (e.g. Java, .NET, Python) to be credible with strong engineers. Excellent communication and the ability to influence at executive level, plus ...

Senior Software Engineer, Full-Stack Applications (Python)

Hiring Organisation
Fitch Ratings
Location
Manchester, Greater Manchester, United Kingdom
Salary
£ 70 K
Apache Airflow for workflow management, or Streamlit for building interactive data applications• Advanced Data Management – Strong SQL design, query optimization, and database architecture expertise• Observability – Experience with observability patterns and tools like Datadog, distributed tracing, monitoring, and logging best practices• DevOps and Infrastructure – Familiarity with ArgoCD for GitOps and Security ...

Software Engineer — Staff / Senior

Location
Greater London, England, United Kingdom
system. AI harnesses in production. Deploying agentic AI into real-world operational settings — acting on real money, tenancies and legal exposure, with the guardrails, observability and correctness that demands. Non-deterministic LLM working within compliant, secure deterministic software. Our stack We build on NestJS + TypeScript on GCP/… build Small, well-factored services. TDD and DDD as defaults. Trunk-based CI/CD — you ship to production and own it, with tests, observability and clean rollbacks. Lean frameworks, readable code, and we move fast because the tests and boundaries let us. What we're looking for 6+ years ...

DevOps Engineer

Location
Belfast City, Northern Ireland, United Kingdom
around the world. Working closely with our engineering teams, you’ll design and maintain cloud infrastructure, improve our CI/CD pipelines, and enhance observability so we can ship high‐quality features quickly and confidently. You’ll bring your initiative as well as your technical skills to solve real operational … improve CI/CD pipelines to support rapid, high‐quality deployments Monitor and improve system availability, performance, and cost‐efficiency Implement and manage observability tools (logging, metrics, tracing) Enhance infrastructure‐as‐code using AWS CDK and related tools Collaborate with engineers to streamline development workflows and deployment strategies Champion DevOps ...

Senior DevOps Engineer

Location
Greater London, England, United Kingdom
based teams on anything related to the technical aspects of our platform. What you'll do Help develop and maintain a robust monitoring and observability framework to ensure system performance, reliability, and early issue detection. Leverage your experience in DevOps to design, maintain, and improve cloud infrastructure. Own and evolve … roles. You’re an IT generalist who’s comfortable shaping and defining the role with us. You have a strong understanding of monitoring and observability tools (e.g., Datadog or Grafana). You have experience applying SRE principles such as SLOs, and error budgets. You have worked with incident management processes ...

Lead Software Engineer - Application Owner

Hiring Organisation
JP Morgan Chase
Location
Bournemouth, Dorset, United Kingdom
Salary
£ 70 K
execution, after-action reviews, and closure of follow-up actions.Define and continuously improve production readiness standards, including release safety and rollback strategy, dependency awareness, observability requirements, and operational runbooks.Contribute hands-on to design and delivery, including system design, code reviews, automation, and complex troubleshooting, with secure-by-design and reliable … applications with strong operational accountability, including controls, resiliency and recovery, and remediation tracking.Strong system design fundamentals and cloud-native operational patterns, including scalability, reliability, observability, and dependency management.Hands-on experience with Go-based services and modern CI/CD practices.Experience operating workloads on AWS and Kubernetes ...

Software Engineer - Data & AI

Hiring Organisation
McLaren Group
Location
Woking, Surrey, UK
Employment Type
Full-time
source systems and downstream consumers at every integration boundary. Own the infrastructure your services run on: infrastructure-as-code, CI/CD, containerisation, and observability and deployment. Contribute to AI evaluation and quality: build and run automated evals and regression tests and use observability tooling to track latency, token usage ...

Lead Software Engineer - Application Owner

Hiring Organisation
JP Morgan Chase
Location
Bournemouth, Dorset, UK
Employment Type
Full-time
after-action reviews, and closure of follow-up actions. Define and continuously improve production readiness standards, including release safety and rollback strategy, dependency awareness, observability requirements, and operational runbooks. Contribute hands-on to design and delivery, including system design, code reviews, automation, and complex troubleshooting, with secure-by-design … with strong operational accountability, including controls, resiliency and recovery, and remediation tracking. Strong system design fundamentals and cloud-native operational patterns, including scalability, reliability, observability, and dependency management. Hands-on experience with Go-based services and modern CI/CD practices. Experience operating workloads on AWS and Kubernetes ...

Senior Lead Software Data Engineer – Corporate Know Your Customer

Location
Auchentibber, Scotland, United Kingdom
scale data processing, microservices, API design, and orchestration frameworks Working knowledge of relational and NoSQL databases, vector stores, and data lake architectures Familiarity with observability tools and frameworks Practical cloud-native experience (AWS, Azure, or GCP) Ability to communicate effectively with senior leaders and executives Commitment to inclusive, collaborative teamwork … catalog services such as Apache Iceberg Experience with LLM orchestration frameworks and model serving infrastructure or managed endpoints Familiarity with AI evaluation and observability practices for LLM workloads Understanding of agentic design patterns and how to constrain agent autonomy in financial workflows Interest in emerging technologies and continuous learning Employer ...

Senior Data Engineer

Location
Greater London, England, United Kingdom
Analytics, Salesforce, Architecture) to define requirements and translate them into scalable technical solutions. Drive continuous improvement of data engineering practices, including CI/CD, observability, testing frameworks, and documentation standards. Provide technical leadership through mentoring, code reviews, and guidance to junior team members, fostering engineering excellence. Ensure compliance with security … Salesforce data models and API integrations. Experience using AWS CDK for infrastructure deployment. Familiarity with orchestration tools such as Airflow. Experience implementing data observability, monitoring and alerting solutions. Knowledge of BI platforms such as Tableau and how data products are consumed by end users. Exposure to MLOps practices, including supporting ...

SRE Engineer

Hiring Organisation
Ricoh
Location
London, UK
Employment Type
Full-time
practices, tooling, and engineering standardsDriving infrastructure‐as‐code and automation across Azure and on‐premImproving image bakery pipelines for secure, repeatable server buildsEmbedding observability using metrics, logs, traces, and effective alertingEnsuring all practices align with ISO 27001 and internal security frameworksManaging automated patching, vulnerability remediation and configuration complianceBuilding dashboards … code (Terraform, ARM/Bicep), configuration management (Ansible, PowerShell DSC), and CI/CD tooling (Azure DevOps, GitHub Actions)Experience with monitoring and observability stacksSolid understanding of OS fundamentals (Windows/Linux), security, networkingBackground in scripting or software development (PowerShell, Python, Go)Experience with containers and orchestration (Docker, Kubernetes ...

Senior AI Engineer

Location
Greater London, England, United Kingdom
passthrough and logic centre for enterprise datasets, consuming other MCP servers and presenting them through one governed interface; and the orchestration, evaluation, and observability services beneath Libros and PRISM, two of the projects we are delivering with Percepta. As a Senior Applied AI Engineer, reporting to the Principal AI Engineer … own. Platforms built jointly with Percepta transfer into our ownership with maintainable designs and a team that can extend them without external help. Evaluation, observability, and control are built into what you ship rather than bolted on before release. Your responsibilities Build and run our AI products and platform Design ...

Platform Engineer - Edinburgh

Location
City of Edinburgh, Scotland, United Kingdom
code generation, testing, documentation, and analysis, while understanding model limitations, protecting client data, and improving delivery quality and speed through pragmatic automation SRE & Observability You’ll bring a reliability mindset to delivery, designing services that are operable by default and measured through meaningful SLIs/SLOs. You’ll help teams … implement pragmatic observability—logging, metrics, and distributed tracing—with actionable alerting, and you’ll contribute to (or lead) incident response and post-incident reviews that drive learning and measurable improvements. We are looking for experience in the following skills Strong experience with the AWS cloud platform and core services. Hands ...