951 to 975 of 6,285 Permanent Observability Jobs

Sr Lead Software Engineer - Java / Python

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
debugs code written by othersLeads technical analysis, estimation, planning, code reviews, architecture sessions, and retrospectives to drive delivery outcomesEstablishes reliability goals and implements observability, resilience patterns, and operational readiness practicesLeads incident response and post-incident reviews to improve production stability and performance; identifies recurring issues and drives automation/remediationUpholds … globally distributed teamsPreferred qualifications, capabilities, and skillsExposure to messaging systems and market protocols (e.g., MQ/Kafka; familiarity with FIX and Solace)Experience with observability stacks and resilience engineering for low-latency/latency-sensitive platformsFamiliarity with PythonExperience operating services in regulated environments with strong auditability and controlsJ.P. Morgan ...

Lead SRE - AWS,Python

Location
Glasgow, Scotland, United Kingdom
delivery velocity Collaborate cross-functionally with software engineering, architecture, and security teams to embed reliability and resiliency principles early in the design process Champion observability practices by building and maintaining monitoring, alerting, and dashboarding solutions that provide actionable insights into system health Mentor and guide junior engineers, fostering a culture … automation and tooling development Experience defining and managing service level indicators, service level objectives, and error budgets in production environments Strong background in observability tooling, including metrics, logging, and distributed tracing platforms Demonstrated experience leading incident response processes, conducting blameless post‐mortems, and driving systemic reliability improvements Experience with container ...

Software Engineer Lead - Site Reliability

Location
Telford, England, United Kingdom
practical improvements across the Customer Digital Platform. You’ll strengthen our DevOps foundations across Azure API platform integration, cloud infrastructure, CI/CD automation, observability, deployment safety and operational readiness, working closely with internal squads, strategic partners and third parties. Job Type: Permanent Location: Birmingham, London, Telford, Edinburgh (Hybrid working … operational readiness. Good awareness of Site Reliability Engineering principles, with the ability to apply them pragmatically to DevOps and platform improvement work. Experience improving observability through monitoring, logging, tracing, alerting, service-health dashboards and actionable telemetry. Experience improving resilience through failover, graceful degradation, recovery testing and practical reliability improvements. Understanding ...

Sr Lead Software Engineer - Java / Python

Location
Greater London, England, United Kingdom
code written by others Leads technical analysis, estimation, planning, code reviews, architecture sessions, and retrospectives to drive delivery outcomes Establishes reliability goals and implements observability, resilience patterns, and operational readiness practices Leads incident response and post-incident reviews to improve production stability and performance; identifies recurring issues and drives automation … teams Preferred qualifications, capabilities, and skills Exposure to messaging systems and market protocols (e.g., MQ/Kafka; familiarity with FIX and Solace) Experience with observability stacks and resilience engineering for low-latency/latency-sensitive platforms Familiarity with Python Experience operating services in regulated environments with strong auditability and controls ...

Lead SRE - AWS,Python

Hiring Organisation
JP Morgan Chase
Location
United Kingdom
Salary
£ 70 K
accelerate delivery velocityCollaborate cross-functionally with software engineering, architecture, and security teams to embed reliability and resiliency principles early in the design processChampion observability practices by building and maintaining monitoring, alerting, and dashboarding solutions that provide actionable insights into system healthMentor and guide junior engineers, fostering a culture of continuous … Java, Bash) for automation and tooling developmentExperience defining and managing service level indicators, service level objectives, and error budgets in production environmentsStrong background in observability tooling, including metrics, logging, and distributed tracing platformsDemonstrated experience leading incident response processes, conducting blameless post-mortems, and driving systemic reliability improvementsExperience with container orchestration ...

Software Engineering Manager - Time & Scheduling

Hiring Organisation
Marks & Spencer
Location
United Kingdom, UK
Employment Type
Full-time
Partnering with product, architecture and engineering leaders to define technical roadmaps, manage dependencies and deliver business outcomes. Driving quality, reliability and operational excellence through observability, resilience, DevOps practices and strong service ownership. Who you areYour skills and experience will include: Previous senior-level, hands-on software engineering experience, with … Messaging Kotlin, Java, Spring and Spring Boot MongoDB, SQL Server and Redis CI/CD, Infrastructure as Code and Automated Testing Dynatrace and Observability ToolingWhat's in it for you? Working at M&S means being part of something bigger - helping to deliver quality, value and service to millions ...

Platform Chapter Lead - Engineering

Location
Greater London, England, United Kingdom
paved paths, less friction — using metrics (e.g. DORA) and real feedback to keep improving it. Keep it reliable and compliant. Oversee performance, resilience and observability for revenue‐critical services through peak traffic, and maintain security and compliance (e.g. PCI‐DSS, GDPR). Bring the business with you. Align platform strategy … balance cost, speed and risk. Strong cloud‐native and modern DevOps background — cloud (ideally AWS) and Kubernetes, CI/CD, infrastructure as code and observability — with enough engineering depth (e.g. Java, .NET, Python) to be credible with strong engineers. Excellent communication and the ability to influence at executive level, plus ...

Senior Software Engineer, Full-Stack Applications (Python)

Hiring Organisation
Fitch Ratings
Location
Manchester, Greater Manchester, United Kingdom
Salary
£ 70 K
Apache Airflow for workflow management, or Streamlit for building interactive data applications• Advanced Data Management – Strong SQL design, query optimization, and database architecture expertise• Observability – Experience with observability patterns and tools like Datadog, distributed tracing, monitoring, and logging best practices• DevOps and Infrastructure – Familiarity with ArgoCD for GitOps and Security ...

Software Engineer — Staff / Senior

Location
Greater London, England, United Kingdom
system. AI harnesses in production. Deploying agentic AI into real-world operational settings — acting on real money, tenancies and legal exposure, with the guardrails, observability and correctness that demands. Non-deterministic LLM working within compliant, secure deterministic software. Our stack We build on NestJS + TypeScript on GCP/… build Small, well-factored services. TDD and DDD as defaults. Trunk-based CI/CD — you ship to production and own it, with tests, observability and clean rollbacks. Lean frameworks, readable code, and we move fast because the tests and boundaries let us. What we're looking for 6+ years ...

DevOps Engineer

Location
Belfast City, Northern Ireland, United Kingdom
around the world. Working closely with our engineering teams, you’ll design and maintain cloud infrastructure, improve our CI/CD pipelines, and enhance observability so we can ship high‐quality features quickly and confidently. You’ll bring your initiative as well as your technical skills to solve real operational … improve CI/CD pipelines to support rapid, high‐quality deployments Monitor and improve system availability, performance, and cost‐efficiency Implement and manage observability tools (logging, metrics, tracing) Enhance infrastructure‐as‐code using AWS CDK and related tools Collaborate with engineers to streamline development workflows and deployment strategies Champion DevOps ...

Lead Software Engineer - Application Owner

Hiring Organisation
JP Morgan Chase
Location
Bournemouth, Dorset, United Kingdom
Salary
£ 70 K
execution, after-action reviews, and closure of follow-up actions.Define and continuously improve production readiness standards, including release safety and rollback strategy, dependency awareness, observability requirements, and operational runbooks.Contribute hands-on to design and delivery, including system design, code reviews, automation, and complex troubleshooting, with secure-by-design and reliable … applications with strong operational accountability, including controls, resiliency and recovery, and remediation tracking.Strong system design fundamentals and cloud-native operational patterns, including scalability, reliability, observability, and dependency management.Hands-on experience with Go-based services and modern CI/CD practices.Experience operating workloads on AWS and Kubernetes ...

Senior DevOps Engineer

Location
Greater London, England, United Kingdom
based teams on anything related to the technical aspects of our platform. What you'll do Help develop and maintain a robust monitoring and observability framework to ensure system performance, reliability, and early issue detection. Leverage your experience in DevOps to design, maintain, and improve cloud infrastructure. Own and evolve … roles. You’re an IT generalist who’s comfortable shaping and defining the role with us. You have a strong understanding of monitoring and observability tools (e.g., Datadog or Grafana). You have experience applying SRE principles such as SLOs, and error budgets. You have worked with incident management processes ...

Software Engineer - Data & AI

Hiring Organisation
McLaren Group
Location
Woking, Surrey, UK
Employment Type
Full-time
source systems and downstream consumers at every integration boundary. Own the infrastructure your services run on: infrastructure-as-code, CI/CD, containerisation, and observability and deployment. Contribute to AI evaluation and quality: build and run automated evals and regression tests and use observability tooling to track latency, token usage ...

Lead Software Engineer - Application Owner

Hiring Organisation
JP Morgan Chase
Location
Bournemouth, Dorset, UK
Employment Type
Full-time
after-action reviews, and closure of follow-up actions. Define and continuously improve production readiness standards, including release safety and rollback strategy, dependency awareness, observability requirements, and operational runbooks. Contribute hands-on to design and delivery, including system design, code reviews, automation, and complex troubleshooting, with secure-by-design … with strong operational accountability, including controls, resiliency and recovery, and remediation tracking. Strong system design fundamentals and cloud-native operational patterns, including scalability, reliability, observability, and dependency management. Hands-on experience with Go-based services and modern CI/CD practices. Experience operating workloads on AWS and Kubernetes ...

Senior Lead Software Data Engineer - Corporate Know Your Customer

Location
Glasgow, Scotland, United Kingdom
scale data processing, microservices, API design, and orchestration frameworks Working knowledge of relational and NoSQL databases, vector stores, and data lake architectures Familiarity with observability tools and frameworks Practical cloud-native experience (AWS, Azure, or GCP) Ability to communicate effectively with senior leaders and executives Commitment to inclusive, collaborative teamwork … catalog services such as Apache Iceberg Experience with LLM orchestration frameworks and model serving infrastructure or managed endpoints Familiarity with AI evaluation and observability practices for LLM workloads Understanding of agentic design patterns and how to constrain agent autonomy in financial workflows Interest in emerging technologies and continuous learning Employer ...

Senior Data Engineer

Location
Greater London, England, United Kingdom
Analytics, Salesforce, Architecture) to define requirements and translate them into scalable technical solutions. Drive continuous improvement of data engineering practices, including CI/CD, observability, testing frameworks, and documentation standards. Provide technical leadership through mentoring, code reviews, and guidance to junior team members, fostering engineering excellence. Ensure compliance with security … Salesforce data models and API integrations. Experience using AWS CDK for infrastructure deployment. Familiarity with orchestration tools such as Airflow. Experience implementing data observability, monitoring and alerting solutions. Knowledge of BI platforms such as Tableau and how data products are consumed by end users. Exposure to MLOps practices, including supporting ...

Software Engineer III - Python

Location
Glasgow, Scotland, United Kingdom
infrastructure-as-code using Terraform within established team patterns across modules, environments, and state management Improve operability of services by adding and using observability tooling including logs, metrics, traces, dashboards, and alerts, and participate in incident response and root-cause analysis Leverage enterprise-authorized AI coding assist tools within … implementing application logic and APIs on top of relational data Experience building APIs and microservices using REST or gRPC, including contracts, security basics, and observability Practical experience delivering LLM-based features as part of software systems, with familiarity with agentic patterns Working knowledge of delivery and operations including CI/ ...

SRE Engineer

Hiring Organisation
Ricoh
Location
London, UK
Employment Type
Full-time
practices, tooling, and engineering standardsDriving infrastructure‐as‐code and automation across Azure and on‐premImproving image bakery pipelines for secure, repeatable server buildsEmbedding observability using metrics, logs, traces, and effective alertingEnsuring all practices align with ISO 27001 and internal security frameworksManaging automated patching, vulnerability remediation and configuration complianceBuilding dashboards … code (Terraform, ARM/Bicep), configuration management (Ansible, PowerShell DSC), and CI/CD tooling (Azure DevOps, GitHub Actions)Experience with monitoring and observability stacksSolid understanding of OS fundamentals (Windows/Linux), security, networkingBackground in scripting or software development (PowerShell, Python, Go)Experience with containers and orchestration (Docker, Kubernetes ...

Senior AI Engineer

Location
Greater London, England, United Kingdom
passthrough and logic centre for enterprise datasets, consuming other MCP servers and presenting them through one governed interface; and the orchestration, evaluation, and observability services beneath Libros and PRISM, two of the projects we are delivering with Percepta. As a Senior Applied AI Engineer, reporting to the Principal AI Engineer … own. Platforms built jointly with Percepta transfer into our ownership with maintainable designs and a team that can extend them without external help. Evaluation, observability, and control are built into what you ship rather than bolted on before release. Your responsibilities Build and run our AI products and platform Design ...

Platform Engineer - Edinburgh

Location
City of Edinburgh, Scotland, United Kingdom
code generation, testing, documentation, and analysis, while understanding model limitations, protecting client data, and improving delivery quality and speed through pragmatic automation SRE & Observability You’ll bring a reliability mindset to delivery, designing services that are operable by default and measured through meaningful SLIs/SLOs. You’ll help teams … implement pragmatic observability—logging, metrics, and distributed tracing—with actionable alerting, and you’ll contribute to (or lead) incident response and post-incident reviews that drive learning and measurable improvements. We are looking for experience in the following skills Strong experience with the AWS cloud platform and core services. Hands ...

Lead Java Developer

Location
Greater London, England, United Kingdom
adoption and ensure successful rollout of new capabilities. Lead root cause analysis on production issues, drive long‐term stability improvements, and strengthen monitoring and observability across the platform. Recommended Experience Strong experience in Core Java, J2EE, Spring Framework Exposure to Python scripting and data analysis Experience in fast moving Capital … such as Kafka, JMS, gRPC etc. Proficient in latency measurement and performance optimization of Java based platforms with focus on JVM tuning Experience with observability stacks like ELK, Prometheus, Grafana, Kiali, Jaeger etc. Sound knowledge for persistence technologies such as relational databases, NoSQL databases, off heap storages and distributed caches ...

Corporate KYC Sr Lead Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
scale data processing, microservices, API design, and orchestration frameworks Working knowledge of relational and NoSQL databases, vector stores, and data lake architectures Familiarity with observability tools and frameworks Practical cloud-native experience (AWS, Azure, or GCP) Ability to communicate effectively with senior leaders and executives Commitment to inclusive, collaborative teamwork … catalog services such as Apache Iceberg Experience with LLM orchestration frameworks and model serving infrastructure or managed endpoints Familiarity with AI evaluation and observability practices for LLM workloads Understanding of agentic design patterns and how to constrain agent autonomy in financial workflows Interest in emerging technologies and continuous learning Employer ...

Lead SRE

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … knowledge of microservice infrastructure components, including service discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes ...

Lead SRE - Chase UK

Location
London, United Kingdom
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … knowledge of microservice infrastructure components, including service discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes ...

Lead SRE

Location
Westminster, West End, United Kingdom
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … knowledge of microservice infrastructure components, including service discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes ...