51 to 75 of 100 Observability Jobs in Glasgow

Senior Lead Site Reliability Engineer

Location
Glasgow, Scotland, United Kingdom
integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch reliability and observability for our most critical platforms. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, you are an integral part of an agile … significant business impact through your capabilities and contributions, and apply deep technical expertise and problem-solving methodologies to tackle a diverse array of reliability, observability, and performance challenges that span multiple technologies and applications. Job responsibilities Regularly provides technical guidance and direction on site reliability practices to support the business ...

DevOps Engineer - Glasgow

Location
Glasgow, Scotland, United Kingdom
production and infrastructure issues. Experience collaborating across multiple technical teams and stakeholders. Strong understanding of security principles within DevOps environments. Desirable Experience with SRE Observability tools and practices. Knowledge of Configuration & Release Management processes. Experience with Cloud Automation DevOps frameworks and tooling. Previous experience within financial services or other highly ...

AWS/DevOps Engineer with .Net/C#

Location
Glasgow, Scotland, United Kingdom
infrastructure. Ensure high system availability, performance, scalability, and reliability. Participate in incident management, root cause analysis, and problem resolution. Implement proactive monitoring, alerting, and observability solutions. Reduce operational overhead through automation and self-healing mechanisms. Support production releases and deployment activities. Design, deploy, and manage AWS-based infrastructure and services. ...

AWS Infrastructure Engineer II — Terraform, EKS

Location
Glasgow, Scotland, United Kingdom
based solutions, Terraform‐driven deployments, and CI/CD pipelines. You will grow into server‐side Java/Kotlin development and contribute to observability, security, and reliability across the Portfolio Management estate. #J-18808-Ljbffr ...

Senior Cloud Platform Engineer - Python/Go Lead

Location
Glasgow, Scotland, United Kingdom
talent to deliver modern cloud solutions and drive AI-enabled tooling across multi-cloud environments. You will lead architecture, CI/CD evolution, observability, and platform modernization while advancing DevEx practices and developer productivity. #J-18808-Ljbffr ...

Lead SRE: AWS & Python for Scalable Reliability

Location
Glasgow, Scotland, United Kingdom
embed reliability into the software development lifecycle and deliver resilient services for millions of users. You will lead incident response, define SLOs, champion observability, and mentor junior engineers, while shaping automation, CI/CD pipelines, and platform tooling to accelerate #J-18808-Ljbffr ...

VP DevOps/SRE Lead — Cloud Infra, Kubernetes & CI/CD

Location
Glasgow, Scotland, United Kingdom
DevOps/SRE Engineer - Vice President to own automation, reliability and production operations of AI/ML platforms. You will build CI/CD, observability, and incident-management practices that keep services stable across international markets. As part of the IPB Tech AIML team, you will lead reliability engineering ...

Lead Site Reliability Engineer - Observability & Resilience

Location
Glasgow, Scotland, United Kingdom
JPMorgan Chase & Co. seeks a Lead Site Reliability Engineer to define the future of reliability for a global firm. You will lead critical resiliency design reviews, break complex problems into actionable work, and mentor engineers ...

Sr Lead Infrastructure Engineer- Devops/AWS

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
machine learning products. As the team takes end-to-end ownership of the platforms it runs, you will build the CI/CD, observability, and incident-management practices that keep those services stable, secure, and performant across international markets. This is a Vice President-level role and an integral part … pipelines, release automation, and deployment toolingEstablishes reliability practices (SLOs, error budgets, runbooks) and leads production incident response and post-incident reviewBuilds and operates observability across the team's AI/ML services (metrics, logging, tracing, alerting)Automates infrastructure provisioning and configuration through infrastructure-as-codeImplements operational security, secrets management ...

Sr Lead Infrastructure Engineer- Devops/AWS

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
machine learning products. As the team takes end-to-end ownership of the platforms it runs, you will build the CI/CD, observability, and incident-management practices that keep those services stable, secure, and performant across international markets. This is a Vice President-level role and an integral part … release automation, and deployment tooling Establishes reliability practices (SLOs, error budgets, runbooks) and leads production incident response and post-incident review Builds and operates observability across the team's AI/ML services (metrics, logging, tracing, alerting) Automates infrastructure provisioning and configuration through infrastructure-as-code Implements operational security, secrets ...

Software Engineer II (AWS Infrastructure)

Location
Glasgow, Scotland, United Kingdom
Contribute to the build and maintenance of CI/CD pipelines (e.g., Jenkins) that apply infrastructure changes safely and reliably Instrument the estate for observability using tools such as Datadog, CloudWatch, and Dynatrace, and use telemetry insights to support improvements to infrastructure hygiene Uses enterprise-authorized AI capabilities within … supporting cluster operations Practical experience working with cloud infrastructure on AWS in a production or near‐production environment Familiarity with production readiness practices including observability (metrics, tracing, logging) and incident management in distributed systems Working knowledge of using enterprise-authorized AI capabilities within the work environment to support software engineering ...

Software Engineer III - AI/ML Platform Reliability

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
enhance the reliability and scalability of AI/ML platforms and applications to accommodate fast-growing demands. Own NFRs and develop tooling for observability, security, resilience, infrastructure management and operations excellence. Build and maintain scalable infrastructure to support the deployment and operation of large-scale AI platforms and apps. Build … architecture. Experience building large scale infrastructure and and cloud-native delivery practice in Google Cloud, AWS, or Azure and Terraform. Extensive experience implementing advanced observability using tools like Open Telemetry, Dynatrace, Grafana, and/or cloud-native services. Systematic problem-solving and troubleshooting skills in a complex system. Hands ...

Software Engineer II (AWS Infrastructure)

Hiring Organisation
Appcast
Location
Glasgow, UK
management processesContribute to the build and maintenance of CI/CD pipelines (e.g., Jenkins) that apply infrastructure changes safely and reliablyInstrument the estate for observability using tools such as Datadog, CloudWatch, and Dynatrace, and use telemetry insights to support improvements to infrastructure hygieneUses enterprise-authorized AI capabilities within the work … workloads and supporting cluster operationsPractical experience working with cloud infrastructure on AWS in a production or near-production environmentFamiliarity with production readiness practices including observability (metrics, tracing, logging) and incident management in distributed systemsWorking knowledge of using enterprise-authorized AI capabilities within the work environment to support software engineering workflows ...

DevOps Engineer

Location
Glasgow, Scotland, United Kingdom
innovative solutions that align with organizational goals and industry standards. Experience & Skills SME Advanced proficiency in Infrastructure Service Incident Management Intermediate proficiency in SRE Observability Intermediate proficiency in Configuration & Release Management Intermediate proficiency in Cloud Automation DevOps Knowledge of continuous integration, delivery and deployment. Knowledge of cloud technologies, container orchestration ...

Linux Engineer

Hiring Organisation
Searchability NS&D
Location
Glasgow, Scotland, United Kingdom
changes while improving resilience and reducing risk. Working within a highly skilled engineering team, you'll help strengthen pre and post-change validation, enhance observability, and support the continuous improvement of automation capabilities. Technology Stack Linux/UNIX Python Ansible Apache Airflow Prometheus Grafana Loki VMware F5 What ...

Application Engineer

Location
Glasgow, Scotland, United Kingdom
interfaces. Strong SQL skills and the ability to analyse and investigate data. Experience supporting batch processing and scheduling technologies. Familiarity with application monitoring and observability tools. Experience working with third-party vendors and technology suppliers. Strong communication skills with the ability to engage technical and business stakeholders. Proactive approach ...

Java Full stack Developer

Hiring Organisation
NEEV LIMITED
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Contract
Contract Rate
£400 per day
enterprise applications using modern Java technologies, cloud-native architectures, and microservices. The ideal candidate will have hands-on experience with Kafka, Kubernetes, API Security, Observability tools, SQL databases, and Spring-based microservices development . Key Responsibilities Design, develop, and maintain scalable Java-based applications using Java 17+, Spring Boot … mechanisms. Build event-driven solutions using Apache Kafka for real-time data processing and messaging. Deploy, manage, and troubleshoot applications on Kubernetes environments. Implement observability solutions using tools such as Splunk, ELK, Grafana, Prometheus, Dynatrace, or AppDynamics. Optimize application performance, scalability, and reliability. Work closely with business stakeholders, architects ...

Google Cloud Platform (GCP) Architect

Location
Glasgow, Scotland, United Kingdom
operational excellence by proactively identifying patterns in system failures, operational metrics, and data; designing and implementing systematic improvements to system reliability, performance, and observability Evaluate and lead vendor/technology assessments - leading sessions with external vendors, startups, and internal teams to drive outcomes-oriented evaluation of architectural designs, technical credentials … automation Frameworks on Kubernetes, including authoring reconciliation loops, admission controllers, webhooks, and custom controllers; strong understanding of Kubernetes internals, API machinery, RBAC, multi-tenancy, observability, and operational best practices for production environments. In-depth Google Cloud development experience -architecting, building, deploying, and operating production cloud-native workloads using core ...

Sr Lead AI Platform Engineer

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
responsibilitiesOwns the design and build of the team's platform: deployment pipelines, model serving, containerisation, orchestration, and environment managementSets the standard for reliability, observability, and operational excellence across the team's production AI/ML servicesBuilds the tooling and paved paths that let AI engineers ship agentic … track record of building deployment and release automationExperience serving, scaling, and monitoring ML models or data-intensive services in production (MLOps)Practical experience with observability tooling (metrics, logging, tracing) and production incident responseExperience operating ML/LLM workloads in production (LLMOps, inference reliability, cost/performance management)Strong communication skills ...

Sr Lead AI Platform Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
Owns the design and build of the team's platform: deployment pipelines, model serving, containerisation, orchestration, and environment management Sets the standard for reliability, observability, and operational excellence across the team's production AI/ML services Builds the tooling and paved paths that let AI engineers ship agentic … record of building deployment and release automation Experience serving, scaling, and monitoring ML models or data-intensive services in production (MLOps) Practical experience with observability tooling (metrics, logging, tracing) and production incident response Experience operating ML/LLM workloads in production (LLMOps, inference reliability, cost/performance management) Strong communication ...

Lead Site Reliability / DevOps Engineer

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
these practices within an application or platformFluency in at least one programming language such as (e.g., Java, Python, Go, etc.)Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency … postmortems for high-availability servicesDeep understanding of distributed system design principles, networking (TCP/IP, DNS, load balancing), Linux internals. Contributions to open-source observability or telemetry projects. Experience working with agent control planes and management protocols. Hands-on knowledge of OpAMP is highly desirable.J.P. Morgan is a global leader ...

GenAI & Agentic AI Lead Engineer

Location
Glasgow, Scotland, United Kingdom
GenAI apps using LLMs, RAG, and tool-using agents on AWS and cloud-native platforms. The role emphasizes modern AI engineering workflows, secure coding, observability, and collaboration with cross-functional teams to turn complex requirements into scalable AI systems. #J-18808-Ljbffr ...

Applied AI ML Lead - Python & Agentic AI

Location
Glasgow, Scotland, United Kingdom
SLMs, RAG, tool-using agents, evaluation, MLOps) and backend/service engineering (Java and/or Python, APIs/microservices, testing, CI/CD, observability, reliability) on AWS and cloud-native platforms. This role values modern AI engineering workflows and tooling such as GitHub Copilot and Claude Code to accelerate …/CD, deployment, monitoring, and maintenance for models/prompts/agents. Implement robust testing (unit/integration), performance benchmarking (latency/cost), and observability (logging/metrics/tracing) for AI services. Collaborate with cross-functional stakeholders to define requirements, success metrics, and rollout plans; communicate complex topics clearly ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
practices within an application or platform Fluency in at least one programming language such as (e.g., Java, Python, Go, etc.) Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency … high-availability services Deep understanding of distributed system design principles, networking (TCP/IP, DNS, load balancing), Linux internals. Contributions to open-source observability or telemetry projects. Experience working with agent control planes and management protocols. Hands-on knowledge of OpAMP is highly desirable. ABOUT US J.P. Morgan ...

Applied AI ML Lead - Python & Agentic AI

Hiring Organisation
Appcast
Location
Glasgow, UK
SLMs, RAG, tool-using agents, evaluation, MLOps) and backend/service engineering (Java and/or Python, APIs/microservices, testing, CI/CD, observability, reliability) on AWS and cloud-native platforms.This role values modern AI engineering workflows and tooling such as GitHub Copilot and Claude Code to accelerate delivery …/CD, deployment, monitoring, and maintenance for models/prompts/agents.Implement robust testing (unit/integration), performance benchmarking (latency/cost), and observability (logging/metrics/tracing) for AI services.Collaborate with cross-functional stakeholders to define requirements, success metrics, and rollout plans; communicate complex topics clearly to technical ...