76 to 100 of 100 Observability Jobs in Glasgow

Lead SRE: AWS Platform & Reliability Leader

Location
Glasgow, Scotland, United Kingdom
Co. in Glasgow seeks a Lead Site Reliability Engineer to define reliability strategy and drive robust, scalable platforms. You will lead incident response, shape observability, and guide AI-assisted reliability workflows across the SDLC. You will mentor peers, conduct resiliency reviews, and partner with product teams to establish SLOs, error ...

Platform Engineer - Glasgow

Location
Glasgow, Scotland, United Kingdom
code generation, testing, documentation, and analysis, while understanding model limitations, protecting client data, and improving delivery quality and speed through pragmatic automation. SRE & Observability You’ll bring a reliability mindset to delivery, designing services that are operable by default and measured through meaningful SLIs/SLOs. You’ll help teams … implement pragmatic observability—logging, metrics, and distributed tracing—with actionable alerting, and you’ll contribute to (or lead) incident response and post‐incident reviews that drive learning and measurable improvements. We are looking for experience in the following skills: Strong experience with the AWS cloud platform and core services. Hands ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … Amazon EKS and Amazon SageMaker, as well as on-prem and local GPU clusters, using reproducible infrastructure as code and continuous delivery pipelines Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads Tune GPU and accelerator ...

Corporate KYC Principle Software Engineer - Executive Director

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
regulated financial services environments Establishes engineering standards for LLM-based applications RAG pipelines, embedding workflows, vector store integrations, and model serving ensuring safety, observability, and reproducibility at scale Drives adoption of advanced technical methods and practices aligned with the latest industry standards and product development methodologies Serves as the function … more disciplines (e.g., cloud, AI/ML, data engineering) Experience in large-scale data processing, microservices, API design, Kafka, Redis, MemCached, observability tools (Dynatrace, Splunk, Grafana), and orchestration frameworks (Airflow, Temporal) Advanced working knowledge of relational and NoSQL databases, vector stores, data lake architectures, and data governance Practical cloud-native ...

Lead SRE - AWS,Python

Location
Glasgow, Scotland, United Kingdom
delivery velocity Collaborate cross-functionally with software engineering, architecture, and security teams to embed reliability and resiliency principles early in the design process Champion observability practices by building and maintaining monitoring, alerting, and dashboarding solutions that provide actionable insights into system health Mentor and guide junior engineers, fostering a culture … automation and tooling development Experience defining and managing service level indicators, service level objectives, and error budgets in production environments Strong background in observability tooling, including metrics, logging, and distributed tracing platforms Demonstrated experience leading incident response processes, conducting blameless post-mortems, and driving systemic reliability improvements Experience with container ...

Software Engineer III - Python

Location
Glasgow, Scotland, United Kingdom
infrastructure-as-code using Terraform within established team patterns across modules, environments, and state management Improve operability of services by adding and using observability tooling including logs, metrics, traces, dashboards, and alerts, and participate in incident response and root-cause analysis Leverage enterprise-authorized AI coding assist tools within … implementing application logic and APIs on top of relational data Experience building APIs and microservices using REST or gRPC, including contracts, security basics, and observability Practical experience delivering LLM-based features as part of software systems, with familiarity with agentic patterns Working knowledge of delivery and operations including CI/ ...

Corporate KYC Sr Lead Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
scale data processing, microservices, API design, and orchestration frameworks Working knowledge of relational and NoSQL databases, vector stores, and data lake architectures Familiarity with observability tools and frameworks Practical cloud-native experience (AWS, Azure, or GCP) Ability to communicate effectively with senior leaders and executives Commitment to inclusive, collaborative teamwork … catalog services such as Apache Iceberg Experience with LLM orchestration frameworks and model serving infrastructure or managed endpoints Familiarity with AI evaluation and observability practices for LLM workloads Understanding of agentic design patterns and how to constrain agent autonomy in financial workflows Interest in emerging technologies and continuous learning Employer ...

Lead Software Engineer - LLM Ops Platform Reliability

Location
Glasgow, Scotland, United Kingdom
Join a team where your engineering expertise directly protects the reliability and resilience of systems that matter at scale. At JPMorganChase, we invest in engineers who think beyond the code — who own outcomes, drive operational ...

Lead Site Reliability Engineer - Operations Excellence

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
Join a team where your engineering expertise directly protects the reliability and resilience of systems that matter at scale. At JPMorganChase, we invest in engineers who think beyond the code — who own outcomes, drive operational ...

ServiceNow AI & Enterprise Automation Lead - Managing Consultant

Location
Glasgow, Scotland, United Kingdom
value* Translate business requirements into AI-enabled workflow solutions**Solution Design & Architecture*** Design and support implementation of:* AI Control Tower (AI lifecycle management, governance, observability)* Agentic AI workflows enabling autonomous execution* Now Assist/GenAI use cases across workflows* Define data, integration, and workflow architectures for AI-enabled ServiceNow solutions ...

Software Engineer III - Full Stack, Global Banking Tech

Location
Glasgow, Scotland, United Kingdom
analysis, analyzing diverse datasets, logs, and telemetry to identify patterns, build visualizations and reporting, and reduce repeat incidents via preventative controls, automation, and enhanced observability Leverages enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity (e.g., code generation/refactoring, unit … agentic frameworks such as Google ADK or LangChain Familiarity with monitoring, tracing, and troubleshooting tools such as log aggregation platforms, API testing tools, and observability dashboards #J-18808-Ljbffr ...

Context Plane Python Engineer

Location
Glasgow, Scotland, United Kingdom
data sources and services across the firm, including enterprise AI and large language model gateways Own quality across your components: automated testing, code reviews, observability, and resilient, secure service design Partner with Corporate Technology AI, product, and data science colleagues to translate concrete use cases into working, measurable capabilities Contribute … working with cloud infrastructure (AWS) and containerized services (Docker/ECS) Ability to own technical components end-to-end - from design through deployment and observability Strong collaboration skills with the ability to work across engineering, product, and data science disciplines Hands‐on experience using enterprise-authorized AI‐assisted software development ...

Senior Lead Software Data Engineer - Corporate Know Your Customer

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
large-scale data processing, microservices, API design, and orchestration frameworksWorking knowledge of relational and NoSQL databases, vector stores, and data lake architecturesFamiliarity with observability tools and frameworksPractical cloud-native experience (AWS, Azure, or GCP)Ability to communicate effectively with senior leaders and executivesCommitment to inclusive, collaborative teamworkStrong problem-solving … table formats and catalog services such as Apache IcebergExperience with LLM orchestration frameworks and model serving infrastructure or managed endpointsFamiliarity with AI evaluation and observability practices for LLM workloadsUnderstanding of agentic design patterns and how to constrain agent autonomy in financial workflowsInterest in emerging technologies and continuous learningEmployer DescriptionJPMorganChase ...

Director of Data Engineering

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
appropriate for regulated financial services environmentsEstablishes engineering standards for LLM-based applications — RAG pipelines, embedding workflows, vector store integrations, and model serving — ensuring safety, observability, and reproducibility at scaleDrives adoption of advanced technical methods and practices aligned with the latest industry standards and product development methodologiesServes as the function … more disciplines (e.g., cloud, AI/ML, data engineering)Experience in large-scale data processing, microservices, API design, Kafka, Redis, MemCached, observability tools (Dynatrace, Splunk, Grafana), and orchestration frameworks (Airflow, Temporal)Advanced working knowledge of relational and NoSQL databases, vector stores, data lake architectures, and data governancePractical cloud-native experience ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelines Implement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloads Tune GPU and accelerator capacity, autoscaling, and cost ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … proprietary large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelinesImplement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloadsTune GPU and accelerator capacity, autoscaling, and cost efficiency ...

Lead Infrastructure Engineer - AWS Cloud Support Engineering

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
adherence to resiliency and security expectations Familiarity with working in a large distributed system across a range of technologies including compute, databases, messaging, observability, and telemetry Knowledge of incident, change, and problem management processes and the controls that govern them Understanding of data-driven decision making and a drive … working in a follow-the-sun or globally distributed on-call support model Familiarity with large-scale cloud migration or modernization initiatives Exposure to observability and telemetry tooling in complex distributed environments ABOUT US J.P. Morgan is a global leader in financial services, providing strategic advice and products ...

Senior Software Engineer

Location
Glasgow, Scotland, United Kingdom
"It feels good to have a career with real purpose." Job Summary Royal London is seeking an experienced Senior Software Engineer to join one of our application delivery teams within the Digital space at Royal ...

Lead SRE - AWS Platform

Hiring Organisation
Appcast
Location
Glasgow, UK
with your team to identify comprehensive service level indicators and partner with stakeholders to establish reasonable service level objectives and error budgetsDesign and implement observability frameworks and alerting strategies, including white and black box monitoring, service level objective-based alerting, and telemetry collection to ensure proactive detection and responseServe … implementing resiliency best practicesFluency in at least one programming language such as Python, Java/Spring Boot, or .NETProficient knowledge and experience in observability, including white and black box monitoring, service level objective alerting, and telemetry collection across large-scale production environmentsProficiency with continuous integration and continuous delivery practices ...

Lead SRE - AWS Platform

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
your team to identify comprehensive service level indicators and partner with stakeholders to establish reasonable service level objectives and error budgets Design and implement observability frameworks and alerting strategies, including white and black box monitoring, service level objective-based alerting, and telemetry collection to ensure proactive detection and response Serve … resiliency best practices Fluency in at least one programming language such as Python, Java/Spring Boot, or .NET Proficient knowledge and experience in observability, including white and black box monitoring, service level objective alerting, and telemetry collection across large-scale production environments Proficiency with continuous integration and continuous delivery ...

Lead Infrastructure Engineer - AWS Cloud Support Engineering

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
outputs, and adherence to resiliency and security expectationsFamiliarity with working in a large distributed system across a range of technologies including compute, databases, messaging, observability, and telemetryKnowledge of incident, change, and problem management processes and the controls that govern themUnderstanding of data-driven decision making and a drive to continue … Terraform certificationsExperience working in a follow-the-sun or globally distributed on-call support modelFamiliarity with large-scale cloud migration or modernization initiativesExposure to observability and telemetry tooling in complex distributed environmentsJ.P. Morgan is a global leader in financial services, providing strategic advice and products to the world's most ...

Lead SRE - AWS Platform

Location
Glasgow, Scotland, United Kingdom
your team to identify comprehensive service level indicators and partner with stakeholders to establish reasonable service level objectives and error budgets Design and implement observability frameworks and alerting strategies, including white and black box monitoring, service level objective-based alerting, and telemetry collection to ensure proactive detection and response Serve … resiliency best practices Fluency in at least one programming language such as Python, Java/Spring Boot, or .NET Proficient knowledge and experience in observability, including white and black box monitoring, service level objective alerting, and telemetry collection across large-scale production environments Proficiency with continuous integration and continuous delivery ...

Director of Software Engineering

Location
Glasgow, Scotland, United Kingdom
hands-on: review code, prototype solutions, and get into the details when it matters Establish engineering standards across code quality, system design, testing, and observability, and hold the team to them Be the person engineers come to when the problem is genuinely hard Team Building & Culture Recruit, develop, andretaina team … Experience writing performance software in Rust Background in space systems, aerospace, or highly constrained real-time environments Experience building data lakes, telemetry platforms, or observability infrastructure at scale A history of leading teams through technical transformations and not justmaintainingthe status quo Spire operates a hybrid work model, and this position ...

Director of Software Engineering

Hiring Organisation
Spire Global
Location
Glasgow, UK
Employment Type
Full-time
infrastructureStay hands-on: review code, prototype solutions, and get into the details when it mattersEstablish engineering standards across code quality, system design, testing, and observability, and hold the team to themBe the person engineers come to when the problem is genuinely hardTeam Building & Culture Recruit, develop, and retain a team … monitoring problemsExperience writing performance software in RustBackground in space systems, aerospace, or highly constrained real-time environmentsExperience building data lakes, telemetry platforms, or observability infrastructure at scaleA history of leading teams through technical transformations and not just maintaining the status quoSpire operates a hybrid work model, and this position will ...

Senior Lead SRE - Reliability & Observability Leader

Location
Glasgow, Scotland, United Kingdom
JPMorgan Chase in the United Kingdom is seeking a Senior Lead Site Reliability Engineer to join an agile team focused on reliability, observability, and performance across critical platforms. You will mentor engineers, lead incident response, and shape SRE strategy while delivering scalable, secure production systems. The role demands deep expertise ...