126 to 150 of 185 Observability Jobs in Scotland

IT Service Manager

Location
Daliburgh, Scotland, United Kingdom
with ServiceNow or a comparable enterprise ITSM platform. Knowledge of operational resilience, business continuity, and disaster recovery frameworks. Experience implementing service automation, self-service, observability, or service experience improvements. Supplier management experience, including SLA monitoring, service reviews, and performance improvement initiatives. Previous people management, coaching, or team development experience. WHAT ...

Operations Team Lead (Production & Reliability)

Location
City of Edinburgh, Scotland, United Kingdom
Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under pressure, clear in communication Systems thinker, fixes root causes, not symptoms How We Think Production is sacred. Clear ownership beats ambiguity. ...

Principal Data Platform Architect

Location
City of Edinburgh, Scotland, United Kingdom
concepts Good knowledge of data governance, controls and architectural standards. Experience with cloud data platforms and technologies, including Snowflake. Strong understanding of data quality, observability, security and risk management. Good awareness of the wider technology landscape and vendors. Strong strategic thinking, problem-solving and commercial judgement. Confident leadership and influencing ...

Identity Security Engineering - SailPoint, Vice President

Hiring Organisation
Hackajob Ltd
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Work From Home
discovery, design, build, testing, parallel run, cutover and operational readiness, with clear ownership of risks,dependenciesand outcomes. Set high engineering standards for documentation, testing, observability, change control,resilienceand secure-by-design delivery. Coach engineers, share technicalknowledgeand create an environment where teams can make good decisions quickly and safely. Identifyopportunities ...

Applied AI ML Lead - Python & Agentic AI

Location
Auchentibber, Scotland, United Kingdom
SLMs, RAG, tool-using agents, evaluation, MLOps) and backend/service engineering (Java and/or Python, APIs/microservices, testing, CI/CD, observability, reliability) on AWS and cloud-native platforms. This role values modern AI engineering workflows and tooling such as GitHub Copilot and Claude Code to accelerate …/CD, deployment, monitoring, and maintenance for models/prompts/agents. Implement robust testing (unit/integration), performance benchmarking (latency/cost), and observability (logging/metrics/tracing) for AI services. Collaborate with cross-functional stakeholders to define requirements, success metrics, and rollout plans; communicate complex topics clearly ...

Senior Data Platform Engineer (Python & Databricks)

Location
City of Edinburgh, Scotland, United Kingdom
ingesting, transforming, and delivering financial data. Design data models that support enterprise reporting, analytics, and downstream integrations. Improve the performance, reliability, maintainability, and observability of existing data workflows. Troubleshoot complex issues across data-processing and application layers. Python and API Development Design, build, and maintain production-grade applications and services … services or other data-intensive industries. Experience with cloud-based data architectures. Familiarity with infrastructure-as-code tools such as Terraform. Experience improving the observability and operational reliability of data pipelines. Familiarity with modern data governance, access-control, and data-quality practices. Experience working in an enterprise environment with strict ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
practices within an application or platform Fluency in at least one programming language such as (e.g., Java, Python, Go, etc.) Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency … high-availability services Deep understanding of distributed system design principles, networking (TCP/IP, DNS, load balancing), Linux internals. Contributions to open-source observability or telemetry projects. Experience working with agent control planes and management protocols. Hands-on knowledge of OpAMP is highly desirable. ABOUT US J.P. Morgan ...

Applied AI ML Lead - Python & Agentic AI

Location
Glasgow, Scotland, United Kingdom
SLMs, RAG, tool-using agents, evaluation, MLOps) and backend/service engineering (Java and/or Python, APIs/microservices, testing, CI/CD, observability, reliability) on AWS and cloud-native platforms. This role values modern AI engineering workflows and tooling such as GitHub Copilot and Claude Code to accelerate …/CD, deployment, monitoring, and maintenance for models/prompts/agents. Implement robust testing (unit/integration), performance benchmarking (latency/cost), and observability (logging/metrics/tracing) for AI services. Collaborate with cross-functional stakeholders to define requirements, success metrics, and rollout plans; communicate complex topics clearly ...

Lead Cloud Engineer

Location
City of Edinburgh, Scotland, United Kingdom
data orchestration toolsets (e.g., dbt, Apache Airflow), ETL/ELT methodologies, real-time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost-control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated ...

Lead SRE: AWS Platform & Reliability Leader

Location
Glasgow, Scotland, United Kingdom
Co. in Glasgow seeks a Lead Site Reliability Engineer to define reliability strategy and drive robust, scalable platforms. You will lead incident response, shape observability, and guide AI-assisted reliability workflows across the SDLC. You will mentor peers, conduct resiliency reviews, and partner with product teams to establish SLOs, error ...

Senior Site Reliability Lead: Resiliency & Telemetry Architect

Location
Glasgow, Scotland, United Kingdom
Commercial & Investment Bank. You will guide resiliency design reviews, mentor engineers, and lead initiatives to improve reliability for large-scale systems. You will champion observability across OpenTelemetry pipelines, manage on-prem/cloud deployments, and collaborate with stakeholders to define SLOs, error budgets, and incident response playbooks. #J-18808-Ljbffr ...

Chief Engineer

Location
Edinburgh, Midlothian, United Kingdom
Chief Engineer and international SDA, you will act as the UK National Design Authority for GCAP Multi-function RF System. Lead the critical low observability technology development programme. Ensure that the jointly developed solution satisfies UK National objectives and aspirations. Lead system and hardware architecture, ensuring alignment across radar, multifunction ...

GCAP MRFS National Chief Engineer & Design Authority

Hiring Organisation
Leonardo DRS
Location
Edinburgh, UK
Employment Type
Full-time
Chief Engineer and international SDA, you will act as the UK National Design Authority for GCAP Multi-function RF System. Lead the critical low observability technology development programme. Ensure that the jointly developed solution satisfies UK National objectives and aspirations. Lead system and hardware architecture, ensuring alignment across radar, multifunction ...

Platform Engineer - Glasgow

Location
Glasgow, Scotland, United Kingdom
code generation, testing, documentation, and analysis, while understanding model limitations, protecting client data, and improving delivery quality and speed through pragmatic automation. SRE & Observability You’ll bring a reliability mindset to delivery, designing services that are operable by default and measured through meaningful SLIs/SLOs. You’ll help teams … implement pragmatic observability—logging, metrics, and distributed tracing—with actionable alerting, and you’ll contribute to (or lead) incident response and post‐incident reviews that drive learning and measurable improvements. We are looking for experience in the following skills: Strong experience with the AWS cloud platform and core services. Hands ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … Amazon EKS and Amazon SageMaker, as well as on-prem and local GPU clusters, using reproducible infrastructure as code and continuous delivery pipelines Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads Tune GPU and accelerator ...

Lead Software Engineer - LLM Ops Platform Reliability

Location
Paisley, Scotland, United Kingdom
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … Amazon EKS and Amazon SageMaker, as well as on-prem and local GPU clusters, using reproducible infrastructure as code and continuous delivery pipelines Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads Tune GPU and accelerator ...

Corporate KYC Principle Software Engineer - Executive Director

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
regulated financial services environments Establishes engineering standards for LLM-based applications RAG pipelines, embedding workflows, vector store integrations, and model serving ensuring safety, observability, and reproducibility at scale Drives adoption of advanced technical methods and practices aligned with the latest industry standards and product development methodologies Serves as the function … more disciplines (e.g., cloud, AI/ML, data engineering) Experience in large-scale data processing, microservices, API design, Kafka, Redis, MemCached, observability tools (Dynatrace, Splunk, Grafana), and orchestration frameworks (Airflow, Temporal) Advanced working knowledge of relational and NoSQL databases, vector stores, data lake architectures, and data governance Practical cloud-native ...

Platform Engineer - Edinburgh

Location
City of Edinburgh, Scotland, United Kingdom
code generation, testing, documentation, and analysis, while understanding model limitations, protecting client data, and improving delivery quality and speed through pragmatic automation SRE & Observability You’ll bring a reliability mindset to delivery, designing services that are operable by default and measured through meaningful SLIs/SLOs. You’ll help teams … implement pragmatic observability—logging, metrics, and distributed tracing—with actionable alerting, and you’ll contribute to (or lead) incident response and post-incident reviews that drive learning and measurable improvements. We are looking for experience in the following skills Strong experience with the AWS cloud platform and core services. Hands ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead SRE - AWS,Python

Location
Glasgow, Scotland, United Kingdom
delivery velocity Collaborate cross-functionally with software engineering, architecture, and security teams to embed reliability and resiliency principles early in the design process Champion observability practices by building and maintaining monitoring, alerting, and dashboarding solutions that provide actionable insights into system health Mentor and guide junior engineers, fostering a culture … automation and tooling development Experience defining and managing service level indicators, service level objectives, and error budgets in production environments Strong background in observability tooling, including metrics, logging, and distributed tracing platforms Demonstrated experience leading incident response processes, conducting blameless post-mortems, and driving systemic reliability improvements Experience with container ...

Senior Lead Software Data Engineer – Corporate Know Your Customer

Location
Auchentibber, Scotland, United Kingdom
scale data processing, microservices, API design, and orchestration frameworks Working knowledge of relational and NoSQL databases, vector stores, and data lake architectures Familiarity with observability tools and frameworks Practical cloud-native experience (AWS, Azure, or GCP) Ability to communicate effectively with senior leaders and executives Commitment to inclusive, collaborative teamwork … catalog services such as Apache Iceberg Experience with LLM orchestration frameworks and model serving infrastructure or managed endpoints Familiarity with AI evaluation and observability practices for LLM workloads Understanding of agentic design patterns and how to constrain agent autonomy in financial workflows Interest in emerging technologies and continuous learning Employer ...

Software Engineer III - Python

Location
Glasgow, Scotland, United Kingdom
infrastructure-as-code using Terraform within established team patterns across modules, environments, and state management Improve operability of services by adding and using observability tooling including logs, metrics, traces, dashboards, and alerts, and participate in incident response and root-cause analysis Leverage enterprise-authorized AI coding assist tools within … implementing application logic and APIs on top of relational data Experience building APIs and microservices using REST or gRPC, including contracts, security basics, and observability Practical experience delivering LLM-based features as part of software systems, with familiarity with agentic patterns Working knowledge of delivery and operations including CI/ ...

Corporate KYC Sr Lead Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
scale data processing, microservices, API design, and orchestration frameworks Working knowledge of relational and NoSQL databases, vector stores, and data lake architectures Familiarity with observability tools and frameworks Practical cloud-native experience (AWS, Azure, or GCP) Ability to communicate effectively with senior leaders and executives Commitment to inclusive, collaborative teamwork … catalog services such as Apache Iceberg Experience with LLM orchestration frameworks and model serving infrastructure or managed endpoints Familiarity with AI evaluation and observability practices for LLM workloads Understanding of agentic design patterns and how to constrain agent autonomy in financial workflows Interest in emerging technologies and continuous learning Employer ...

Software Engineer III - Full Stack, Global Banking Tech

Location
Glasgow, Scotland, United Kingdom
analysis, analyzing diverse datasets, logs, and telemetry to identify patterns, build visualizations and reporting, and reduce repeat incidents via preventative controls, automation, and enhanced observability Leverages enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity (e.g., code generation/refactoring, unit … agentic frameworks such as Google ADK or LangChain Familiarity with monitoring, tracing, and troubleshooting tools such as log aggregation platforms, API testing tools, and observability dashboards ABOUT US J.P. Morgan is a global leader in financial services, providing strategic advice and products to the world's most prominent corporations, governments ...

Lead Cloud Engineer

Location
City of Edinburgh, Scotland, United Kingdom
data orchestration toolsets (e.g., dbt, Apache Airflow), ETL/ELT methodologies, real‐time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost‐control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated ...