151 to 175 of 185 Observability Jobs in Scotland

Linux Automation Engineer — Hybrid (Glasgow)

Location
Paisley, Scotland, United Kingdom
London, to support cross-site initiatives. You will develop automation for patching, upgrades and changes across Linux, VMware and F5 environments, and contribute to observability, validation, and continuous improvement of the platform. #J-18808-Ljbffr ...

Senior Linux Automation Engineer (Hybrid | Travel Paid)

Location
Glasgow, Scotland, United Kingdom
reliable, infrastructure changes in Linux, VMware, and F5 environments. The role focuses on designing automation that accelerates patching and upgrades, strengthens change validation, improves observability, and supports continual automation improvements within a #J-18808-Ljbffr ...

ServiceNow AI & Enterprise Automation Lead - Managing Consultant

Location
Glasgow, Scotland, United Kingdom
value* Translate business requirements into AI-enabled workflow solutions**Solution Design & Architecture*** Design and support implementation of:* AI Control Tower (AI lifecycle management, governance, observability)* Agentic AI workflows enabling autonomous execution* Now Assist/GenAI use cases across workflows* Define data, integration, and workflow architectures for AI-enabled ServiceNow solutions ...

Associate/Vice President, AI Infrastructure Engineer

Hiring Organisation
Hackajob Ltd
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Work From Home
aligned with firmwide risk and compliance standards Implement and maintain infrastructure as code and automation to ensure repeatable, auditable platform provisioning Build and operate observability, monitoring, and alerting solutions for AI platforms, ensuring availability, performance, and cost transparency Collaborate with Security and Risk partners to integrate identity, access controls, data … cloud compute, networking, storage, and security services. Understanding of ML platform operations and governance concepts, including model deployment strategies, lifecycle management, monitoring/observability, and Disaster Recovery Experience supporting LLMs, generative AI platforms, or model serving infrastructure. Experience supporting AI and machine learning workloads, with exposure to managed compute ...

Context Plane Python Engineer

Location
Glasgow, Scotland, United Kingdom
data sources and services across the firm, including enterprise AI and large language model gateways Own quality across your components: automated testing, code reviews, observability, and resilient, secure service design Partner with Corporate Technology AI, product, and data science colleagues to translate concrete use cases into working, measurable capabilities Contribute … working with cloud infrastructure (AWS) and containerized services (Docker/ECS) Ability to own technical components end-to-end - from design through deployment and observability Strong collaboration skills with the ability to work across engineering, product, and data science disciplines Hands‐on experience using enterprise-authorized AI‐assisted software development ...

Senior Lead Software Data Engineer - Corporate Know Your Customer

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
large-scale data processing, microservices, API design, and orchestration frameworksWorking knowledge of relational and NoSQL databases, vector stores, and data lake architecturesFamiliarity with observability tools and frameworksPractical cloud-native experience (AWS, Azure, or GCP)Ability to communicate effectively with senior leaders and executivesCommitment to inclusive, collaborative teamworkStrong problem-solving … table formats and catalog services such as Apache IcebergExperience with LLM orchestration frameworks and model serving infrastructure or managed endpointsFamiliarity with AI evaluation and observability practices for LLM workloadsUnderstanding of agentic design patterns and how to constrain agent autonomy in financial workflowsInterest in emerging technologies and continuous learningEmployer DescriptionJPMorganChase ...

Platform Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
container orchestration, infrastructure as code and automation, you will help deliver secure, scalable and resilient services. You will also investigate complex technical issues, improve observability and reduce operational risk and manual effort. We are a multidisciplinary team looking for candidates with a broad mix of skills and experience. … server administration Cloud platforms, virtualisation and containers Infrastructure as code and configuration management Programming and scripting, such as Python, Bash or Go Monitoring, observability and SRE practices Infrastructure, networking and performance troubleshooting Secure, resilient and scalable system design Technical documentation and operational guidance Technical leadership and mentoring Beneficial skills include ...

Director, Artificial Intelligence & Machine Learning Engineer

Location
City of Edinburgh, Scotland, United Kingdom
deployment, monitoring, and adoption across the firm. Translate emerging AI capabilities into robust, scalable solutions that deliver measurable business impact. Define evaluation, testing, and observability approaches for AI systems, including non-deterministic and agentic applications. Optimize AI systems and workloads for latency, throughput, compute efficiency, reliability, security, and cost. Partner … fast-moving environments, ideating quickly on research prototypes. Generative AI, large language models, AI agents, and architectures for production AI systems. Evaluation, experimentation, observability, debugging, and testing approaches for AI/ML systems, including non-deterministic and agentic applications. Machine learning and optimization frameworks such as PyTorch, TensorFlow, or JAX. ...

Operations Team Lead: Reliability & Scale

Location
Aberdeen City, Scotland, United Kingdom
system that scales while ensuring reliability across live customer-facing systems. You will shape processes, lead incidents, and drive a culture focused on observability and proactive reliability engineering. You will define SLIs/SLOs, improve MTTR, and build a high-performing team focused on reducing outages and preventing recurrence. #J ...

Production Reliability Leader

Location
City of Edinburgh, Scotland, United Kingdom
team, and moving from reactive firefighting to proactive reliability engineering. This hands-on role focuses on monitoring, incident management, on-call rotations, and driving observability with SLIs/SLOs while maintaining a blameless culture and strong ownership. #J-18808-Ljbffr ...

Backend Engineer

Location
City of Edinburgh, Scotland, United Kingdom
using containers and modern cloud/platform technologies Implement and maintain CI/CD pipelines for automated testing, deployment and release management Establish strong observability, monitoring, logging and alerting capabilities Contribute to technical design reviews, engineering standards and best practices Work closely with AI Engineers, Data Scientists, Enterprise Architects, Security … OAuth, SAML and SSO API Gateways and enterprise integrations Cloud-native development CI/CD and DevOps practices Distributed systems and high-availability architectures Observability, monitoring and operational tooling Enterprise Integration Platforms AI/Agent Orchestration Platforms Why Join? This is an opportunity to work on a next-generation enterprise ...

Director of Data Engineering

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
appropriate for regulated financial services environmentsEstablishes engineering standards for LLM-based applications — RAG pipelines, embedding workflows, vector store integrations, and model serving — ensuring safety, observability, and reproducibility at scaleDrives adoption of advanced technical methods and practices aligned with the latest industry standards and product development methodologiesServes as the function … more disciplines (e.g., cloud, AI/ML, data engineering)Experience in large-scale data processing, microservices, API design, Kafka, Redis, MemCached, observability tools (Dynatrace, Splunk, Grafana), and orchestration frameworks (Airflow, Temporal)Advanced working knowledge of relational and NoSQL databases, vector stores, data lake architectures, and data governancePractical cloud-native experience ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelines Implement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloads Tune GPU and accelerator capacity, autoscaling, and cost ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … proprietary large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelinesImplement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloadsTune GPU and accelerator capacity, autoscaling, and cost efficiency ...

AI Platform Engineer

Location
Stirling, Scotland, United Kingdom
across the business adopt AI solutions safely and effectively. This is a hands-on engineering role that combines feature delivery with responsibility for resilience, observability, governance and measurable outcomes. You’ll work closely with Product Owners, Architects, Engineers, Security teams and business stakeholders to deliver secure, scalable and reliable platform … applications, APIs, services or platforms. Experience using logs, metrics, traces and telemetry to improve platform performance and reliability. Experience designing solutions with resilience, scalability, observability and operability in mind. Evidence of improving the dependability of production systems through practical engineering changes. A track record of delivering technology capabilities into production ...

Lead Infrastructure Engineer - AWS Cloud Support Engineering

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
adherence to resiliency and security expectations Familiarity with working in a large distributed system across a range of technologies including compute, databases, messaging, observability, and telemetry Knowledge of incident, change, and problem management processes and the controls that govern them Understanding of data-driven decision making and a drive … working in a follow-the-sun or globally distributed on-call support model Familiarity with large-scale cloud migration or modernization initiatives Exposure to observability and telemetry tooling in complex distributed environments ABOUT US J.P. Morgan is a global leader in financial services, providing strategic advice and products ...

Lead Software Engineer (Container Platforms)

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
usability, and self-service for engineering consumers Create and maintain delivery workflows that standardize engineering practices and reduce operational toil across the platform Improve observability and operational readiness through monitoring, logging, tracing, alerting, runbook development, and on-call practices Partner with security and risk stakeholders to implement secure-by-default … Familiarity with infrastructure-as-code and automation practices, including tools such as Terraform, Helm, Kustomize, Argo CD, Flux, or continuous integration systems Experience with observability stacks and site reliability engineering practices, including service level indicators and objectives, incident response, and post-incident reviews Exposure to regulated environments and implementing security ...

Platform Engineer

Location
City of Edinburgh, Scotland, United Kingdom
container orchestration, infrastructure as code and automation, you will help deliver secure, scalable and resilient services. You will also investigate complex technical issues, improve observability and reduce operational risk and manual effort.We are a multidisciplinary team looking for candidates with a broad mix of skills and experience. … server administration* Cloud platforms, virtualisation and containers* Infrastructure as code and configuration management* Programming and scripting, such as Python, Bash or Go* Monitoring, observability and SRE practices* Infrastructure, networking and performance troubleshooting* Secure, resilient and scalable system design* Technical documentation and operational guidance* Technical leadership and mentoringBeneficial skills include:* Strong ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Location
Auchentibber, Scotland, United Kingdom
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You’ll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you’ll find meaningful impact and growth here. … large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelines Implement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloads Tune GPU and accelerator capacity, autoscaling, and cost ...

Lead SRE - AWS Platform

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
your team to identify comprehensive service level indicators and partner with stakeholders to establish reasonable service level objectives and error budgets Design and implement observability frameworks and alerting strategies, including white and black box monitoring, service level objective-based alerting, and telemetry collection to ensure proactive detection and response Serve … resiliency best practices Fluency in at least one programming language such as Python, Java/Spring Boot, or .NET Proficient knowledge and experience in observability, including white and black box monitoring, service level objective alerting, and telemetry collection across large-scale production environments Proficiency with continuous integration and continuous delivery ...

Lead Infrastructure Engineer - AWS Cloud Support Engineering

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
outputs, and adherence to resiliency and security expectationsFamiliarity with working in a large distributed system across a range of technologies including compute, databases, messaging, observability, and telemetryKnowledge of incident, change, and problem management processes and the controls that govern themUnderstanding of data-driven decision making and a drive to continue … Terraform certificationsExperience working in a follow-the-sun or globally distributed on-call support modelFamiliarity with large-scale cloud migration or modernization initiativesExposure to observability and telemetry tooling in complex distributed environmentsJ.P. Morgan is a global leader in financial services, providing strategic advice and products to the world's most ...

Data Engineer IV

Location
Glasgow, Scotland, United Kingdom
About Planet DDS We’re on a mission to fix dental software - and we’re not playing small. Our platform replaces clunky, outdated systems with modern, cloud-based, AI-powered technology built to actually work ...

AWS SRE Devops engineer

Location
Glasgow, Scotland, United Kingdom
planning, failure testing, and resilience validation. Define and manage service health metrics (SLIs/SLOs/SLAs) to drive measurable improvements in reliability. Build observability solutions to monitor AWS, Snowflake, and Databricks workloads. Collaborate with engineering teams to embed reliability best practices throughout platform development. Analyse incidents and proactively address … principles and practical experience defining SLAs, SLOs, and error budgets. Demonstrated AWS expertise (e.g., EC2, S3, IAM, VPC, CloudWatch) in production environments. Experience with observability tools, monitoring, and alerting practices. Proficient in automation, Infrastructure as Code (Terraform, CloudFormation, or CDK), and scripting (Python/Bash). Exposure to Snowflake ...

Lead SRE - AWS Platform

Location
Glasgow, Scotland, United Kingdom
your team to identify comprehensive service level indicators and partner with stakeholders to establish reasonable service level objectives and error budgets Design and implement observability frameworks and alerting strategies, including white and black box monitoring, service level objective-based alerting, and telemetry collection to ensure proactive detection and response Serve … resiliency best practices Fluency in at least one programming language such as Python, Java/Spring Boot, or .NET Proficient knowledge and experience in observability, including white and black box monitoring, service level objective alerting, and telemetry collection across large-scale production environments Proficiency with continuous integration and continuous delivery ...

Data Architect

Location
City of Edinburgh, Scotland, United Kingdom
platforms and delivery teams. The role translates enterprise and product architecture principles into concrete solution designs covering data ingestion, processing, storage, APIs, events, security, observability and operational controls. Working closely with platform engineers, software engineers, data engineers, modellers, analysts and governance teams, the Data Architect ensures that data solutions … robust. Security and Privacy by Design - Embed appropriate access controls, encryption, masking, segregation, retention, monitoring and audit requirements into technical designs. Data Quality and Observability - Define technical implementation patterns for validation, reconciliation, lineage capture, monitoring, alerting and incident investigation. Performance and Scalability - Assess and optimise technical designs for latency, throughput ...