101 to 125 of 185 Observability Jobs in Scotland

Sr Lead Infrastructure Engineer- Devops/AWS

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
machine learning products. As the team takes end-to-end ownership of the platforms it runs, you will build the CI/CD, observability, and incident-management practices that keep those services stable, secure, and performant across international markets. This is a Vice President-level role and an integral part … pipelines, release automation, and deployment toolingEstablishes reliability practices (SLOs, error budgets, runbooks) and leads production incident response and post-incident reviewBuilds and operates observability across the team's AI/ML services (metrics, logging, tracing, alerting)Automates infrastructure provisioning and configuration through infrastructure-as-codeImplements operational security, secrets management ...

Sr Lead Infrastructure Engineer- Devops/AWS

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
machine learning products. As the team takes end-to-end ownership of the platforms it runs, you will build the CI/CD, observability, and incident-management practices that keep those services stable, secure, and performant across international markets. This is a Vice President-level role and an integral part … release automation, and deployment tooling Establishes reliability practices (SLOs, error budgets, runbooks) and leads production incident response and post-incident review Builds and operates observability across the team's AI/ML services (metrics, logging, tracing, alerting) Automates infrastructure provisioning and configuration through infrastructure-as-code Implements operational security, secrets ...

Senior Backend Engineer

Hiring Organisation
Inspire People
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
part of a multidisciplinary agile team, you will design, build and run platform services that underpin critical digital products, helping development teams improve observability, monitoring, CI/CD and service resilience. As a Senior Backend Engineer (Site Reliability), you will: * Design, build and operate reliable, secure and scalable cloud platform … services supporting critical digital products. * Develop and maintain platform tooling, automation, observability, monitoring and CI/CD capabilities. * Build software solutions using Python and modern engineering practices. * Write clean, maintainable code and infrastructure-as-code solutions to support service delivery. * Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets ...

Software Engineer II (AWS Infrastructure)

Location
Auchentibber, Scotland, United Kingdom
Contribute to the build and maintenance of CI/CD pipelines (e.g., Jenkins) that apply infrastructure changes safely and reliably Instrument the estate for observability using tools such as Datadog, CloudWatch, and Dynatrace, and use telemetry insights to support improvements to infrastructure hygiene Uses enterprise-authorized AI capabilities within … supporting cluster operations Practical experience working with cloud infrastructure on AWS in a production or near-production environment Familiarity with production readiness practices including observability (metrics, tracing, logging) and incident management in distributed systems Working knowledge of using enterprise-authorized AI capabilities within the work environment to support software engineering ...

Software Engineer III - AI/ML Platform Reliability

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
enhance the reliability and scalability of AI/ML platforms and applications to accommodate fast-growing demands. Own NFRs and develop tooling for observability, security, resilience, infrastructure management and operations excellence. Build and maintain scalable infrastructure to support the deployment and operation of large-scale AI platforms and apps. Build … architecture. Experience building large scale infrastructure and and cloud-native delivery practice in Google Cloud, AWS, or Azure and Terraform. Extensive experience implementing advanced observability using tools like Open Telemetry, Dynatrace, Grafana, and/or cloud-native services. Systematic problem-solving and troubleshooting skills in a complex system. Hands ...

AWS SRE DevOps Engineer: Cloud Reliability & Automation

Location
Glasgow, Scotland, United Kingdom
automation, resilience, and operational excellence across critical data services. You will automate infrastructure, drive DR planning, define SLIs/SLOs/SLAs, and implement observability to improve reliability and performance across services. #J-18808-Ljbffr ...

Senior SRE Lead for AI/ML Data Platform

Location
Cumbernauld, Scotland, United Kingdom
engineers, implement best-practice SRE processes, and collaborate with cross-functional teams to deliver secure, scalable solutions. The role emphasizes automation, observability, and incident response. #J-18808-Ljbffr ...

Linux Engineer

Hiring Organisation
Searchability NS&D
Location
Glasgow, Scotland, United Kingdom
changes while improving resilience and reducing risk. Working within a highly skilled engineering team, you'll help strengthen pre and post-change validation, enhance observability, and support the continuous improvement of automation capabilities. Technology Stack Linux/UNIX Python Ansible Apache Airflow Prometheus Grafana Loki VMware F5 What ...

Application Engineer

Location
Glasgow, Scotland, United Kingdom
interfaces. Strong SQL skills and the ability to analyse and investigate data. Experience supporting batch processing and scheduling technologies. Familiarity with application monitoring and observability tools. Experience working with third-party vendors and technology suppliers. Strong communication skills with the ability to engage technical and business stakeholders. Proactive approach ...

Senior Software Engineer C Linux Networking - Remote

Hiring Organisation
Saxon Recruitment Solutions
Location
Edinburgh, UK
Employment Type
Full-time
router leveraging a software dataplane. It provides a broad set of routing protocols and other network features, as well as cutting edge configuration and observability capabilities. What you'll needTechnical leadership for a significant development or project, with some demonstrable experience operating as a Subject Matter ExpertDemonstrable experience designing ...

TSB | Senior Operations Analyst | Grade D | CIO | Edinburgh | 12 Month Fixed Term Contract

Location
City of Edinburgh, Scotland, United Kingdom
role involves monitoring the bank's core business services through industry leading toolsets and the ongoing development of these to ensure that service observability is maximised and continually improved. Inquisitive, problem-solving skills are required as the role involves the triage and reaction to anomalies within our technology to ensure ...

Support Engineer

Location
City of Edinburgh, Scotland, United Kingdom
create automated tests and submit code fixes for engineering review. Experience supporting AI-enabled or data-intensive applications. Familiarity with incident-management and observability tools. How we work We’re an in-person team based in Edinburgh. You should expect to work from the office as your default. We support ...

Java Full stack Developer

Hiring Organisation
NEEV LIMITED
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Contract
Contract Rate
£400 per day
enterprise applications using modern Java technologies, cloud-native architectures, and microservices. The ideal candidate will have hands-on experience with Kafka, Kubernetes, API Security, Observability tools, SQL databases, and Spring-based microservices development . Key Responsibilities Design, develop, and maintain scalable Java-based applications using Java 17+, Spring Boot … mechanisms. Build event-driven solutions using Apache Kafka for real-time data processing and messaging. Deploy, manage, and troubleshoot applications on Kubernetes environments. Implement observability solutions using tools such as Splunk, ELK, Grafana, Prometheus, Dynatrace, or AppDynamics. Optimize application performance, scalability, and reliability. Work closely with business stakeholders, architects ...

Site Reliability Engineer (SRE) - Cloud Kubernetes Platform

Hiring Organisation
Intuition IT Solutions Ltd
Location
Glasgow, United Kingdom
Employment Type
Contract
Contract Rate
GBP Annual
experience managing Kubernetes environments on public cloud platforms such as AKS, EKS, or GKE, along with solid knowledge of SRE, DevOps, infrastructure automation, and observability practices. You will be responsible for maintaining the availability, reliability, scalability, and performance of cloud-native infrastructure and CI/CD platforms. Key Responsibilities Manage … Cause Analysis (RCA) . Implement preventive and corrective measures to improve platform reliability. Automate operational activities using Python or other Scripting languages. Improve platform observability, monitoring, alerting, and performance. Support incident management and on-call rotations . Contribute to GitOps-based deployment practices using tools such as Flux . Required ...

Sr Lead AI Platform Engineer

Location
Auchentibber, Scotland, United Kingdom
Owns the design and build of the team's platform: deployment pipelines, model serving, containerisation, orchestration, and environment management Sets the standard for reliability, observability, and operational excellence across the team's production AI/ML services Builds the tooling and paved paths that let AI engineers ship agentic … record of building deployment and release automation Experience serving, scaling, and monitoring ML models or data-intensive services in production (MLOps) Practical experience with observability tooling (metrics, logging, tracing) and production incident response Experience operating ML/LLM workloads in production (LLMOps, inference reliability, cost/performance management) Strong communication ...

Lead Site Reliability / DevOps Engineer

Location
Auchentibber, Scotland, United Kingdom
practices within an application or platform Fluency in at least one programming language such as (e.g., Java, Python, Go, etc.) Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency … high-availability services Deep understanding of distributed system design principles, networking (TCP/IP, DNS, load balancing), Linux internals. Contributions to open-source observability or telemetry projects. Experience working with agent control planes and management protocols. Hands-on knowledge of OpAMP is highly desirable. #J-18808-Ljbffr ...

Software Engineer III - AI/ML Platform Reliability

Location
Auchentibber, Scotland, United Kingdom
enhance the reliability and scalability of AI/ML platforms and applications to accommodate fast-growing demands. Own NFRs and develop tooling for observability, security, resilience, infrastructure management and operations excellence. Build and maintain scalable infrastructure to support the deployment and operation of large-scale AI platforms and apps. Build … architecture. Experience building large scale infrastructure and and cloud-native delivery practice in Google Cloud, AWS, or Azure and Terraform. Extensive experience implementing advanced observability using tools like Open Telemetry, Dynatrace, Grafana, and/or cloud-native services. Systematic problem-solving and troubleshooting skills in a complex system. Hands ...

Google Cloud Platform (GCP) Architect

Location
Glasgow, Scotland, United Kingdom
operational excellence by proactively identifying patterns in system failures, operational metrics, and data; designing and implementing systematic improvements to system reliability, performance, and observability Evaluate and lead vendor/technology assessments - leading sessions with external vendors, startups, and internal teams to drive outcomes-oriented evaluation of architectural designs, technical credentials … automation Frameworks on Kubernetes, including authoring reconciliation loops, admission controllers, webhooks, and custom controllers; strong understanding of Kubernetes internals, API machinery, RBAC, multi-tenancy, observability, and operational best practices for production environments. In-depth Google Cloud development experience -architecting, building, deploying, and operating production cloud-native workloads using core ...

Sr Lead AI Platform Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
Owns the design and build of the team's platform: deployment pipelines, model serving, containerisation, orchestration, and environment management Sets the standard for reliability, observability, and operational excellence across the team's production AI/ML services Builds the tooling and paved paths that let AI engineers ship agentic … record of building deployment and release automation Experience serving, scaling, and monitoring ML models or data-intensive services in production (MLOps) Practical experience with observability tooling (metrics, logging, tracing) and production incident response Experience operating ML/LLM workloads in production (LLMOps, inference reliability, cost/performance management) Strong communication ...

Lead Site Reliability / DevOps Engineer

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
these practices within an application or platformFluency in at least one programming language such as (e.g., Java, Python, Go, etc.)Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency … postmortems for high-availability servicesDeep understanding of distributed system design principles, networking (TCP/IP, DNS, load balancing), Linux internals. Contributions to open-source observability or telemetry projects. Experience working with agent control planes and management protocols. Hands-on knowledge of OpAMP is highly desirable.J.P. Morgan is a global leader ...

Hybrid Linux Automation Engineer – Travel Expensed

Location
Milton, Scotland, United Kingdom
automation across large-scale enterprise infrastructure, including Linux, VMware, and F5 environments. You will build automation to accelerate patching and changes, strengthen validation, improve observability with Prometheus, Grafana, and Airflow, and work with Python and Ansible within a collaborative engineering team. #J-18808-Ljbffr ...

Senior AI Engineer: Lead Multi-Agent Systems & Production

Location
Glasgow, Scotland, United Kingdom
GlobalLogic is seeking a Senior/Lead AI Engineer to own and drive core agentic frameworks, evaluation pipelines, and observability tooling for safe, trustworthy, scalable AI solutions in enterprise contexts. You will provide technical leadership and architectural direction for AI engineering, designing high-performance multi-agent systems and mentoring engineers ...

Senior AI Infra Engineer – LLM Ops & Reliability

Location
Auchentibber, Scotland, United Kingdom
serving stacks, optimize performance and costs, and lead incident response across cloud and on-prem GPU clusters. The role emphasizes secure software delivery, observability, and strong engineering fundamentals. You will collaborate with engineering to deliver scalable AI platforms, drive reliability, and build reusable patterns for robust production deployments, with ...

Senior AI/ML Platform Reliability Engineer

Location
Auchentibber, Scotland, United Kingdom
platforms, with a focus on reliability and security. As part of the Reliability Engineering team, you will own non-functional requirements, build tooling for observability and resilience, and partner across teams to unblock high-impact AI use cases. This role emphasizes production-grade code, on-call readiness, and shaping ...

Linux Engineer

Location
Glasgow, Scotland, United Kingdom
changes while improving resilience and reducing risk. Working within a highly skilled engineering team, you'll help strengthen pre and post-change validation, enhance observability, and support the continuous improvement of automation capabilities. Python Ansible Grafana Loki VMware F5 What We're Looking For We're looking for engineers with ...