1 to 25 of 33 Observability Jobs in Scotland

Lead SRE - Azure and GCP

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
with good understanding of REST APIs Hands-on experience with cloud-based technologies and tools especially in deployment, monitoring and operations, such as Google Observability, Azure Monitor, Data Dog, Prometheus, Splunk, Elasticsearch and Grafana. Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g. ...

Senior Manager of SRE

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
drive outcomes-oriented probing of architectural designs, technical credentials, and applicability for use within existing systems and information architecture. Drives continuous improvement in system observability, alerting, and capacity planning. Collaborates with engineering and data teams to optimize infrastructure and deployment processes, focusing on automation and operational excellence. Performs platform design ...

Java Software Engineer - VP

Hiring Organisation
Henderson Scott
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
Data: Proficiency with Docker, Kubernetes, and relational databases (MS SQL, Sybase). Delivery Automation & Monitoring: Strong focus on automated testing, automated release pipelines, and observability tools (Grafana, Prometheus). Mindset: Delivery-focused problem solver with a hands-on approach and strong stakeholder communication skills. Desirable Experience Background in Equity Swaps ...

Software Engineer III - Data Engineering - Corporate Know Your Customer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
depth in disciplines such as cloud, AI/ML, or data engineering Experience in large-scale data processing, microservices, API design, Kafka, Redis, MemCached, observability tools (Dynatrace, Splunk, Grafana), and orchestration frameworks (Airflow, Temporal) Advanced working knowledge of relational and NoSQL databases, vector stores, data lake architectures, and data governance ...

Lead Software Data Engineer - Corporate Know Your Customer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
depth in disciplines such as cloud, AI/ML, or data engineering Experience in large-scale data processing, microservices, API design, Kafka, Redis, MemCached, observability tools (Dynatrace, Splunk, Grafana), and orchestration frameworks (Airflow, Temporal) Advanced working knowledge of relational and NoSQL databases, vector stores, data lake architectures, and data governance ...

Google Cloud Platform Architect

Hiring Organisation
Sanderson
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Contractor
Contract Rate
£400 - £500 per day
expertise in Go and/or Python Kubernetes platform engineering experience, including CRDs and API extensions. Strong knowledge of Kubernetes internals, multi-tenancy and observability Proven experience with CI/CD, automation, infrastructure as code and deployment pipelines. Strong understanding of SDLC, Agile delivery, resiliency and security Practical ...

Google Cloud Platform Architect

Hiring Organisation
Sanderson Recruitment
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Contract
Contract Rate
£400 - £500 per day
expertise in Go and/or Python Kubernetes platform engineering experience, including CRDs and API extensions. Strong knowledge of Kubernetes internals, multi-tenancy and observability Proven experience with CI/CD, automation, infrastructure as code and deployment pipelines. Strong understanding of SDLC, Agile delivery, resiliency and security Practical ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
Fluency in Python & deep knowledge of software applications and technical processes with emerging depth in one or more technical disciplines Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, etc. Proficiency in continuous ...

Data Engineer

Hiring Organisation
Tria
Location
Glasgow, City of Glasgow, United Kingdom
Employment Type
Permanent
Salary
£50000 - £60000/annum
also be useful. We're particularly interested in people who have worked with modern data platforms and understand the principles of ELT, automation, testing, observability, data quality and infrastructure-as-code . What You'll Bring A strong understanding of modern data engineering practices A genuine interest in building well ...

Lead Site Reliability Engineer - Chief Technology Office

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
.NET) Deep knowledge of software applications and technical processes with emerging depth in one or more technical disciplines Proficiency and hands-on experience in observability practices including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, or Splunk Proficiency ...

Software Engineer III - Cloud Data Platform (AWS/Databricks, Terraform)

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
testing Implement security best practices across the platform, including least-privilege access controls, secrets management, encryption, network segmentation, and auditability Improve platform reliability through observability tooling logging, metrics, and tracing alongside alerting, incident response practices, and performance and cost optimization Collaborate with stakeholders to translate business and technical requirements into ...

Software Developer

Hiring Organisation
McGregor Boyall Associates Limited
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Contract
OpenShift , RabbitMQ , PostgreSQL , Jenkins and GitLab CI/CD. Proven experience designing APIs, event-driven architectures and cloud-native applications. Strong understanding of DevOps, observability tools, Grafana/Kibana and software quality practices. Experience modernising legacy systems and delivering software within Agile environments. Excellent communication, collaboration and problem-solving skills. ...

Python Developer

Hiring Organisation
Diana Duggan UK Limited
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Contractor
Contract Rate
£450 per day
opportunities for increased automation and reduced manual intervention. Troubleshoot complex issues across application, automation, infrastructure, and deployment layers. Contribute to improvements in deployment reliability, observability, resilience, and operational support. Participate in code reviews and contribute to engineering standards, design decisions, and best practices. Collaborate with developers, QA engineers, DevOps, architects ...

Senior Site Reliability Engineer

Hiring Organisation
GCS
Location
Glasgow, City of Glasgow, United Kingdom
Employment Type
Permanent
Salary
£75000 - £95000/annum Bonus
drive reliability, scalability and performance across critical banking systems. This role combines hands-on SRE engineering with technical leadership, with a strong focus on observability, automation, continuous improvement and optimisation. Responsibilities: * Build and maintain reliable, scalable and secure infrastructure platforms and solutions. * Apply SRE and software engineering practices to improve … lead complex troubleshooting and root cause analysis. * Develop automation using programming and scripting to reduce manual intervention and improve efficiency. * Develop and improve observability, monitoring, instrumentation and performance capabilities. * Use data and reliability metrics to drive continuous improvement and optimisation. * Lead technical discussions, blameless retrospectives and problem-solving activities. * Work ...

Business Systems Engineer

Hiring Organisation
Head Resourcing
Location
Edinburgh, Ingliston, City of Edinburgh, United Kingdom
Employment Type
Permanent
Salary
£60000 - £70000/annum
Stack PHP/Laravel TypeScript, Node.js, NestJS React MySQL & PostgreSQL AWS (Lambda, ECS, RDS, S3, SQS, SNS) Docker & CI/CD pipelines Monitoring and observability tools What We're Looking For Strong experience in either PHP/Laravel or TypeScript/Node.js . Proven ownership of production software and business ...

Sr Lead Infrastructure Engineer - Devops AWS

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
machine learning products. As the team takes end-to-end ownership of the platforms it runs, you will build the CI/CD, observability, and incident-management practices that keep those services stable, secure, and performant across international markets. This is a Vice President-level role and an integral part … release automation, and deployment tooling Establishes reliability practices (SLOs, error budgets, runbooks) and leads production incident response and post-incident review Builds and operates observability across the team's AI/ML services (metrics, logging, tracing, alerting) Automates infrastructure provisioning and configuration through infrastructure-as-code Implements operational security, secrets ...

Lead Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
build, and operate REST and gRPC APIs and microservices, defining clear contracts using OpenAPI and Protobuf while ensuring backward compatibility, authentication, rate limiting, and observability Apply resilience engineering patterns - including timeouts, retries, and circuit breakers - to ensure reliable, production-grade service behavior Build and maintain well-tested, maintainable Python services … query optimization, indexing, and transaction management Demonstrated experience building AI solutions using large language models in production environments, including quality assurance, safety controls, evaluation, observability, and cost management Strong API and microservices engineering experience, including service design, contract definition, security patterns, performance tuning, and distributed system observability Hands-on multi ...

Infrastructure Automation Engineer

Hiring Organisation
Searchability NS&D
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Full-Time
Salary
£45,000 - £55,000 per annum
changes while improving resilience and reducing risk. Working within a highly skilled engineering team, you'll help strengthen pre- and post-change validation, enhance observability, and support the continuous improvement of automation capabilities. Technology Stack Linux/UNIX Python Ansible Apache Airflow Prometheus Grafana Loki VMware F5 What ...

Linux Engineer

Hiring Organisation
Searchability NS&D
Location
Glasgow, Scotland, United Kingdom
changes while improving resilience and reducing risk. Working within a highly skilled engineering team, you'll help strengthen pre and post-change validation, enhance observability, and support the continuous improvement of automation capabilities. Technology Stack Linux/UNIX Python Ansible Apache Airflow Prometheus Grafana Loki VMware F5 What ...

Principal AI Engineer

Hiring Organisation
Blend
Location
Edinburgh, Scotland, United Kingdom
perform reliably at scale. Recent, personal experience architecting and building production AI systems. Expect to discuss retrieval strategy, caching, context economics, serving constraints, evaluation, observability and what went wrong in practice. Strong systems thinking. You can reason through unfamiliar platforms and problems rather than relying on expertise in a single ...

DevOps Engineer

Hiring Organisation
CGI
Location
Glasgow City, United Kingdom
Employment Type
Full Time
clients to deliver with confidence. As a DevOps Engineer, you'll work across Microsoft Azure and AWS, using cloud-native technologies, automation and observability to improve application delivery and platform performance. You'll have the freedom to take ownership of challenges, introduce creative improvements and make a measurable impact … hands-on with Kubernetes and Jenkins to create reliable, repeatable delivery processes and resolve challenges across applications and infrastructure. You'll also help strengthen observability across our environments, using monitoring, metrics and visualisation to identify issues early and improve platform performance. Supported by collaborative development, cloud, platform and operations teams ...

Google Cloud Platform (GCP) Architect

Hiring Organisation
Diana Duggan UK Limited
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Contractor
Contract Rate
£450 per day
operational excellence by proactively identifying patterns in system failures, operational metrics, and data; designing and implementing systematic improvements to system reliability, performance, and observability Evaluate and lead vendor/technology assessments - leading sessions with external vendors, startups, and internal teams to drive outcomes-oriented evaluation of architectural designs, technical credentials … automation Frameworks on Kubernetes, including authoring reconciliation loops, admission controllers, webhooks, and custom controllers; strong understanding of Kubernetes internals, API machinery, RBAC, multi-tenancy, observability, and operational best practices for production environments. In-depth Google Cloud development experience -architecting, building, deploying, and operating production cloud-native workloads using core ...

Sr Lead AI Platform Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
Owns the design and build of the team's platform: deployment pipelines, model serving, containerisation, orchestration, and environment management Sets the standard for reliability, observability, and operational excellence across the team's production AI/ML services Builds the tooling and paved paths that let AI engineers ship agentic … record of building deployment and release automation Experience serving, scaling, and monitoring ML models or data-intensive services in production (MLOps) Practical experience with observability tooling (metrics, logging, tracing) and production incident response Experience operating ML/LLM workloads in production (LLMOps, inference reliability, cost/performance management) Strong communication ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
practices within an application or platform Fluency in at least one programming language such as (e.g., Java, Python, Go, etc.) Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency … high-availability services Deep understanding of distributed system design principles, networking (TCP/IP, DNS, load balancing), Linux internals. Contributions to open-source observability or telemetry projects. Experience working with agent control planes and management protocols. Hands-on knowledge of OpAMP is highly desirable. ABOUT US J.P. Morgan ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … Amazon EKS and Amazon SageMaker, as well as on-prem and local GPU clusters, using reproducible infrastructure as code and continuous delivery pipelines Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads Tune GPU and accelerator ...