1 to 25 of 28 Observability Jobs in Lanarkshire

Software Engineer - DevOps

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
issues and partnering with development teams to resolve them. Maintain and monitor asset inventory across environments. Monitor, troubleshoot, and remediate issues using Splunk and observability/monitoring platforms such as Datadog, Dynatrace, or Grafana. Support cost rationalization efforts; partner with architects to gather, clarify, and translate technical requirements. Proficient with ...

Lead SRE- Azure & GCP

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
with good understanding of REST APIs Hands-on experience with cloud-based technologies and tools especially in deployment, monitoring and operations, such as Google Observability, Azure Monitor, Data Dog, Prometheus, Splunk, Elasticsearch and Grafana. Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g. ...

Senior Manager of SRE

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
drive outcomes-oriented probing of architectural designs, technical credentials, and applicability for use within existing systems and information architecture. Drives continuous improvement in system observability, alerting, and capacity planning. Collaborates with engineering and data teams to optimize infrastructure and deployment processes, focusing on automation and operational excellence. Performs platform design ...

Senior AI Engineer

Hiring Organisation
83zero Limited
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
Salary
£70,000
Experience with CI/CD using GitHub, GitLab or Jenkins Agile engineering experience Experience with AI agents, tool calling, embeddings, prompt engineering or LLM observability is beneficial React/TypeScript and Terraform/IaC experience is beneficial Experience taking GenAI POCs into production is highly beneficial Why this role ...

Java Software Engineer - VP

Hiring Organisation
Henderson Scott
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
Data: Proficiency with Docker, Kubernetes, and relational databases (MS SQL, Sybase). Delivery Automation & Monitoring: Strong focus on automated testing, automated release pipelines, and observability tools (Grafana, Prometheus). Mindset: Delivery-focused problem solver with a hands-on approach and strong stakeholder communication skills. Desirable Experience Background in Equity Swaps ...

Software Engineer III - Data Engineering- Corporate Know Your Customer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
depth in disciplines such as cloud, AI/ML, or data engineering Experience in large-scale data processing, microservices, API design, Kafka, Redis, MemCached, observability tools (Dynatrace, Splunk, Grafana), and orchestration frameworks (Airflow, Temporal) Advanced working knowledge of relational and NoSQL databases, vector stores, data lake architectures, and data governance ...

Lead Software Data Engineer - Corporate Know Your Customer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
depth in disciplines such as cloud, AI/ML, or data engineering Experience in large-scale data processing, microservices, API design, Kafka, Redis, MemCached, observability tools (Dynatrace, Splunk, Grafana), and orchestration frameworks (Airflow, Temporal) Advanced working knowledge of relational and NoSQL databases, vector stores, data lake architectures, and data governance ...

Python Developer - GenAI

Hiring Organisation
83zero Limited
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
Salary
£55,000
Experience with CI/CD using GitHub, GitLab or Jenkins *Agile engineering experience *Experience with AI agents, tool calling, embeddings, prompt engineering or LLM observability is beneficial *React/TypeScript and Terraform/IaC experience is beneficial *Experience taking GenAI POCs into production is highly beneficial Why this role ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
Fluency in Python & deep knowledge of software applications and technical processes with emerging depth in one or more technical disciplines Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, etc. Proficiency in continuous ...

Lead Site Reliability Engineer - Chief Technology Office

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
.NET) Deep knowledge of software applications and technical processes with emerging depth in one or more technical disciplines Proficiency and hands-on experience in observability practices including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, or Splunk Proficiency ...

Lead Site Reliability Engineer - Chief Technology Office

Hiring Organisation
JP Morgan Chase
Location
Glasgow, Lanarkshire, United Kingdom
Salary
£ 70 K
Boot, .NET)Deep knowledge of software applications and technical processes with emerging depth in one or more technical disciplinesProficiency and hands-on experience in observability practices including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, or SplunkProficiency ...

Associate Director, Data Science/Gen AI Lead - ER&I

Hiring Organisation
Deloitte
Location
Glasgow, Lanarkshire, United Kingdom
Salary
£ 70 K
/GenAI governance & ethics (bias detection, explainability). GenAI Platform & Infrastructure Architecture (Cloud, Lakehouse). GenAI ModelOps & Performance Monitoring. AI-driven business intelligence & reporting. Observability & FinOps for AI/GenAI. Cloud Infrastructure, Networking, & Security for AI.Aligning GenAI Architectures Across Organizations: Experience aligning GenAI architecture blueprints across business units and geographies ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
practices within an application or platform Proficiency in at least one programming language such as Python, Java/Spring Boot, or .NET Experience in observability practices such as white and black box monitoring, service level objective alerting, and telemetry collection Proficient knowledge of software applications and technical processes within ...

Site Reliability Engineer II - AI & Corporate Risk Tech

Hiring Organisation
JP Morgan Chase
Location
Glasgow, Lanarkshire, United Kingdom
Salary
£ 70 K
site reliability practices within an application or platformProficiency in at least one programming language such as Python, Java/Spring Boot, or .NETExperience in observability practices such as white and black box monitoring, service level objective alerting, and telemetry collectionProficient knowledge of software applications and technical processes within a given ...

Mathematical Software Engineer

Hiring Organisation
Spire Global
Location
Glasgow, Lanarkshire, United Kingdom
Salary
£ 70 K
orbital maneuver planning, or ground-system contact scheduling strategies.Iterate on features based on stakeholder feedback and evolving requirements from the business.Build and refine observability dashboards and alerts.Support live operations of constellation scheduling software.Improve the team's development environment and processes.Investigate and mitigate issues that arise in production systems.Key Skills:Strong ...

Sr Lead Infrastructure Engineer- Devops/AWS

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
machine learning products. As the team takes end-to-end ownership of the platforms it runs, you will build the CI/CD, observability, and incident-management practices that keep those services stable, secure, and performant across international markets. This is a Vice President-level role and an integral part … release automation, and deployment tooling Establishes reliability practices (SLOs, error budgets, runbooks) and leads production incident response and post-incident review Builds and operates observability across the team's AI/ML services (metrics, logging, tracing, alerting) Automates infrastructure provisioning and configuration through infrastructure-as-code Implements operational security, secrets ...

Java Full stack Developer

Hiring Organisation
NEEV LIMITED
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Contract
Contract Rate
£400 per day
enterprise applications using modern Java technologies, cloud-native architectures, and microservices. The ideal candidate will have hands-on experience with Kafka, Kubernetes, API Security, Observability tools, SQL databases, and Spring-based microservices development . Key Responsibilities Design, develop, and maintain scalable Java-based applications using Java 17+, Spring Boot … mechanisms. Build event-driven solutions using Apache Kafka for real-time data processing and messaging. Deploy, manage, and troubleshoot applications on Kubernetes environments. Implement observability solutions using tools such as Splunk, ELK, Grafana, Prometheus, Dynatrace, or AppDynamics. Optimize application performance, scalability, and reliability. Work closely with business stakeholders, architects ...

Sr Lead AI Platform Engineer

Hiring Organisation
JP Morgan Chase
Location
Glasgow, Lanarkshire, United Kingdom
Salary
£ 70 K
Tech.Job responsibilitiesOwns the design and build of the team's platform: deployment pipelines, model serving, containerisation, orchestration, and environment managementSets the standard for reliability, observability, and operational excellence across the team's production AI/ML servicesBuilds the tooling and paved paths that let AI engineers ship agentic … track record of building deployment and release automationExperience serving, scaling, and monitoring ML models or data-intensive services in production (MLOps)Practical experience with observability tooling (metrics, logging, tracing) and production incident responseExperience operating ML/LLM workloads in production (LLMOps, inference reliability, cost/performance management)Strong communication skills ...

Sr Lead AI Platform Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
Owns the design and build of the team's platform: deployment pipelines, model serving, containerisation, orchestration, and environment management Sets the standard for reliability, observability, and operational excellence across the team's production AI/ML services Builds the tooling and paved paths that let AI engineers ship agentic … record of building deployment and release automation Experience serving, scaling, and monitoring ML models or data-intensive services in production (MLOps) Practical experience with observability tooling (metrics, logging, tracing) and production incident response Experience operating ML/LLM workloads in production (LLMOps, inference reliability, cost/performance management) Strong communication ...

Cloud Infrastructure Engineer - Go production experience

Hiring Organisation
Uniting People
Location
Glasgow, Lanarkshire, United Kingdom
Employment Type
Contract
Contract Rate
GBP 475 Daily
experience: ForgeRock, Okta, Ping, Auth0 Kafka/event-streaming: producing/consuming events, setting up topics/consumers, reacting to events in a service Observability: Prometheus/Grafana/OTel ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
practices within an application or platform Fluency in at least one programming language such as (e.g., Java, Python, Go, etc.) Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency … high-availability services Deep understanding of distributed system design principles, networking (TCP/IP, DNS, load balancing), Linux internals. Contributions to open-source observability or telemetry projects. Experience working with agent control planes and management protocols. Hands-on knowledge of OpAMP is highly desirable. ABOUT US Our client ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … Amazon EKS and Amazon SageMaker, as well as on-prem and local GPU clusters, using reproducible infrastructure as code and continuous delivery pipelines Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads Tune GPU and accelerator ...

Corporate KYC Principle Software Engineer - Executive Director

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
regulated financial services environments Establishes engineering standards for LLM-based applications RAG pipelines, embedding workflows, vector store integrations, and model serving ensuring safety, observability, and reproducibility at scale Drives adoption of advanced technical methods and practices aligned with the latest industry standards and product development methodologies Serves as the function … more disciplines (e.g., cloud, AI/ML, data engineering) Experience in large-scale data processing, microservices, API design, Kafka, Redis, MemCached, observability tools (Dynatrace, Splunk, Grafana), and orchestration frameworks (Airflow, Temporal) Advanced working knowledge of relational and NoSQL databases, vector stores, data lake architectures, and data governance Practical cloud-native ...

Corporate KYC Sr Lead Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
scale data processing, microservices, API design, and orchestration frameworks Working knowledge of relational and NoSQL databases, vector stores, and data lake architectures Familiarity with observability tools and frameworks Practical cloud-native experience (AWS, Azure, or GCP) Ability to communicate effectively with senior leaders and executives Commitment to inclusive, collaborative teamwork … catalog services such as Apache Iceberg Experience with LLM orchestration frameworks and model serving infrastructure or managed endpoints Familiarity with AI evaluation and observability practices for LLM workloads Understanding of agentic design patterns and how to constrain agent autonomy in financial workflows Interest in emerging technologies and continuous learning Employer ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelines Implement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloads Tune GPU and accelerator capacity, autoscaling, and cost ...