22 of 22 OpenTelemetry Jobs in Scotland

Software Engineer III - Java

Location
Glasgow, Scotland, United Kingdom
Docker and Kubernetes Understanding of Infrastructure as Code tools (e.g., Terraform) and environment provisioning Experience with observability tools and concepts (logging, metrics, tracing, OpenTelemetry) Exposure to integrating AI capabilities (semantic search, retrieval-based Q&A, content summarization) with attention to quality, privacy, and governance Understanding of modern application infrastructure design ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
JumpStart, SageMaker Endpoints, and Amazon Bedrock for managed model hosting Experience with online LLM quality monitoring (e.g., hallucination, toxicity, drift detection) and tracing via OpenTelemetry conventions Contributions to open-source LLM serving or inference projects (e.g., vLLM, llm-d, Ray, KServe, Triton) ABOUT US J.P. Morgan is a global leader ...

Lead Software Engineer - LLM Ops Platform Reliability

Location
Paisley, Scotland, United Kingdom
JumpStart, SageMaker Endpoints, and Amazon Bedrock for managed model hosting Experience with online LLM quality monitoring (e.g., hallucination, toxicity, drift detection) and tracing via OpenTelemetry conventions Contributions to open-source LLM serving or inference projects (e.g., vLLM, llm-d, Ray, KServe, Triton) ABOUT US J.P. Morgan is a global leader ...

Python/AIML Developer - Software Engineer III

Location
Glasgow, Scotland, United Kingdom
accuracy, hallucination and toxicity metrics, guardrails and content moderation, red-teaming, and embedding evals into CI/CD. Experience with AI/ML observability (OpenTelemetry, Phoenix, or equivalent tracing) and with model governance, auditability, and control requirements in a regulated financial-services environment. Familiarity with financial risk domain concepts - market ...

Senior Lead Site Reliability / DevOps Engineer

Location
Auchentibber, Scotland, United Kingdom
members and stakeholders to define comprehensive service level indicators, service level objectives, and error budgets Designs, implements, and maintains operational reliability for large-scale OpenTelemetry pipelines on hybrid on-prem/cloud environments, supporting telemetry ingestion, processing, and export to backends such as InfluxDB, Prometheus, Elasticsearch, and OpenSearch Drives … assessment, refactoring, and incremental migration of custom legacy telemetry collection code to standardized OpenTelemetry instrumentation, reducing technical debt while maintaining system stability Actively contributes to the engineering community as an advocate of firmwide frameworks, tools, and practices, and influences peers and project decision-makers to consider the use and application ...

Senior Lead Site Reliability Engineer

Location
Glasgow, Scotland, United Kingdom
members and stakeholders to define comprehensive service level indicators, service level objectives, and error budgets Designs, implements, and maintains operational reliability for large-scale OpenTelemetry pipelines on hybrid on-prem/cloud environments, supporting telemetry ingestion, processing, and export to backends such as InfluxDB, Prometheus, Elasticsearch, and OpenSearch Drives … assessment, refactoring, and incremental migration of custom legacy telemetry collection code to standardized OpenTelemetry instrumentation, reducing technical debt while maintaining system stability Actively contributes to the engineering community as an advocate of firmwide frameworks, tools, and practices, and influences peers and project decision-makers to consider the use and application ...

Senior DevOps Engineer

Location
City of Edinburgh, Scotland, United Kingdom
availability, scalability, and performance through automation and engineering improvements. Monitoring & Observability Implement enterprise monitoring, alerting, logging, and observability solutions. Grafana ELK/Elastic Stack OpenTelemetry Proactively identify issues before they impact customers and engineering teams. Develop dashboards and operational reporting for platform health and reliability. DevSecOps & Security Embed security best … working with API Gateway technologies such as: Strong understanding of: REST APIs OpenAPI Specifications OAuth2 JWT API Security Standards Experience with: Grafana ELK Stack OpenTelemetry Scripting & Automation Proficiency in one or more of: Python [optional] Bash Go PowerShell Networking & Security You'll thrive in this role if you: Have ...

Lead Site Reliability / DevOps Engineer

Location
Auchentibber, Scotland, United Kingdom
losses Documents and shares knowledge within your organization via internal forums and communities of practice Design, implement, and maintain operational reliability for large-scale Opentelemetry pipelines on hybrid on-prem/cloud environments. Support telemetry ingestion, processing, and export to backends such as InfluxDB, Prometheuse, Elasticsearch, OpenSearch optimizing for performance … real-time monitoring, logging, and alerting. Assess, refactor, and incrementally migrate custom legacy telemetry collection code to standardized OpenTelemetry instrumentation, reducing technical debt while maintaining system stability. Required qualifications, capabilities, and skills Formal training or certification on software engineering concepts and advanced applied experience Deep proficiency in reliability, scalability, performance ...

Cloud Infrastructure Engineer - Go production experience

Hiring Organisation
Uniting People
Location
Glasgow, Lanarkshire, United Kingdom
Employment Type
Contract
Contract Rate
GBP 430 Daily
Kafka/event-streaming: producing/consuming events, setting up topics/consumers, reacting to events in a service Observability: Prometheus/Grafana/OTel ...

Cloud Infra Engineer

Hiring Organisation
DNS Info Ltd
Location
Glasgow, Glasgow City, City of Glasgow, United Kingdom
Employment Type
Contract
Contract Rate
£350 - £400/day Inside IR 35
Kafka/event-streaming: producing/consuming events, setting up topics/consumers, reacting to events in a service Observability: Prometheus/Grafana/OTel ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
losses Documents and shares knowledge within your organization via internal forums and communities of practice Design, implement, and maintain operational reliability for large-scale Opentelemetry pipelines on hybrid on-prem/cloud environments. Support telemetry ingestion, processing, and export to backends such as InfluxDB, Prometheuse, Elasticsearch, OpenSearch optimizing for performance … real-time monitoring, logging, and alerting. Assess, refactor, and incrementally migrate custom legacy telemetry collection code to standardized OpenTelemetry instrumentation, reducing technical debt while maintaining system stability. Required qualifications, capabilities, and skills Formal training or certification on software engineering concepts and advanced applied experience Deep proficiency in reliability, scalability, performance ...

Software Engineer III - AI/ML Platform Reliability

Location
Auchentibber, Scotland, United Kingdom
Are you passionate about building resilient, scalable systems that power the future of AI? At JPMorganChase, we're pushing the boundaries of what's possible with artificial intelligence and machine learning - and we need engineers ...

Java Engineer, Associate

Hiring Organisation
Hackajob Ltd
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Work From Home
About this role At BlackRock, technology is the foundation of our business. As a Associate, Java Back-End Engineer, youll lead by example architecting, coding, and mentoring teams to build resilient systems that power our ...

Senior Lead SRE - Reliability & Observability Leader

Location
Glasgow, Scotland, United Kingdom
will mentor engineers, lead incident response, and shape SRE strategy while delivering scalable, secure production systems. The role demands deep expertise in cloud, automation, OpenTelemetry, and instrumentation, with a track record of improving service levels and reducing toil in large-scale #J-18808-Ljbffr ...

Lead Cloud Engineer

Location
City of Edinburgh, Scotland, United Kingdom
security and cost controls, real‐time cost‐anomaly detection, and cost estimation into the Software Development Life Cycle (SDLC) as a standard. Understanding of OpenTelemetry or other telemetry approaches. Internal Platform Support: Assist the Capability Lead and architects in defining and delivering an internal collection of reusable components, reference architectures …/ELT methodologies, real‐time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost‐control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated Industries: Experience working in highly regulated sectors such ...

Lead Cloud Engineer

Location
City of Edinburgh, Scotland, United Kingdom
security and cost controls, real-time cost-anomaly detection, and cost estimation into the Software Development Life Cycle (SDLC) as a standard. Understanding of OpenTelemetry or other telemetry approaches. Internal Platform Support: Assist the Capability Lead and architects in defining and delivering an internal collection of reusable components, reference architectures …/ELT methodologies, real-time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost-control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated Industries: Experience working in highly regulated sectors such ...

Senior Site Reliability Lead: Resiliency & Telemetry Architect

Location
Glasgow, Scotland, United Kingdom
Bank. You will guide resiliency design reviews, mentor engineers, and lead initiatives to improve reliability for large-scale systems. You will champion observability across OpenTelemetry pipelines, manage on-prem/cloud deployments, and collaborate with stakeholders to define SLOs, error budgets, and incident response playbooks. #J-18808-Ljbffr ...

Senior Site Reliability & DevOps Leader

Location
Glasgow, Scotland, United Kingdom
/cloud environments. In this role, you will conduct resiliency reviews, break complex problems into actionable work, and own the design and operation of OpenTelemetry pipelines, ensuring robust telemetry, monitoring, and rapid incident response to protect #J-18808-Ljbffr ...

Lead Site Reliability Engineer - Observability & Resilience

Location
Glasgow, Scotland, United Kingdom
reliability for a global firm. You will lead critical resiliency design reviews, break complex problems into actionable work, and mentor engineers across large-scale OpenTelemetry pipelines in hybrid environments. You will guide incident responses, drive data‐driven improvements to SLOs and error budgets, and influence cross‐team decisions with ...

Staff Software Engineer

Location
Uddingston, Scotland, United Kingdom
Are you the engineer other engineers trust when the stakes are highest? Do you thrive on solving complex technical challenges, setting technical direction and helping teams deliver scalable, reliable systems that support business growth? Are ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
JOB DESCRIPTION Help shape how AI systems run reliably in production at scale. In this role, you'll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Location
Auchentibber, Scotland, United Kingdom
Help shape how AI systems run reliably in production at scale. In this role, you’ll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge ...