151 to 175 of 213 OpenTelemetry Jobs

Observability SRE AVP: Cloud Observability & Migrations

Location
Greater London, England, United Kingdom
observability and resiliency initiatives in a large-scale environment. You will migrate monitoring tooling to Google Cloud Observability and Grafana, implement OpenTelemetry instrumentation, and author reusable deployment solutions for OpenShift/Kubernetes and VM environments. The role requires deep SRE knowledge, strong collaboration with application teams, and a focus ...

Software Engineer III - AI/ML Platform Reliability

Hiring Organisation
JP Morgan Chase
Location
Glasgow, Lanarkshire, United Kingdom
Salary
£ 80 K
Are you passionate about building resilient, scalable systems that power the future of AI? At JPMorganChase, we're pushing the boundaries of what's possible with artificial intelligence and machine learning — and we need engineers ...

Software Engineer III - AI ML Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Bournemouth, Dorset, South West, United Kingdom
Employment Type
Permanent
hackajob is partnering directly with JPMorganChase to hire for this role. JOB DESCRIPTION Are you passionate about building resilient, scalable systems that power the future of AI? At JPMorganChase, we're pushing the boundaries of ...

AI Platform SRE - Infra, Security & Observability

Location
United Kingdom
agents for enterprise customers. You will join a growing SRE team and help shape the reliability culture from the ground up, including Terraform modules, OpenTelemetry pipelines, and #J-18808-Ljbffr ...

Observability AVP: SRE & Cloud Telemetry Lead

Location
Greater London, England, United Kingdom
role focuses on reliability, performance, and resiliency across applications and services. You will migrate monitoring tools to Google Cloud Observability and Grafana using OpenTelemetry, author reusable deployment solutions, and onboard engineering teams to new telemetry standards. This is a senior technical role within Production Management. #J-18808-Ljbffr ...

Senior Lead SRE - Reliability & Observability Leader

Location
Glasgow, Scotland, United Kingdom
will mentor engineers, lead incident response, and shape SRE strategy while delivering scalable, secure production systems. The role demands deep expertise in cloud, automation, OpenTelemetry, and instrumentation, with a track record of improving service levels and reducing toil in large-scale #J-18808-Ljbffr ...

Staff AI Engineer

Location
Greater London, England, United Kingdom
retrieval‐augmented generation (RAG), memory structures, and context window optimisation. Production deployment. Experience deploying and monitoring agentic applications in production environments, including scaling, monitoring (OTEL), cost, and token usage. End‐to‐end ownership. You’ve taken agentic systems from concept through to production and are comfortable being accountable ...

Engineering Manager, Runtime Platform, Robot Software

Hiring Organisation
wayve
Location
London, United Kingdom
Salary
£ 80 K
haveExperience in automotive, robotics, or another safety-relevant real-time domain.Familiarity with profiling toolchains (pprof/gperftools, perf, Nsight, NVLumo) and observability stacks (OpenTelemetry, Grafana, Datadog).Experience with NVIDIA (Orin/Thor) and/or Qualcomm compute platforms.Experience with a micro-kernel and/or real-time OS (e.g. ...

Senior Observability Engineer

Location
Greater London, England, United Kingdom
those systems behave. This is a high-impact, hands-on role at the centre of our reliability engineering agenda: building on Honeycomb and OpenTelemetry, driving instrumentation across a globally distributed microservices estate, and partnering directly with development teams to turn telemetry data into better, faster software.**About the team**This … distributed trading environment.* Define platform standards, data models, and integration patterns for telemetry collection, storage, and querying across the estate.**Instrumentation and telemetry*** Drive OpenTelemetry adoption across engineering teams, providing hands-on guidance and reusable instrumentation patterns for services built in Java, Python, and C++.* Partner with development teams ...

Senior Observability Engineer

Location
City Of London, England, United Kingdom
into how those systems behave. This is a high, hands-on role at the centre of our reliability engineering agenda: building on Honeycomb and OpenTelemetry, driving instrumentation across a globally distributed microservices estate, and partnering directly with development teams to turn telemetry data into better, faster software. About the team … distributed trading environment. Define platform standards, data models, and integration patterns for telemetry collection, storage, and querying across the estate. Instrumentation and telemetry – Drive OpenTelemetry adoption across engineering teams, providing hands-on guidance and reusable instrumentation patterns for services built in Java, Python, and C++. Partner with development teams ...

DevOps and Automation Engineer (Contract)

Location
Greater London, England, United Kingdom
environment provisioning using Infrastructure-as-Code (Terraform) and pipeline-driven automation.* Implement advanced observability and monitoring: Use platforms such as Datadog, Prometheus, Grafana, and OpenTelemetry to provide real-time insights into system health, deployments, and business metrics.* Embed security and compliance by design: Integrate security into every stage … innovation sprints to generate new automation ideas and create user stories for rapid prototyping.* Use technologies like Terraform, Ansible, Azure DevOps, Github, OctopusDeploy, Kubernetes, OpenTelemetry, Datadog, Grafana, Prometheus, low-code automation platforms (e.g., Power Automate, UiPath) to help evolve the team's capabilities.* Coordinate planned outage and environment refreshes ...

Lead Developer

Location
United Kingdom
test, container image scanning, and blue-green deployment automation Own infrastructure-as-code using Bicep or Terraform, and embed observability from day one using OpenTelemetry alongside Application Insights Champion test-driven development, including meaningful unit testing with xUnit, integration testing with Playwright and consumer-driven contract testing between services … microservices migration Firsthand experience with .NET Aspire for local multi-service orchestration Experience with infrastructure-as-code: Bicep (preferred) or Terraform Knowledge of OpenTelemetry, distributed tracing, and structured logging at scale Experience with Azure API Gateways - versioning, rate limiting, developer portals Understanding of event-driven architecture patterns: sagas, outbox pattern ...

Lead DevOps Engineer

Hiring Organisation
Elliptic
Location
London, United Kingdom
Salary
£ 80 K
Elliptic has helped trace and disrupt over $21.8 billion in illicit crypto laundered across blockchains, from sanctioned nation states to organized crime networks hiding funds through token swaps and unregulated exchanges. It’s how compliance ...

Software Engineer III - AI/ML Platform Reliability

Location
Glasgow, Scotland, United Kingdom
Are you passionate about building resilient, scalable systems that power the future of AI? At JPMorganChase, we're pushing the boundaries of what's possible with artificial intelligence and machine learning - and we need engineers ...

Lead Site Reliability Engineer - Observability & Resilience

Location
Glasgow, Scotland, United Kingdom
reliability for a global firm. You will lead critical resiliency design reviews, break complex problems into actionable work, and mentor engineers across large-scale OpenTelemetry pipelines in hybrid environments. You will guide incident responses, drive data‐driven improvements to SLOs and error budgets, and influence cross‐team decisions with ...

Staff Software Engineer, Observability & Profiling

Hiring Organisation
Humanloop
Location
London, United Kingdom
Salary
> £ 150 K
experience and performance engineering craftExperience profiling or instrumenting accelerator workloadsExperience operating metrics systems at very high cardinality, or large-scale telemetry storage backendsExperience with OpenTelemetry instrumentation, collector pipelines, and tail-based sampling strategiesInterest in applying AI/LLMs to operational workflows such as automated root cause analysis, anomaly detection ...

Software Engineer — Observability Instrumentation

Hiring Organisation
G Research
Location
London, United Kingdom
Salary
£ 80 K
strong, customer-focused Software Engineer to help make observability easier to adopt across the organisation. This role focuses on the producer side: instrumentation patterns, OpenTelemetry SDKs and the collector configurations that help teams emit consistent, high-quality telemetry.This role is suited to someone who enjoys creating leverage through shared engineering … instrumentation across C#, Python and Kubernetes-based services, and ongoing development of our observability capabilities.Key responsibilitiesKey responsibilities of the role include:Extending and maintaining OpenTelemetry SDKs, libraries and CollectorsBuilding and supporting instrumentation paths across Kubernetes and non-Kubernetes workloadsEmbedding observability standards across platform and application teamsLeveraging AI and AI assisted ...

Lead Cloud Engineer

Location
Leeds, England, United Kingdom
security and cost controls, real‐time cost‐anomaly detection, and cost estimation into the Software Development Life Cycle (SDLC) as a standard. Understanding of OpenTelemetry or other telemetry approaches. Internal Platform Support: Assist the Capability Lead and architects in defining and delivering an internal collection of reusable components, reference architectures …/ELT methodologies, real‐time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost‐control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated Industries: Experience working in highly regulated sectors such ...

Lead Cloud Engineer

Location
Manchester, England, United Kingdom
security and cost controls, real‐time cost‐anomaly detection, and cost estimation into the Software Development Life Cycle (SDLC) as a standard. Understanding of OpenTelemetry or other telemetry approaches. Internal Platform Support: Assist the Capability Lead and architects in defining and delivering an internal collection of reusable components, reference architectures …/ELT methodologies, real‐time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost‐control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated Industries: Experience working in highly regulated sectors such ...

Lead Cloud Engineer

Location
Greater London, England, United Kingdom
security and cost controls, real‐time cost‐anomaly detection, and cost estimation into the Software Development Life Cycle (SDLC) as a standard. Understanding of OpenTelemetry or other telemetry approaches. Internal Platform Support: Assist the Capability Lead and architects in defining and delivering an internal collection of reusable components, reference architectures …/ELT methodologies, real‐time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost‐control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated Industries: Experience working in highly regulated sectors such ...

Lead Cloud Engineer

Location
City of Edinburgh, Scotland, United Kingdom
security and cost controls, real‐time cost‐anomaly detection, and cost estimation into the Software Development Life Cycle (SDLC) as a standard. Understanding of OpenTelemetry or other telemetry approaches. Internal Platform Support: Assist the Capability Lead and architects in defining and delivering an internal collection of reusable components, reference architectures …/ELT methodologies, real‐time streaming (e.g., AWS Kinesis, Apache Kafka), Vector databases, and RAG architectures. Observability & FinOps: Experience implementing modern observability tooling (OpenTelemetry) alongside automated cost‐control systems (such as Karpenter, Infracost, OpenCost, or Cloud Custodian). Domain & Sector Experience Regulated Industries: Experience working in highly regulated sectors such ...

Senior Cloud Engineer

Hiring Organisation
Entrust Corporation
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 80 K
Join us at Entrust At Entrust, we’re shaping the future of identity centric security solutions. From our comprehensive portfolio of solutions to our flexible, global workplace, we empower careers, foster collaboration, and build solutions ...

Senior Backend Engineer - Databases - Loki Query | UK | Remote

Hiring Organisation
Grafana Labs
Location
United Kingdom
Salary
£ 70 K
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source ...

Staff Platform Engineer

Location
Greater London, England, United Kingdom
Staff Platform Engineer Department: Engineering Employment Type: Full Time Location: London Description Hybrid: 2 days per week in our Tower Bridge office (Tuesday/Wednesday). RVU is a group of online brands that include ...