1 to 25 of 147 OpenTelemetry Jobs in London

Observability Engineer - Assistant Vice President

Location
Greater London, England, United Kingdom
drive the migration of applications from existing monitoring tools (Geneos ITRS, Prometheus, ELK, Splunk, AppDynamics, etc.) to Google Cloud Observability (GCO) and Grafana using OpenTelemetry (OTel) as the instrumentation standard. You will act as a hands-on technical authority, authoring reusable deployment solutions, configuring telemetry collectors, and providing direct technical … transparency, innovation, and technical excellence that encourages continuous improvement and automation. Collaborative Enablement: Partner with development and SRE teams to drive the adoption of OpenTelemetry (OTel) and Google Cloud Observability (GCO) and Grafana standards. Regulatory Compliance: Operate effectively within a highly regulated environment, ensuring all observability and deployment solutions comply ...

ML Ops Engineer

Location
Greater London, England, United Kingdom
telemetry to track unit economics and throughput for training and serving AI models. Implement end-to-end observability using tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow. Your Skills Hands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated ...

ML Ops Engineer

Hiring Organisation
Anaplan
Location
London, UK
Employment Type
Full-time
telemetry to track unit economics and throughput for training and serving AI models. Implement end-to-end observability using tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow. Your SkillsHands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated to AI/ ...

Senior DevOps Engineer (Azure)

Hiring Organisation
Darktrace
Location
London, UK
Employment Type
Full-time
supporting technologies like ArgoCD and Helm. You should also be familiar and well-versed in monitoring, logging, and observability tools including: Prometheus; Grafana; Loki; OpenTelemetry; and the ELK stack. Amongst this, you should be able to demonstrate: Proven experience as a DevOps Engineer, with a solid background in building ...

Strategic DevSecOps Consultant

Hiring Organisation
CloudBees
Location
London, UK
Employment Type
Full-time
.Familiarity with AI-enabled software development, agentic workflows, large language models (LLMs), or AI governance practices. Experience with observability and telemetry platforms such as OpenTelemetry, Splunk, Dynatrace, Datadog, AppDynamics, Grafana, or similar technologies. Experience working with large-scale enterprise architecture, governance, compliance, and regulated environments. Thought leadership experience through technical ...

Software Engineering Tech Lead (SRE + AI)

Location
Greater London, England, United Kingdom
agents, MCP tool integrations, and deterministic evaluation pipelines for automated operational decision support. Telemetry & Insights: Architect ingestion and correlation pipelines across distributed logs, metrics, OpenTelemetry traces, change events, and runbooks to accelerate Mean Time to Detection (MTTD) and Resolution (MTTR). Safe Production Automation: Develop proactive anomaly detection and Human … Hands‐on experience building LLM pipelines, AI Agents, Model Context Protocol (MCP) servers/clients, RAG architectures, and evaluation frameworks. Observability & Telemetry: Experience with OpenTelemetry (OTel), Prometheus, Grafana, Splunk, ThousandEyes, or distributed tracing systems. Cloud & Infrastructure: Expertise in public cloud providers (AWS, GCP, Azure), Terraform/IaC, and GitOps/ ...

Principal AI Platform Engineer (Python)

Location
Greater London, England, United Kingdom
test, package, and release software, and the practices around them, such as automated testing and staged rollouts. Observability: experience instrumenting systems with Prometheus, Grafana, OpenTelemetry, or an equivalent stack, and using that data to diagnose failures in distributed systems. Security (critical): a working grasp of secrets management, identity and access ...

Cloud Native Specialist

Location
Greater London, England, United Kingdom
OpenShift) and container orchestration.* Demonstrate how Dynatrace provides automated, code-level visibility into microservices without manual instrumentation or sidecar overhead.* Advocate for OpenTelemetry (OTel) integration and explain how Dynatrace extends the value of open-source telemetry in a production-grade environment.2. Domain Execution:* Lead technical discovery and high-stakes Proof ...

Analytics Services Platform Engineer

Location
Greater London, England, United Kingdom
technologies including EMR, MSK, Athena, Redshift, Glue and MWAA Experience with CI/CD and observability tools such as Jenkins, ArgoCD, Prometheus, Grafana and OpenTelemetry Strong problem‐solving skills and a systematic approach to diagnosing and resolving issues Highly Desirable Skills Experience with streaming frameworks such as Flink, Kafka Streams ...

Platform Engineer

Location
Greater London, England, United Kingdom
/CD systems (GitHub Actions preferable, ArgoCD etc). Experience with monitoring, alerting and logging stacks (the Grafana stack: Prometheus, Loki, Tempo; plus OpenTelemetry). A working understanding of networking and distributed systems. A working understanding of security and an interest in DevSecOps. An ability to work through ambiguity ...

Site Reliability Engineer - Service Assurance Systems

Location
Greater London, England, United Kingdom
systems at scale. Familiarity with infrastructure-as-code tools such as Terraform or Ansible. Experience with log aggregation and analysis platforms such as the OTEL Stack or AWS CloudWatch Logs Insights. Exposure to Kubernetes or other container orchestration platforms. Experience working in an Agile or DevOps team environment. EEO Statement ...

Site Reliability Engineering (SRE) / Observability Technical Lead

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
roles, with leadership responsibilities. Proven expertise with Application Performance Monitoring (APM) tools such as New Relic, Datadog, AppDynamics, or Dynatrace. Hands-on experience with OpenTelemetry (OTel) for distributed tracing and observability instrumentation. Strong proficiency in Infrastructure as Code (IaC) using Terraform. Solid understanding of cloud platforms including ...

Platform Compute Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
storage, networking, IAM)Infrastructure-as-Code (e.g., Terraform, Pulumi)Monitoring and alerting (e.g., Prometheus, Grafana, Datadog, Zabbix)Logging and tracing (e.g., ELK stack, Fluentd, OpenTelemetry, JaegerIdentity and access management (e.g., LDAP, Kerberos, OAuth2)Cloud-native services (e.g., S3, EBS, GKE, EKS, Cloud Functions)Git-based workflows and version controlAgile methodologies ...

Database Platform Engineer

Location
Greater London, England, United Kingdom
services relevant to data platforms such as RDS, Aurora, S3, EC2 or EKS Familiarity with modern observability stacks such as Prometheus, Grafana, Elk or OTel Desirable: experience with cloud‐native and distributed SQL databases such as Aurora, YugabyteDB or TiDB Desirable: knowledge of data streaming and integration tools such ...

AWS DevOps Engineer

Location
Greater London, England, United Kingdom
pipelines using GitHub Actions or GitLab CI/CD, plus ArgoCD (GitOps). Experience with monitoring/observability tools (Prometheus, Grafana, ELK stack, Datadog, OpenTelemetry) – including metrics, logs, traces, dashboards, alerting, and integration with AWS services (e.g., CloudWatch, X‐Ray). Solid understanding of AWS EKS IAM concepts, especially ...

Staff Engineer - Money, Risk & Payment Ancillaries Yuno Totalmente remoto · Mundial ayer

Location
Greater London, England, United Kingdom
gRPC, REST Frameworks — Spring Boot, Spring WebFlux; Go standard library Messaging — Apache Kafka, SQS Databases — PostgreSQL, Redis Infrastructure — AWS, Kubernetes, Docker, Terraform Observability — Datadog, OpenTelemetry CI/CD — GitHub Actions, ArgoCD Version Control — Git/GitHub What We Offer at Yuno Competitive Compensation Remote Work — you can work from everywhere ...

Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President

Location
Greater London, England, United Kingdom
custom agentic loops) to coordinate multi-step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root-cause analysis, anomaly detection, and semantic log clustering. Self-healing Infrastructure Engineering: Experience designing closed-loop, self-healing ...

Camunda Architect / Lead Architect (Camunda 8 Preferred)

Location
Greater London, England, United Kingdom
SaaS and self-managed deployments at scale Exposure to cloud platforms such as: AWS Azure GCP Experience with observability tooling such as: Prometheus Grafana OpenTelemetry Experience with: enterprise integration patterns API management workflow/task UI customisation Experience in regulated industries such as: Banking Insurance Healthcare Ideal Profile This role ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
secure and fit for purpose. Technology Environment AWS Terraform Docker Kafka/AWS MSK Couchbase MongoDB Redis ClickHouse GitHub Actions/GitLab CI Grafana OpenTelemetry Python Golang What We're Looking For Strong experience building and operating AWS-based infrastructure platforms. Expert-level Terraform experience within enterprise production environments. Extensive ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, UK
Employment Type
Full-time
latency, cost, hallucination/quality metrics) using tools such as Arize, WhyLabs, or Langfuse. Logging & tracing: centralized logging (ELK/OpenSearch) and distributed tracing (OpenTelemetry) across data, training, and inference pipelines for end-to-end root-cause analysis. Integration — AI Stack, Enterprise Networks & Service Provider EnvironmentsPlatform integration: API-based ...

Platform Engineer - Monitoring, Observability & SIEM (MONSO)

Location
City of Westminster, England, United Kingdom
pipelines, code reviews, and self-service enablement. Modern Observability & Telemetry: Strong background in Splunk (SPL, dashboards, alerts, data ingestion, forwarders) and also configuring OpenTelemetry collectors and pipelines; knowledge of Prometheus and Grafana or similar tools. Kubernetes (Power User): Strong, hands‐on experience deploying and operating workloads, stateful appli‐cations, Helm ...

Senior Forward Deployment Engineer

Location
Greater London, England, United Kingdom
failover, and production‐recovery exercises, highlighting skills in system reliability and continuity planning. Experience with enterprise observability tools such as Splunk, ELK, Grafana, Prometheus, OpenTelemetry, AppDynamics, or Dynatrace, reflecting proficiency in monitoring and diagnostics. Experience modernizing monolithic or legacy enterprise applications into maintain #J-18808-Ljbffr ...

Expert Forward Deployment Engineer

Hiring Organisation
Luxoft
Location
London, UK
Employment Type
Full-time
failover, and production-recovery exercises, highlighting skills in system reliability and continuity planning. Experience with enterprise observability tools such as Splunk, ELK, Grafana, Prometheus, OpenTelemetry, AppDynamics, or Dynatrace, reflecting proficiency in monitoring and diagnostics. Experience modernizing monolithic or legacy enterprise applications into maintain OtherLanguages English: C1 Advanced Seniority Senior London ...

DevOps, AI and Automation Engineer

Location
Greater London, England, United Kingdom
DevOps, GitHub Actions, CheckMarx, JFrog Artifactory, Octopus Deploy* Infrastructure & Automation: Terraform, Ansible* Containers & Platforms: Docker, Kubernetes, Rancher, Azure, Windows, Linux* Observability: Datadog, Grafana, Prometheus, OpenTelemetry* AI Enablement: Azure AI Foundry, Model Context Protocol (MCP)**Expected Outcomes*** Establish scalable and secure CI/CD processes for AI-enabled platforms and services. ...

Devops SRE

Location
Greater London, England, United Kingdom
Performance Strong security mindset with a proven track record of designing secure, resilient cloud‐native systems. Experience implementing observability stacks including Prometheus , Dynatrace , and OpenTelemetry . Deep understanding of Linux internals , system performance tuning, and troubleshooting. Familiarity with Aqua Security for container runtime protection. CI/CD & Automation Tooling Hands ...