1 to 25 of 111 Permanent OpenTelemetry Jobs in London

Observability Engineer - Assistant Vice President

Location
Greater London, England, United Kingdom
drive the migration of applications from existing monitoring tools (Geneos ITRS, Prometheus, ELK, Splunk, AppDynamics, etc.) to Google Cloud Observability (GCO) and Grafana using OpenTelemetry (OTel) as the instrumentation standard. You will act as a hands‐on technical authority, authoring reusable deployment solutions, configuring telemetry collectors, and providing direct technical … transparency, innovation, and technical excellence that encourages continuous improvement and automation. Collaborative Enablement: Partner with development and SRE teams to drive the adoption of OpenTelemetry (OTel) and Google Cloud Observability (GCO) and Grafana standards. Regulatory Compliance: Operate effectively within a highly regulated environment, ensuring all observability and deployment solutions comply ...

Observability Engineer - Assistant Vice President

Location
Greater London, England, United Kingdom
drive the migration of applications from existing monitoring tools (Geneos ITRS, Prometheus, ELK, Splunk, AppDynamics, etc.) to Google Cloud Observability (GCO) and Grafana using OpenTelemetry (OTel) as the instrumentation standard. You will act as a hands-on technical authority, authoring reusable deployment solutions, configuring telemetry collectors, and providing direct technical … transparency, innovation, and technical excellence that encourages continuous improvement and automation. Collaborative Enablement: Partner with development and SRE teams to drive the adoption of OpenTelemetry (OTel) and Google Cloud Observability (GCO) and Grafana standards. Regulatory Compliance: Operate effectively within a highly regulated environment, ensuring all observability and deployment solutions comply ...

ML Ops Engineer

Location
Greater London, England, United Kingdom
telemetry to track unit economics and throughput for training and serving AI models. Implement end-to-end observability using tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow. Your Skills Hands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated ...

ML Ops Engineer

Hiring Organisation
Anaplan
Location
London, United Kingdom
Salary
£ 80 K
benchmarking and telemetry to track unit economics and throughput for training and serving AI models.Implement end-to-end observability using tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow.Your SkillsHands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated to AI/ ...

Senior DevOps Engineer (Azure)

Hiring Organisation
Darktrace
Location
London, United Kingdom
Salary
£ 80 K
supporting technologies like ArgoCD and Helm. You should also be familiar and well-versed in monitoring, logging, and observability tools including: Prometheus; Grafana; Loki; OpenTelemetry; and the ELK stack. Amongst this, you should be able to demonstrate:Proven experience as a DevOps Engineer, with a solid background in building ...

Strategic DevSecOps Consultant

Hiring Organisation
CloudBees
Location
London, UK
Employment Type
Full-time
.Familiarity with AI-enabled software development, agentic workflows, large language models (LLMs), or AI governance practices. Experience with observability and telemetry platforms such as OpenTelemetry, Splunk, Dynatrace, Datadog, AppDynamics, Grafana, or similar technologies. Experience working with large-scale enterprise architecture, governance, compliance, and regulated environments. Thought leadership experience through technical ...

Strategic DevSecOps Consultant

Hiring Organisation
CloudBees
Location
London, United Kingdom
Salary
£ 80 K
IDPs).Familiarity with AI-enabled software development, agentic workflows, large language models (LLMs), or AI governance practices.Experience with observability and telemetry platforms such as OpenTelemetry, Splunk, Dynatrace, Datadog, AppDynamics, Grafana, or similar technologies.Experience working with large-scale enterprise architecture, governance, compliance, and regulated environments.Thought leadership experience through technical publications, conference ...

Senior Cloud Engineer, AI Platform SRE

Location
Greater London, England, United Kingdom
Built and maintained CI/CD with GitHub Actions, GitLab CI, Argo CD, Jenkins or similar. Observability: Hands‐on with Datadog, Prometheus, Grafana or OpenTelemetry, and opinionated about what's worth alerting on. Incident management: Calm, methodical instincts under pressure, and a habit of fixing the class of problem rather ...

Cloud Native Specialist

Location
Greater London, England, United Kingdom
OpenShift) and container orchestration.* Demonstrate how Dynatrace provides automated, code-level visibility into microservices without manual instrumentation or sidecar overhead.* Advocate for OpenTelemetry (OTel) integration and explain how Dynatrace extends the value of open-source telemetry in a production-grade environment.2. Domain Execution:* Lead technical discovery and high-stakes Proof ...

Platform Compute Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, United Kingdom
Salary
£ 100 K
storage, networking, IAM)Infrastructure-as-Code (e.g., Terraform, Pulumi)Monitoring and alerting (e.g., Prometheus, Grafana, Datadog, Zabbix)Logging and tracing (e.g., ELK stack, Fluentd, OpenTelemetry, JaegerIdentity and access management (e.g., LDAP, Kerberos, OAuth2)Cloud-native services (e.g., S3, EBS, GKE, EKS, Cloud Functions)Git-based workflows and version controlAgile methodologies ...

Site Reliability Engineer - Service Assurance Systems

Location
Greater London, England, United Kingdom
systems at scale. Familiarity with infrastructure-as-code tools such as Terraform or Ansible. Experience with log aggregation and analysis platforms such as the OTEL Stack or AWS CloudWatch Logs Insights. Exposure to Kubernetes or other container orchestration platforms. Experience working in an Agile or DevOps team environment. EEO Statement ...

Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President

Location
Greater London, England, United Kingdom
custom agentic loops) to coordinate multi-step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root-cause analysis, anomaly detection, and semantic log clustering. Self-healing Infrastructure Engineering: Experience designing closed-loop, self-healing ...

Site Reliability Engineering (SRE) / Observability Technical Lead

Hiring Organisation
NTT DATA
Location
London, United Kingdom
Salary
£ 80 K
DevOps roles, with leadership responsibilities.Proven expertise with Application Performance Monitoring (APM) tools such as New Relic, Datadog, AppDynamics, or Dynatrace.Hands-on experience with OpenTelemetry (OTel) for distributed tracing and observability instrumentation.Strong proficiency in Infrastructure as Code (IaC) using Terraform.Solid understanding of cloud platforms including AWS, GCP, or Azure.Experience with automation ...

Site Reliability Engineering (SRE) / Observability Technical Lead

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
roles, with leadership responsibilities. Proven expertise with Application Performance Monitoring (APM) tools such as New Relic, Datadog, AppDynamics, or Dynatrace. Hands-on experience with OpenTelemetry (OTel) for distributed tracing and observability instrumentation. Strong proficiency in Infrastructure as Code (IaC) using Terraform. Solid understanding of cloud platforms including ...

Database Platform Engineer

Location
Greater London, England, United Kingdom
services relevant to data platforms such as RDS, Aurora, S3, EC2 or EKS Familiarity with modern observability stacks such as Prometheus, Grafana, Elk or OTel Desirable: experience with cloud‐native and distributed SQL databases such as Aurora, YugabyteDB or TiDB Desirable: knowledge of data streaming and integration tools such ...

Senior Forward Deployment Engineer

Location
Greater London, England, United Kingdom
failover, and production-recovery exercises, highlighting skills in system reliability and continuity planning. - Experience with enterprise observability tools such as Splunk, ELK, Grafana, Prometheus, OpenTelemetry, AppDynamics, or Dynatrace, reflecting proficiency in monitoring and diagnostics. - Experience modernizing monolithic or legacy enterprise applications into maintain #J-18808-Ljbffr ...

Expert Forward Deployment Engineer

Hiring Organisation
Luxoft
Location
London, United Kingdom
Salary
£ 80 K
failover, and production-recovery exercises, highlighting skills in system reliability and continuity planning. Experience with enterprise observability tools such as Splunk, ELK, Grafana, Prometheus, OpenTelemetry, AppDynamics, or Dynatrace, reflecting proficiency in monitoring and diagnostics. Experience modernizing monolithic or legacy enterprise applications into maintain OtherLanguages English: C1 Advanced Seniority Senior London ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, United Kingdom
Salary
£ 70 K
usage, latency, cost, hallucination/quality metrics) using tools such as Arize, WhyLabs, or Langfuse.Logging & tracing: centralized logging (ELK/OpenSearch) and distributed tracing (OpenTelemetry) across data, training, and inference pipelines for end-to-end root-cause analysis.Integration — AI Stack, Enterprise Networks & Service Provider EnvironmentsPlatform integration: API-based and event ...

Enterprise Architect - AI

Location
Greater London, England, United Kingdom
latency, cost, hallucination/quality metrics) using tools such as Arize, WhyLabs, or Langfuse. Logging & tracing: centralized logging (ELK/OpenSearch) and distributed tracing (OpenTelemetry) across data, training, and inference pipelines for end-to-end root‐cause analysis. Integration — AI Stack, Enterprise Networks & Service Provider Environments Platform integration: API-based ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, UK
Employment Type
Full-time
latency, cost, hallucination/quality metrics) using tools such as Arize, WhyLabs, or Langfuse. Logging & tracing: centralized logging (ELK/OpenSearch) and distributed tracing (OpenTelemetry) across data, training, and inference pipelines for end-to-end root-cause analysis. Integration — AI Stack, Enterprise Networks & Service Provider EnvironmentsPlatform integration: API-based ...

Senior Platform Engineer

Hiring Organisation
9fin
Location
London, United Kingdom
Salary
£ 80 K
Good working knowledge of AWS services including ECS, EC2, Lambda, VPC, IAM, Route53, CloudFront, S3, RDSGood understanding of monitoring and logging solutions. We use OpenTelemetry, AWS Cloudwatch and SigNoz so experience with them is a bonus.Basic SRE knowledge, and experience in alerting and incident management platforms (eg. incident.io, Pagerduty)Proven ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
working knowledge of AWS services including ECS, EC2, Lambda, VPC, IAM, Route53, CloudFront, S3, RDS Good understanding of monitoring and logging solutions. We use OpenTelemetry, AWS Cloudwatch and SigNoz so experience with them is a bonus. Basic SRE knowledge, and experience in alerting and incident management platforms (eg. incident.io, Pagerduty ...

Devops SRE

Location
Greater London, England, United Kingdom
Performance Strong security mindset with a proven track record of designing secure, resilient cloud‐native systems. Experience implementing observability stacks including Prometheus , Dynatrace , and OpenTelemetry . Deep understanding of Linux internals , system performance tuning, and troubleshooting. Familiarity with Aqua Security for container runtime protection. CI/CD & Automation Tooling Hands ...

Senior Software Engineer (Infrastructure)

Location
Greater London, England, United Kingdom
Experience developing production‐ready infrastructure management tooling with either Python or Golang Familiarity with at least one of the following: Observability Tools (e.g. Prometheus, OpenTelemetry, Grafana) Databases (e.g. Postgres, DuckDB) Event Streaming platforms (e.g. Kafka) Container Orchestration (e.g. Docker, Kubernetes) Familiarity with cloud platforms such as AWS, Azure ...

Principal Software Engineer - Platform Engineering - Accelerator Business

Location
Greater London, England, United Kingdom
/CD, application resiliency, and secure delivery (e.g., SLSA framework and GitOps). Deep experience with Observability and Monitoring tools (e.g., Prometheus, Grafana, OTEL). Expertise in performance optimisation of distributed systems (e.g., caching, network latency). Practical experience with Service Mesh technologies (e.g., Istio, Linkerd, Cilium). Demonstrated success ...