26 to 50 of 74 Remote/Hybrid OpenTelemetry Jobs

Senior Developer Experience Engineer

Hiring Organisation
9fin
Location
London, United Kingdom
Salary
£ 80 K
infrastructure as code. We currently use AWS and Terraform. Building internal developer platforms, developer portals, CLIs or service templates. Observability using tools such as OpenTelemetry, CloudWatch or SigNoz. Reliability engineering, incident response or production operations. Production support or on-call experience for your own team's services. Experience working closely ...

SRE Observability Technical Lead - Vice President

Hiring Organisation
Citigroup
Location
London, United Kingdom
Salary
£ 80 K
solutions to improve the service reliability and/or increase productivity and efficiencyHands-on experience in observability tools and stacks such as Grafana, Prometheus, OpenTelemetry, ELK, Splunk, and similar platforms.Deep understanding of SLIs, SLOs, Error Budgets, and telemetry best practices in high-availability environments.Proven ability to troubleshoot integration issues ...

Lead DevSecOps Engineer

Location
Greater London, England, United Kingdom
platforms (e.g., EC2 to EKS, or cross-cloud) with a focus on data integrity and minimal downtime Ability to implement standardized telemetry pipelines (e.g., OpenTelemetry, Prometheus, or ELK) that provide developers with out-of-the-box visibility into their services Familiarity with automated policy enforcement and compliance-as-code (e.g. ...

typescript developer for authentication platforms

Location
Greater London, England, United Kingdom
GitHub Actions, and writing unit and integration tests with Jest Familiarity with security principles including IAM, encryption and networking, alongside observability tools such as OpenTelemetry, Honeycomb or Grafana Nice to have: Exposure to identity or MFA platforms such as Auth0 or Transmit Security, knowledge of microservices architecture, API gateways such ...

Site Reliability Engineer

Location
United Kingdom
delivery lifecycles. An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. Proficiency in shell scripting for automation and system management. Experience with Infrastructure as Code, including Terraform and Ansible. Knowledge ...

Lead AI Engineer

Location
Greater London, England, United Kingdom
experience using frameworks like Autogen and LangGraph . Solid grounding in MLOps , containerisation (Docker, Kubernetes), and vector databases. Understanding of agent monitoring tools (Langfuse, OpenTelemetry). Strong software engineering best practices (testing, CI/CD, code reviews). Excellent communicator able to work with cross‐functional teams and clients. Desire ...

Senior Associate Engineer (Digital Product Team)

Hiring Organisation
Canada Life
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
regularly to production. An appreciation of observability, using metrics, logs, and traces (for example Azure Monitor/Application Insights, Datadog, Grafana/Prometheus, or OpenTelemetry) to understand system health, and an interest in SLIs/SLOs and alerting. Comfort with, or enthusiasm to learn, Infrastructure as Code to provision ...

SRE Observability Technical Lead - Vice President

Location
Belfast City District, Northern Ireland, United Kingdom
Observability Engineering, or platform infrastructure roles focused on operational telemetry. Hands‐on experience in observability tools and stacks such as Grafana, Prometheus, OpenTelemetry, ELK, Splunk, and similar platforms. Deep understanding of SLIs, SLOs, Error Budgets, and telemetry best practices in high‐availability environments. Proven ability to troubleshoot integration issues ...

SRE Observability Technical Lead - Vice President

Hiring Organisation
Citigroup
Location
Belfast, Down, United Kingdom
Salary
£ 60 K
Experience in SRE, Observability Engineering, or platform infrastructure roles focused on operational telemetry.Hands-on experience in observability tools and stacks such as Grafana, Prometheus, OpenTelemetry, ELK, Splunk, and similar platforms.Deep understanding of SLIs, SLOs, Error Budgets, and telemetry best practices in high-availability environments.Proven ability to troubleshoot integration issues ...

Senior Software Developer (Python)

Location
Greater London, England, United Kingdom
Location: London, London. Hours: 37.5 hrs/week. Pay Rate: Unknown. Better content. Better products. And better careers. Working in Tech, Product or Data at the company is about building the next and the new. ...

Site Reliability Engineer

Location
West of England, England, United Kingdom
systems integration Version-controlled automation and operational tooling Experience with any of the following would be particularly useful: ServiceNow, Halo, Jira Service Management, OpenTelemetry, distributed tracing, Slack/Teams automation, datacentre or colocation environments, GPU infrastructure, DCIM, IPAM, virtualisation platforms or LLM-assisted operational automation. This ...

Site Reliability Engineer, K8s (Remote International)

Hiring Organisation
PulsePoint
Location
United Kingdom
Salary
£ 60 K
data infrastructureBare-metal and cloud environmentsGitOps and infrastructure automationModern observability and reliability engineering practicesTechnologies commonly used across the environment include Kubernetes, ArgoCD, Puppet, Terraform, OpenTelemetry, Prometheus, Alertmanager, Kafka, Redis and Ceph.Experience with every technology is not required.Who we’re looking forSuccess in this role is not measured by the number ...

Python Backend Developer

Location
Greater London, England, United Kingdom
frontend work, and Go for select infrastructure Tools: RabbitMQ and Kafka for messaging, PostgreSQL and Redis for data storage Environment: Linux servers Observability: OpenTelemetry, Prometheus, Grafana and Zabbix Must-Haves: Strong background in software development, with strong experience with Python. A degree in Computer Science or a numerical subject from ...

Senior Developer Experience Engineer

Location
Greater London, England, United Kingdom
infrastructure as code. We currently use AWS and Terraform. Building internal developer platforms, developer portals, CLIs or service templates. Observability using tools such as OpenTelemetry, CloudWatch or SigNoz. Reliability engineering, incident response or production operations. Experience in fintech or another regulated environment. Benefits We’re a scaling start ...

Database Reliability Engineer

Hiring Organisation
Starling Bank
Location
London, United Kingdom
Salary
£ 80 K
ensuring rigorous data integrity and mobilityA Security & Observability Mindset: You believe security is paramount. You focus on building deep observability (Prometheus/Grafana/OpenTelemetry/Humio) and automated guardrails so the fleet is secure by design without requiring manual interventionEngineering via Code: While you are a systems expert, your ...

AI Harness Engineers

Hiring Organisation
Capgemini
Location
Greater London, United Kingdom
Employment Type
Full Time
data platform, Neo4j Enterprise as the semantic knowledge graph, an agent memory plane serving episodic and precedent memory over MCP, MCP-native connectors, OpenTelemetry and Grafana for observability, all on CNCF-conformant Kubernetes with Helm and Argo CD, deployable to any hyperscaler or on-prem. A tool-for-tool match ...

Hybrid Lead Observability Engineer | Platform & SRE

Location
United Kingdom
seeking a Lead Observability Engineer/Senior Software Engineer in Nottingham, hybrid role. You will monitor and enhance our observability estate, designing platforms with OTel, Grafana and Prometheus while collaborating across teams. You'll work with Golang/Java, AWS services (Fargate/Lambda), and Kubernetes, driving best practices, code ...

Platform Engineer

Hiring Organisation
itecopeople
Location
London, United Kingdom
Employment Type
Permanent
Salary
£54000 - £65000/annum
Code Develop and maintain GitOps CI/CD pipelines Manage Kubernetes networking, service mesh and gateway technologies Improve platform observability using Grafana, Prometheus and OpenTelemetry Maintain platform security, resilience and automation Troubleshoot production platform issues and drive continuous improvement Work closely with architects to turn high-level designs into robust … platform engineering within production environments Terraform and Infrastructure as Code Docker, GitOps and CI/CD pipelines Linux and Bash scripting Grafana, Prometheus and OpenTelemetry Kubernetes networking and service mesh technologies Production platform operations, troubleshooting and automation You'll also be able to demonstrate: Experience owning technical implementation decisions rather ...

Lead Observability Engineer - Hybrid Cloud Telemetry

Location
United Kingdom
will design and operate the observability estate across our UK platforms, mentoring peers and shaping architecture. The role emphasizes building secure, scalable systems with OTel, Grafana, Prometheus, Golang, Java, AWS, Fargate, and Kubernetes. You will collaborate across teams in an outcome-based culture with flexible work-from-home days. #J ...

Staff Software Engineer – Identity and Access, Identity Squads | UK | Remote

Hiring Organisation
Grafana Labs
Location
United Kingdom
Salary
£ 70 K
working with OpenFGA and KiFederate IDP.Experience working in security critical environments.Experience contributing to or maintaining Open Source projects.Familiarity with observability tooling (e.g., Grafana, Prometheus, OpenTelemetry).Compensation & Rewards:In the UK, the Base compensation range for this role is 100,000 - 124,000. Actual compensation may vary based on level, experience ...

SRE Managing Consultant - Cloud Operating Model

Location
Manchester, England, United Kingdom
/SLOs, incident management, observability, and continuous improvement across cloud and hybrid platforms.* Exposure to modern observability tooling and ecosystems (e.g. Datadog, Dynatrace, Prometheus, OpenTelemetry, Loki), with a strong understanding of how metrics, logs, and traces are applied to inform reliability strategy, incident management, and operational decision‐making.## **Security Check ...

SRE Managing Consultant - Cloud Operating Model

Location
Greater London, England, United Kingdom
/SLOs, incident management, observability, and continuous improvement across cloud and hybrid platforms.* Exposure to modern observability tooling and ecosystems (e.g. Datadog, Dynatrace, Prometheus, OpenTelemetry, Loki), with a strong understanding of how metrics, logs, and traces are applied to inform reliability strategy, incident management, and operational decision‐making.## **Security Check ...

Senior Site Reliability Engineer

Hiring Organisation
GCS
Location
Glasgow, City of Glasgow, United Kingdom
Employment Type
Permanent
Salary
£75000 - £95000/annum Bonus
Role Title: Senior Site Reliability Engineer Location: Knutsford or Glasgow - Hybrid (2 days per week onsite) Role Category: Permanent Overview: We're recruiting for an experienced Senior Site Reliability Engineer to drive reliability, scalability and ...

Front End Angular Developer

Hiring Organisation
Vanguard
Location
Manchester, Greater Manchester, United Kingdom
Salary
£ 70 K
Boot, JavaS3 Buckets, lambda, ECSKongCypress, TypeScriptTechnology we use: • Front-End: Angular, JavaScript, TypeScript, AEM (Hybrid), GraphQL, SSR, NgRX, Jest, SCSS• Backend: Spring Boot, Java, OpenTelemetry, HoneyComb, Liquibase, REST API's.• Cloud/AWS: S3 Buckets, lambda, ECS, SQS/SNS, CloudFront, IAM, Secrets Manager, CloudFormation• Networking: AKAMAI, NGINX, Kong, Okta ...

Site Reliability Engineer- Spacetime UK

Location
Greater London, England, United Kingdom
mature our observability stack, moving from cloud-native tools to a robust, scalable, and insightful platform built on best-in-class technologies (Prometheus, OpenTelemetry, etc.). If you are an SRE who thrives on platform-building challenges and wants to be relied upon to build a production-grade observability stack … build Aalyria's centralized observability platform, integrating and scaling tools for metrics (e.g. Prometheus), logging (e.g. Loki), and distributed tracing (e.g. Tempo/OpenTelemetry). Define, implement, and manage a robust framework of Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for our core products, ensuring ...