101 to 125 of 258 OpenTelemetry Jobs in the UK

Senior Infrastructure Engineer I

Location
Greater London, England, United Kingdom
Golang. Deep understanding of container and orchestration internals (e.g., Kubernetes, Docker). Hands-on experience extending and integrating open-source tools like Prometheus, OpenTelemetry, and Grafana. Proven ability to package and scale developer-facing infrastructure utilities with a focus on improving system reliability. Desirable Experience building cloud-agnostic solutions running ...

Lead AI Software Engineer

Hiring Organisation
Wells Fargo
Location
London, UK
Employment Type
Full-time
Experience with RAG-based retrieval systems, vector search, embeddings, or enterprise knowledge retrieval patterns. Experience with agent observability and tracing tools such as LangSmith, OpenTelemetry, Arize, or similar platforms. Exposure to Splunk or enterprise observability frameworks. Experience with Docker, OpenShift, Kubernetes, or similar containerization and deployment platforms. Experience with Copilot ...

Database Reliability Engineer

Location
Manchester, England, United Kingdom
rigorous data integrity and mobility A Security & Observability Mindset: You believe security is paramount. You focus on building deep observability (Prometheus/Grafana/OpenTelemetry/Humio) and automated guardrails so the fleet is secure by design without requiring manual intervention Engineering via Code: While you are a systems expert ...

Sr. Observability Engineer – Kings Cross, London

Location
Greater London, England, United Kingdom
across our hybrid and cloud-native environments.* Innovate & Automate: Spearhead the evaluation, selection, and implementation of cutting-edge observability tools and platforms (e.g., Dynatrace, OpenTelemetry, Prometheus, Grafana). Architect and build robust, automated observability pipelines. Take an active part in documenting and defining processes and best practice.* Optimize & Analyze: Conduct … large-scale monitoring and observability solutions.* Expert-Level Tooling: Deep expertise with modern observability platforms (e.g., Dynatrace, AWS Cloudwatch, Prometheus, Grafana, ELK Stack, Splunk, OpenTelemetry).* Cloud & Infrastructure: Advanced knowledge of major cloud platforms (AWS, Azure, GCP), containerization (Docker, Kubernetes), and Infrastructure as Code (Terraform, Ansible).* Programming & Automation: Strong ...

Senior .NET Backend Developer

Hiring Organisation
Oscar Associates (UK) Limited
Location
York, North Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£70,000
Kubernetes. PostgreSQL, Redis or Elasticsearch. GraphQL. RabbitMQ or other messaging technologies. CI/CD pipelines and modern DevOps practices. Observability tooling such as Grafana, OpenTelemetry or Prometheus. Experience using AI-assisted development tools within the software development lifecycle. What's on Offer Hybrid working (2 days per week in York ...

Site Reliability Engineer

Location
United Kingdom
delivery lifecycles. An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. Proficiency in shell scripting for automation and system management. Experience with Infrastructure as Code, including Terraform and Ansible. Knowledge ...

Site Reliability Engineer

Location
Milton Keynes, England, United Kingdom
planning Proactively identify and prioritise reliability improvements Experience & Skills Required Hands‐on experience with Azure Monitoring (Application Insights, Alerts, Action Groups) Strong knowledge of OpenTelemetry (including Kubernetes) Scripting/automation using PowerShell and/or Azure CLI Experience with Terraform and GitHub Actions Ability to define SLIs/SLOs ...

Site Reliability Engineer, K8s (Remote International)

Hiring Organisation
PulsePoint
Location
United Kingdom, UK
Employment Type
Full-time
data infrastructureBare-metal and cloud environmentsGitOps and infrastructure automationModern observability and reliability engineering practicesTechnologies commonly used across the environment include Kubernetes, ArgoCD, Puppet, Terraform, OpenTelemetry, Prometheus, Alertmanager, Kafka, Redis and Ceph. Experience with every technology is not required. Who we're looking forSuccess in this role is not measured ...

ML / Backend Engineer @ Sqwish

Location
Cambridge, England, United Kingdom
PyO3 Experience with FastAPI, Pydantic, SQLAlchemy, Alembic, pytest, mypy, or Ruff Familiarity with Postgres, Redis, event-driven systems, queues, or streaming architectures Experience with OpenTelemetry, Prometheus, Grafana, Loki, Tempo, or structured logging Comfort with Docker, Kubernetes, Helm, Terraform, GitHub Actions, or release automation Exposure to LLM infrastructure, model routing, embeddings ...

Lead AI Engineer

Location
Greater London, England, United Kingdom
experience using frameworks like Autogen and LangGraph . Solid grounding in MLOps , containerisation (Docker, Kubernetes), and vector databases. Understanding of agent monitoring tools (Langfuse, OpenTelemetry). Strong software engineering best practices (testing, CI/CD, code reviews). Excellent communicator able to work with cross‐functional teams and clients. Desire ...

Python Backend Developer

Location
Greater London, England, United Kingdom
frontend work, and Go for select infrastructure Tools: RabbitMQ and Kafka for messaging, PostgreSQL and Redis for data storage Environment: Linux servers Observability: OpenTelemetry, Prometheus, Grafana and Zabbix Must-Haves: Strong background in software development, with strong experience with Python. A degree in Computer Science or a numerical subject from ...

Database Reliability Engineer

Location
Southampton, England, United Kingdom
challenge of multi‐tenant, multi‐region, multi‐cloud scenarios with rigorous data integrity. Security & Observability mindset: build deep observability (Prometheus/Grafana/OpenTelemetry/Humio) and guardrails for secure operation. Engineering via code: deliver backend services in Java with clean relational modeling and performant DDL. Interview Process Stage ...

Senior AI Engineer| London

Hiring Organisation
Infosys Technologies
Location
London, UK
Employment Type
Full-time
Docker, Kubernetes). Preferred Delivered AI projects within Agile frameworks Experience on Gen AI Feedback Analysis, topic modelling, sentiment analysis Knowledge of AgentOps and OpenTelemetry Understanding of Network Security Concepts, Network Telemetry and Analytics Understanding of Cloud computing and Virtualization Exposure to APM/Observability tools (Dynatrace, AppDynamics, Datadog, Splunk ...

Staff Software Engineer - International Pricing

Location
Greater London, England, United Kingdom
Confluent, IBM MQ/MQFTE and SFTP. Databases: MongoDB and SQL Server on Azure. API & Integration: Apigee, REST APIs and Windows Services. Observability: Dynatrace, OpenTelemetry, PagerDuty and Helix. What’s in it for you? Working at M&S means being part of something bigger – helping to deliver quality, value ...

Database Reliability Engineer

Hiring Organisation
Starling Bank
Location
London, UK
Employment Type
Full-time
ensuring rigorous data integrity and mobilityA Security & Observability Mindset: You believe security is paramount. You focus on building deep observability (Prometheus/Grafana/OpenTelemetry/Humio) and automated guardrails so the fleet is secure by design without requiring manual interventionEngineering via Code: While you are a systems expert, your ...

Senior Lead SRE: Reliability, Observability & Resiliency

Location
Auchentibber, Scotland, United Kingdom
members and stakeholders to define comprehensive service level indicators, service level objectives, and error budgets Designs, implements, and maintains operational reliability for large-scale OpenTelemetry pipelines on hybrid on-prem/cloud environments, supporting telemetry ingestion, processing, and export to backends such as InfluxDB, Prometheus, Elasticsearch, and OpenSearch Drives … assessment, refactoring, and incremental migration of custom legacy telemetry collection code to standardized OpenTelemetry instrumentation, reducing technical debt while maintaining system stability Actively contributes to the engineering community as an advocate of firmwide frameworks, tools, and practices, and influences peers and project decision-makers to consider the use and application ...

Senior Lead Site Reliability / DevOps Engineer

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
team members and stakeholders to define comprehensive service level indicators, service level objectives, and error budgetsDesigns, implements, and maintains operational reliability for large-scale OpenTelemetry pipelines on hybrid on-prem/cloud environments, supporting telemetry ingestion, processing, and export to backends such as InfluxDB, Prometheus, Elasticsearch, and OpenSearchDrives the assessment … refactoring, and incremental migration of custom legacy telemetry collection code to standardized OpenTelemetry instrumentation, reducing technical debt while maintaining system stabilityActively contributes to the engineering community as an advocate of firmwide frameworks, tools, and practices, and influences peers and project decision-makers to consider the use and application of leading ...

DevOps Engineer

Hiring Organisation
Infinity Quest
Location
Greater Edinburgh Area, United Kingdom
engineering improvements. - Monitoring & Observability - Implement enterprise monitoring, alerting, logging, and observability solutions. Utilize tools including: - Prometheus - Grafana - Google Cloud Monitoring - ELK/Elastic Stack - OpenTelemetry - Proactively identify issues before they impact customers and engineering teams. - Develop dashboards and operational reporting for platform health and reliability. - DevSecOps & Security - Embed security best … Management - Kong (desirable) - Strong understanding of: - REST APIs - OpenAPI Specifications - OAuth2 - JWT - API Security Standards - Observability & Monitoring Experience with: - Prometheus - Grafana - ELK Stack - OpenTelemetry - Google Cloud Monitoring - Scripting & Automation Proficiency in one or more of: - Python [optional] - Bash - Go - PowerShell - Networking & Security You'll thrive in this role ...

Sr. Manager, Site Reliability

Location
Manchester, England, United Kingdom
Engineering and Support. Select and stand up the primary observability platform, preferring extension of existing Omnicell contracts (DataDog, IBM/Instana, Prometheus/Grafana, OpenTelemetry, or other tooling already in use) over net‐new procurement. Define the instrumentation standards all new services must meet. Partner with the VP to migrate … Docker, Helm, and Service Mesh technologies (Istio, Linkerd). Hands‐on experience designing modern observability platforms using tools such as DataDog, Prometheus, Grafana, OpenTelemetry, Elasticsearch/Kibana, or equivalent — with an opinion about what a good telemetry stack looks like. Familiarity with integrating AI/ML‐based anomaly detection, alerting ...

Site Reliability Engineer

Location
West of England, England, United Kingdom
systems integration Version-controlled automation and operational tooling Experience with any of the following would be particularly useful: ServiceNow, Halo, Jira Service Management, OpenTelemetry, distributed tracing, Slack/Teams automation, datacentre or colocation environments, GPU infrastructure, DCIM, IPAM, virtualisation platforms or LLM-assisted operational automation. This ...

Site Reliability Engineer

Location
Manchester, England, United Kingdom
progressing the performance and availability of our critical systems. Using your engineering expertise, you will implement solutions that enhance reliability, including service instrumentation with OpenTelemetry and improved logging practices. You will leverage AI tools and LLM platforms in your daily work to reduce toil, drive autonomous operations, and optimise system ...

Software Engineer, SRE

Location
Manchester, England, United Kingdom
progressing the performance and availability of our critical systems. Using your engineering expertise, you will implement solutions that enhance reliability, including service instrumentation with OpenTelemetry and improved logging practices. You will leverage AI tools and LLM platforms in your daily work to reduce toil, drive autonomous operations, and optimise system ...

Python/AIML Developer - Software Engineer III

Location
Glasgow, Scotland, United Kingdom
accuracy, hallucination and toxicity metrics, guardrails and content moderation, red-teaming, and embedding evals into CI/CD. Experience with AI/ML observability (OpenTelemetry, Phoenix, or equivalent tracing) and with model governance, auditability, and control requirements in a regulated financial-services environment. Familiarity with financial risk domain concepts - market ...

Staff Software Engineer - AI Agents (Satori)

Location
Belfast City District, Northern Ireland, United Kingdom
TypeScript/Node a plus. Hands‐on work with agent frameworks and eval tooling. Track record of debugging distributed systems from traces; experience with OpenTelemetry’s GenAI semantic conventions a plus. Cloud‐native development, automating deployments, familiarity with container systems; AWS experience a plus. Excellent time management skills, comfortable coordinating ...

Software Engineer - AI Agents (Satori)

Location
Belfast City District, Northern Ireland, United Kingdom
level usage at enterprise scale. Development experience with Python; TypeScript/Node a plus. Track record of debugging distributed systems from traces; experience with OpenTelemetry's GenAI semantic conventions a plus. Cloud-native development, automating deployments, familiarity with container systems; AWS experience a plus. Excellent time management skills, comfortable coordinating ...