1,551 to 1,575 of 1,869 Observability Jobs

Senior Observability Engineer: Honeycomb & OpenTelemetry

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Group is seeking a Senior Observability Engineer in London to own and evolve the observability platform, primarily centered around Honeycomb and OpenTelemetry. You will instrument services in Java, Python, and C++, collaborate with SRE and development teams, and drive reliability through SLOs and incident response. This role requires … years in observability or related field, with UK working hours considered. #J-18808-Ljbffr ...

Senior Data Architect

Hiring Organisation
Hilti
Location
Allen, Texas, United States
Employment Type
Permanent
Salary
USD Annual
streaming patterns (e.g.,MSK/Kinesis,DMS,Glue). Partner with Solution/Platform Architects to ensuremicroservicesandevent drivendesigns are data efficient and observability ready. Data Governance, Quality & Observability Implementdata governance(policies, stewardship, data classification),cataloging(Glue Data Catalog),lineage(Open Lineage compatible), andquality(rules, thresholds, SLAs/SLOs). Embedmonitoring …/tokenization, RBAC/ABAC. AI/LLM enablement:Amazon Bedrock(Knowledge Bases, Guardrails, Agents), embeddings, chunking, retrieval,prompt design,token optimization, evaluation loops. Observability & FinOps: Cloud logs/metrics, lineage/quality SLAs,cost controls(storage/compute), workload rightsizing. DevOps/MLOps/IaC:Git,CI/CDfor ...

Observability Engineer (Dynatrace) — Telemetry & Performance

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Computacenter is seeking a Monitoring & Observability Engineer (Dynatrace) to design, implement and manage observability across customer IT estates in the UK. You will collect telemetry, diagnose issues and drive proactive improvements across teams. The role requires strong Dynatrace/Grafana/Splunk experience, scripting skills, cloud familiarity (Azure/ ...

Observability Architect

Hiring Organisation
Jobleads-UK
Location
Bristol, England, United Kingdom
Responsibilities Lead the assessment, design, and optimisation of the observability strategy for the co-location migration programme. Review the current observability architecture across infrastructure, networks, middleware, databases, and applications. Assess existing logging, metrics, distributed tracing, and monitoring capabilities to determine readiness for the co-location migration. Develop an observability strategy … alerting for infrastructure failures, application degradation, latency increases, replication issues, and capacity constraints. Support operational readiness activities including rehearsals and production cutover monitoring. Ensure observability solutions meet financial services regulatory requirements for auditability, log retention, security, and data governance. Validate access controls and security monitoring for observability platforms. Support evidence ...

SRE - Site Reliability Engineer - Observability & Performance

Hiring Organisation
Sanderson Recruitment
Location
Bristol, Somerset, United Kingdom
Employment Type
Contract
Contract Rate
GBP 550 - 600 Daily
Observability and Performance Up to £600 per day outside IR35 6 month initial contract Bristol - Largely remote I'm currently working with a client who is looking for an SRE to implement and enhance observability across Java applications, middleware and Linux infrastructure using Grafana click apply for full job details ...

Senior Observability & Performance Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Herbert Smith Freehills Kramer in London is seeking a Performance & Observability Engineer to drive reliability across the technology stack. You will move from basic monitoring to full observability, delivering real-time performance insights and proactive optimisation for complex legal applications. Working with SRE and DevOps teams, you will design … observability, instrument services with Grafana/Nexthink, implement SLIs/SLOs, and automate incident detection. #J-18808-Ljbffr ...

SRE Consultant

Hiring Organisation
Akkodis
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£90000 - £100000/annum
hold currently). The Role As a Site Reliability Engineer (SRE) you will lead site reliability engineering initiatives with a strong emphasis on observability, ensuring high performance and reliability of applications & infrastructure. Provide strategic insights to shape the overall SRE strategy while collaborating on the design and implementation of scalable … solutions. Establish effective monitoring, alerting and incident response strategies to maintain system availability and promote continuous improvement by collaborating with team members to deliver observability best practices and SRE methodologies. The Responsibilities Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to measure and maintain system ...

Dynatrace/Observability Engineer £528 per day INSIDE

Hiring Organisation
Hays
Location
Telford, Shropshire, West Midlands, United Kingdom
Employment Type
Contract
Contract Rate
Up to £528 per day + £528 per day INSIDE
Title: Dynatrace/Observability Engineer Rate: £528 per day INSIDE Clearance Required: SC Eligible Duration: 6 months Location: Telford - 2 days min per month Mandatory Certifications: Dynatrace Associate Certification As a Dynatrace/Observability Engineer, you will be responsible for designing, implementing, and supporting monitoring solutions across a range … insight, and proactive incident management. Key Responsibilities: Translate high-level monitoring and non-functional requirements (NFRs) into actionable configurations in Dynatrace. Deliver full-stack observability solutions, including application-aware network performance monitoring (NPM), synthetics, log analytics, and infrastructure metrics. Collaborate with architects and project teams to integrate monitoring into solution ...

DevOps Engineer

Hiring Organisation
Fruition Group
Location
Leeds, West Yorkshire, Yorkshire, United Kingdom
Employment Type
Contract
Contract: Inside IR35 We're seeking an experienced Senior DevOps Engineer to join a small, highly skilled engineering team delivering a large-scale enterprise observability platform as they move away from Splunk This is an opportunity to work on a critical cloud platform supporting the migration of numerous services onto … modern monitoring and logging solution. What you'll be doing * Support and enhance a large-scale observability platform. * Help engineering teams onboard and migrate their services. * Build and maintain dashboards, log pipelines and alerting. * Develop and manage cloud infrastructure using Terraform across Azure and AWS. * Produce technical documentation and operational ...

Site Reliability Engineer

Hiring Organisation
BC Forward
Location
Buffalo, New York, United States
Employment Type
Permanent
Salary
USD 1,100 Annual
Engineer to ensure the reliability, scalability, performance, and operational excellence of critical banking platforms and applications. The ideal candidate will have strong experience in observability, automation, incident management, Azure, and Infrastructure as Code and a proven ability to design, implement, and mature SRE practices across the SDLC while leading complex … reliability, availability, performance, and operational maturity through automation and engineering excellence. Define, implement, and monitor SLOs, SLIs, and error budgets for critical services. Develop observability strategies using Dynatrace, OpenTelemetry, distributed tracing, metrics, logs, dashboards, and alerting. Design and maintain end-to-end monitoring that provides actionable insights into application, infrastructure ...

Cloud Advisory Senior Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
security and cost outcomes: Embed zero trust IAM, security-by-design, HA/DR, and operational controls. Define SLOs/SLIs and bake in observability (metrics, logs, traces). Apply AIOps to reduce noise, accelerate incident triage, and improve reliability, and embed FinOps to manage performance and run-cost value …/DR, containers/orchestration, API management, and iPaaS, plus modern engineering patterns such as microservices, event‐driven architecture, and DDD. Operational excellence, observability and AIOps: Translate NFRs into pragmatic architecture decisions, define SLOs/SLIs, and design modern observability (metrics, logs, traces). Apply AIOps for alert reduction, anomaly ...

Lead/Senior Site Reliability Engineering

Hiring Organisation
Aceolution
Location
United Kingdom
alerts from multiple data sources and alerting workflows. ● Write libraries and APIs that give engineers self-service access to our monitoring, logging, and other observability systems. ● Use Terraform to deploy public and private cloud infrastructure. You are an ideal candidate if you: ● Have 5+ years’ experience designing, deploying and operating … quality and customer experience. ● Are curious, learn fast and feel comfortable diving into unfamiliar code and systems to solve problems. ● Understand the value of observability and can work with other teams to help them better monitor their services. ● Are willing to be part of a production on-call rotation. ● Have ...

Senior Software Engineer – Order Fulfilment

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
security. Champion engineering excellence through clean, maintainable, and well-tested code, promoting best practices across teams. Drive operational excellence by designing systems with observability, monitoring, resilience, and supportability at their core. Use tools such as Dynatrace to improve platform visibility, monitoring, and alerting capabilities. Collaborate closely with Product Managers, Engineering … enterprise services, APIs, and integrations across distributed systems. Passionate about software quality, maintainability, testing, and long-term sustainability. Experienced in CI/CD practices, observability, production support, and operational excellence. Comfortable designing systems with resilience, recoverability, monitoring, and supportability in mind. Pragmatic and commercially aware, capable of balancing technical ambition ...

Senior Enterprise Tools Engineer

Hiring Organisation
Jobleads-UK
Location
United Kingdom
serving as a senior escalation point, you will improve monitoring accuracy, reduce alert noise, validate automation workflows, and contribute to AIOps tuning and observability standards. You will help transition enterprise tool operations from reactive issues handling toward proactive, automation-driven reliability practices that improve uptime, user communication, and service maturity. … Responsibilities** **General Reliability Operations** * Monitor observability and AIOps platforms to detect anomalies, performance degradation, and emerging issues across enterprise systems.* Perform advanced incident triage and event correlation to identify root cause and reduce duplicate or misrouted incidents.* Lead or contribute to post-incident reviews, identifying systemic fixes and automation opportunities. ...

Senior Backend Platform Developer, Contact Center Systems

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
services for the company's contact center systems. This is a software engineering role first: where you will design APIs, serverless services, data flows, observability tooling, and AI-enabled workflows that support our customer and agent experience. The contact center is the domain, but the core need is strong development … integrations across internal systems, customer data, and third-party platforms. Technical patterns, implementation standards, and reusable services for the broader contact center systems roadmap. Observability pipelines using OpenSearch, CloudWatch, Kibana, Snowflake, and structured logging. Automation and AI-enabled workflows that improve routing, agent support, troubleshooting, and operational visibility. Migration tooling ...

Director of Software Engineering - Executive Director

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
drive impact, we've got an opportunity just for you. As a Director of Software Engineering at JPMorgan Chase within the Engineer's Observability Platforms team, you lead a technical area and drive impact within teams, technologies, and projects firm wide. Utilize your in-depth knowledge of software, applications, technical … driver of AIOps innovation and solution delivery. Job responsibilities Leads technology and process implementations to achieve functional technology objectives in the Observability Platforms space, providing essential services for Site Reliability Engineers, Operations and Engineers across the whole firm Delivers technical solutions that can be leveraged across multiple businesses and domains ...

DevOps Engineer TLNT1 NI

Hiring Organisation
Ocho
Location
Belfast, UK
engineering team. Key Responsibilities Build out infrastructure-as-code practices across an on-premise environment. Own and improve CI/CD pipelines. Build out observability: dashboards, alerting and automated checks (Grafana, InfluxDB). Work closely with development teams to support and streamline release processes. Partner with an existing DevOps engineer … Experience with Nomad or a comparable container orchestration tool (Kubernetes experience is transferable). Experience with HashiCorp Consul and Terraform. Experience with monitoring/observability tooling such as Grafana and InfluxDB. Solid grounding in CI/CD pipeline design and infrastructure-as-code. Comfortable operating in an on-premise ...

Senior Enterprise Tools Engineer

Hiring Organisation
Jobleads-UK
Location
United Kingdom
serving as a senior escalation point, you will improve monitoring accuracy, reduce alert noise, validate automation workflows, and contribute to AIOps tuning and observability standards. You will help transition enterprise tool operations from reactive issues handling toward proactive, automation-driven reliability practices that improve uptime, user communication, and service maturity. … Responsibilities** **General Reliability Operations** * Monitor observability and AIOps platforms to detect anomalies, performance degradation, and emerging issues across enterprise systems.* Perform advanced incident triage and event correlation to identify root cause and reduce duplicate or misrouted incidents.* Lead or contribute to post-incident reviews, identifying systemic fixes and automation opportunities. ...

AI Engineer

Hiring Organisation
Parkside
Location
London, United Kingdom
Employment Type
Permanent
Salary
£50000 - £90000/annum
regulated industry demands Implement retrieval systems from ingestion and chunking through to vector stores and retrieval optimisation Ship production-grade code with proper observability, error handling, testing and CI/CD Help design guardrails and failure handling so AI systems behave safely with real customers and real money involved … prompt engineering Exposure to agent frameworks (LangGraph, Claude Agent SDK, OpenAI SDK) or equivalent custom implementations An interest in LLM evaluation, debugging and observability Cloud platform experience (AWS, GCP or Azure) is a plus at junior level and expected at senior level The bar scales with the level. For senior ...

Principal Platform Engineer (SRE/Cloud)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
team enables engineering teams to build scalable and reliable services by supporting a foundation of tools and common services — including our managed Kubernetes clusters, observability tooling, and much more — while setting best practices around incident management and cost controls and empowering teams to build it, run it. The Principal Engineers … architectural vision that will deliver on Beamery's strategy Engage with other Principal Engineers in setting and advocating company-wide standards for operational excellence, observability, reliability and incident response Take a whole-company view of major incidents, identifying recurring themes and turning them into company-level investments Own and evolve ...

Lead Architect- Data & Database Systems

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
optimize queries for performance and correctness. Monitor and tune database performance, including query tuning, indexing, partitioning, and caching. Build and maintain monitoring, alerting, and observability for database systems. Support production incidents and perform root cause analysis; participate in on-call rotations. Assist application teams with data access patterns, migrations … MySQL/MariaDB, Microsoft SQL Server, or Oracle. Performance tuning experience, including indexing strategies, query profiling, and execution plan interpretation. Experience with monitoring and observability tooling such as Prometheus, Grafana, Datadog, Dynatrace or New Relic. Familiarity with versioned schema migrations using tools such as Flyway, Liquibase, Alembic, or sqitch. Good ...

Software Engineer, GPU Infrastructure- ChatGPT Engineering

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
large-scale GPU infrastructure supporting ChatGPT inference. Build internal platforms, tooling, and AI-powered agents that automate fleet operations and reduce operational overhead. Improve observability, reliability, and operational efficiency across thousands of GPUs. Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response. Identify infrastructure bottlenecks and implement … software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure. Understand infrastructure observability, monitoring, capacity planning, and incident management. Enjoy identifying cross-team pain points and building reusable platforms that improve developer productivity. Are comfortable working across software ...

SRE Technical Lead

Hiring Organisation
83zero Limited
Location
Wokingham, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
senior technical escalation point for major incidents and high-risk releases. Lead blameless post-incident reviews and ensure measurable service improvements. Define and oversee observability, monitoring and capacity management practices. Ensure SRE approaches align with security, governance and compliance requirements. Mentor and coach senior engineers, helping to improve SRE maturity … OpenShift. Experience designing and supporting hybrid and multi-cloud platforms. Experience with service mesh technologies such as Istio. Strong hands-on experience with observability tooling including Prometheus, Grafana, Loki, Tempo and OpenTelemetry. Infrastructure as Code and GitOps expertise using tools such as Helm, Kustomize, ArgoCD and Tekton. Experience building ...

Senior Software Engineer - Permanent - London/Hybrid - £70,000 - 85,000

Hiring Organisation
Robson Bale Ltd
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP 70,000 - 85,000 Annual
Datadog, Kibana and Heap to investigate issues, understand customer impact and improve service health. Partner with Software Engineering, SRE and Platform teams to improve observability, reliability and operational readiness. Contribute to service onboarding, post-incident reviews and continuous improvement initiatives. What We're Looking For Experience in Application Operations, Production … Back End language. Practical Kubernetes knowledge, including deployments, logging, monitoring and safe rollbacks. MySQL operational experience, including performance investigation and query analysis. Experience with observability platforms such as Datadog and log analysis tools such as Kibana. Strong communication skills and a passion for automation, continuous improvement and engineering quality. Nice ...

SR. GOOGLE CLOUD PLATFORM ENGINEER

Hiring Organisation
Widenet Consulting
Location
Renton, Washington, United States
Employment Type
Permanent
Salary
USD 9,000 Hourly
Workspace, and other cloud systems Automate role-based access, group management, and policy enforcement Ensure consistent identity propagation across: Infrastructure CI/CD systems Observability platforms Reduce reliance on manual access management and exception handling Automate Platform Workflows & Guardrails Implement Infrastructure-as-Code (IaC) and automation for GCP and Workspace … automation coverage Integrate Cross-Platform Capabilities Ensure GCP and Workspace integrate seamlessly with: CI/CD platforms (e.g., GitLab) Infrastructure tools (e.g., Terraform Cloud) Observability platforms (e.g., Datadog) Enable consistent patterns across multi-cloud and multi-tool environments Drive Platform Adoption & Consistency Partner with engineering teams to identify friction ...