1,551 to 1,575 of 1,915 Observability Jobs

Senior Backend Engineer - Databases Pyroscope | UK | Remote

Hiring Organisation
Jobleads-UK
Location
United Kingdom
United Kingdom (Remote) Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations … cost that makes sense. Turn Pyroscope into a platform capability inside Grafana: bi-directional trace-to-profile correlation, integration with Kubernetes Monitoring and App Observability, and profiles surfaced where engineers already start their investigations. Prepare Pyroscope for an agent-driven world: APIs, CLI, and docs designed so AI agents ...

Senior AI Enablement Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
. You will help build the paved roads that allow engineers to use AI tools safely and effectively: standards, reusable workflows, automation, evaluation approaches, observability, documentation, and internal platform capabilities. The goal is to make AI‐assisted engineering reliable, measurable, secure, and aligned with enterprise software delivery standards. You will … testing, code review, onboarding, and knowledge retrieval. Contribute to AI governance implementation by helping translate policy and security expectations into usable engineering workflows. Support observability and measurement for AI adoption, including usage insights, effectiveness, quality signals, and operational risks. Partner with platform, architecture, AppSec, and infrastructure teams to ensure ...

Software Engineer - Synthetic Monitoring | UK | Remote

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … Monitoring as code and with confidence Break down complex, ambiguous problems into incremental deliverables and iterate quickly based on feedback Ensure quality through testing, observability of your own systems, and documentation : our checks are something customers alert on, so reliability is a feature Be a part of the team ...

Agentic AI Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
AIOps. The role will focus on designing, building, deploying and operating enterprise AI agents and multi-agent workflows on AWS, with strong emphasis on observability, reliability, cost control, security and continuous optimization in production environments. Required Skills Agentic AI, Engineering, AWS Platform, AIOps/LLMOps, Dev Key Responsibilities AI Agent … DynamoDB and SQS/SNS. Create reusable libraries, patterns and accelerators to standardize AI agent development across teams. AIOps, Production Monitoring & Operations Establish monitoring, observability and operational governance for production AI workloads. Track agent performance, model latency, cost, prompt effectiveness, error rates and quality signals. Define alerting, incident response ...

Senior Backend Software Engineer (AI Infrastructure / Artifact Management) – Developer Services

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
automated remediation, advancing AI from assisted analysis toward closed-loop resolution; Architect and optimize large-scale distributed systems to continuously improve throughput, reliability, observability, and developer experience. Work closely with engineering, security, infrastructure, and AI business teams to understand real-world requirements and drive complex projects from design to adoption … solid engineering skills; Familiarity with distributed systems and hands‐on experience in one or more areas such as storage, computing, task scheduling, caching, messaging, observability, or reliability engineering; Strong problem‐solving, system design, and execution skills, with the ability to diagnose issues across complex system paths and drive long‐term ...

SRE Technical Lead

Hiring Organisation
Adecco
Location
Reading, Berkshire, United Kingdom
Employment Type
Permanent
Salary
GBP 70,000 - 90,000 Annual
remediation Act as the technical escalation point for major incidents and high-risk releases Lead blameless post-incident reviews and ensure continuous improvement Establish observability and capacity management practices using modern tooling Identify and eliminate systemic reliability risks and operational inefficiencies Collaborate with engineering, platform, security, and operations teams across … Experience working in multi-cloud or hybrid cloud environments Strong understanding of SRE principles (SLOs, SLAs, error budgets, reliability engineering) Hands-on experience with observability tooling (eg, Prometheus, Grafana, OpenTelemetry, Loki, Tempo) Strong knowledge of Infrastructure as Code and GitOps (eg, Helm, Kustomize, ArgoCD, Tekton) Experience with CI/ ...

Senior Observability Solution Architect – Pre-Sales

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
leading observability platform in Greater London is seeking an experienced Solution Engineer to join their team. This role involves collaborating with account executives on technical sales cycles, delivering impactful presentations, and overseeing technical aspects of the process. The ideal candidate will have a minimum of 5 years in a customer ...

Staff Analytics Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
that improve performance, developer experience, cost efficiency, or operational maturity. Owning and evolving core platform components, including CI/CD, testing strategies, environment management, observability, and infrastructure as code. Acting as the technical escalation point for complex, cross‐cutting platform issues and guiding teams toward robust, scalable solutions. Driving Snowflake … performance and cost optimisation, informed by real workloads and modelling patterns. Implementing and maturing data SLAs/SLOs, data observability, lineage, and quality frameworks to ensure trusted analytics at scale. Collaborating with data product and engineering teams to enable safe, scalable ingestion and well‐defined data contracts. Influencing how teams ...

Senior Cloud Engineer - Contract

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Engineering builds and runs the new Azure cloud that everything else at Flagstone depends on. It's the foundation for our security tooling, our observability, and the hosting for our AI. The team works in infrastructure-as-code (Terraform and Bicep), owns the landing zones and hub-and-spoke networking … roll out golden paths that cut delivery cost and speed up engineering squads. Stand up our AI platform foundations: an AI gateway, an observability stack, and hosting for AI tools including Flagstone Concierge, with model access through AWS Bedrock. Build and maintain infrastructure pipelines with security scanning, plan validation ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You’ll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you’ll find meaningful impact and growth here. … large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelines Implement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloads Tune GPU and accelerator capacity, autoscaling, and cost ...

Senior DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
platform engineering and operational backbone of our solutions. Responsibilities Define and maintain Auriga’s DevOps and platform engineering roadmap across infrastructure, CI/CD, observability, reliability and security. Design scalable platform architectures across cloud, on-premises, hybrid and customer-site environments. Establish reusable standards, deployment patterns and reference architectures across … deployments across development, test, staging, production and customer environments. Design and support customer-isolated, multi-site and multi-tenant deployment models where required. Implement observability, monitoring, alerting, SLOs and operational health practices for production and customer-facing systems. Lead incident response, post-incident reviews and continuous improvement of platform reliability ...

Global DevOps Lead

Hiring Organisation
Stott & May Professional Search Limited
Location
United Kingdom
Employment Type
Permanent, Work From Home
Salary
£95,000
with engineering, cloud, and operations teams to deliver a modern, automated, and scalable platform. You'll drive DevOps strategy across infrastructure, CI/CD, observability, SRE, and cloud optimisation while influencing senior stakeholders across the business. Key Responsibilities - Define and implement a global DevOps operating model, including governance, standards … initiatives. - Partner with engineering and cloud teams to establish clear ownership across DevOps and infrastructure. - Lead the implementation and optimisation of enterprise monitoring and observability using Datadog. - Build scalable deployment pipelines that improve release quality and speed. - Establish and monitor DORA metrics, driving improvements in deployment frequency, lead time, change ...

Site Reliability Engineer (SRE) - Cloud & Automation

Hiring Organisation
Spencer Rose Ltd
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP 60,000 - 70,000 Annual
implementation of SRE practices across the organisation, working closely with infrastructure teams to optimise deployment processes and embed automation and operational excellence. Enhance observability and reliability , defining and implementing SLAs, SLOs and SLIs to improve alerting, monitoring, and capacity planning. Identify and eliminate toil , developing frameworks to analyse recurring issues … beneficial). Experience supporting and building multi-environment, multi-region cloud platforms (AWS or GCP), using IaC and GitOps workflows. Hands-on experience with observability/APM tooling such as Grafana, Datadog or Dynatrace. Background working in regulated financial services or banking environments. Excellent troubleshooting, analytical and communication skills, able ...

Vice President, DevOps Production Services

Hiring Organisation
Jobleads-UK
Location
Manchester, England, United Kingdom
enterprise applications and ensure platform stability, resiliency, and availability. Monitor application health, system performance, batch jobs, interfaces, and alerts using enterprise monitoring and observability tools. Investigate, troubleshoot, and resolve production incidents within defined SLAs. Perform root cause analysis (RCA) for recurring issues and drive permanent fixes. Analyze production logs, identify … Cloud experience preferred. Knowledge of automation/scripting using Python, Shell, or PowerShell. Exposure to DevOps/SRE practices, CI/CD pipelines, and observability tooling. Strong communication skills with the ability to provide concise incident and executive status updates. #J-18808-Ljbffr ...

Full Stack Engineer (Contract) – Leeds

Hiring Organisation
Jobleads-UK
Location
Leeds, England, United Kingdom
secure, scalable and maintainable applications Create automated unit and integration tests Contribute to CI/CD pipelines and continuous delivery Implement logging, monitoring and observability best practices Support production issues and continuous improvement initiatives Participate in peer reviews and Agile ceremonies Produce and maintain technical documentation The Team … testing Strong troubleshooting and problem-solving skills Excellent communication and collaborative approach Nice to Have Cloud-native development experience Contract testing and UI automation Observability (logging, metrics and tracing) Financial Services experience JIRA or similar Agile tooling To Be Considered... Please either apply by clicking online or emailing your ...

Azure DevOps Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
modules aligned to business and security requirements Implement and enforce Azure security controls and policies Lead platform security initiatives across Azure environments Contribute to observability, alerting, and Site Reliability Engineering practices Collaborate with engineering teams to deliver resilient and scalable solutions Security & Platform Focus Areas Implement perimeter security using Azure … Terraform Kubernetes certification and experience with AKS Deep understanding of DevOps and platform engineering principles Strong knowledge of cloud security best practices Experience with observability, monitoring, and SRE concepts Why Apply Fully remote within the UK Work with a mission‐driven, highly respected health tech organisation Modern cloud environment with ...

Senior Backend Java Developer (Java/Spring Boot)

Hiring Organisation
HTC Global Services Inc
Location
Dearborn, Michigan, United States
Employment Type
Permanent
Salary
USD Annual
infrastructure and deployment processes to enhance resiliency and reliability. Support application security practices, including data protection through encryption and anonymization. Troubleshoot production issues using observability and debugging tools. Required Qualifications Bachelor's degree. 6+ years of overall IT experience. 4+ years of software development experience. 5+ years of experience with … backend services. Experience building or maintaining test frameworks and testing tools. Strong understanding of software quality practices and continuous improvement processes. Knowledge of observability, debugging, and production operations. Experience with Docker, CI/CD pipelines, and cloud computing concepts. Experience with databases such as PostgreSQL, MySQL, or MongoDB. ...

Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Watford, England, United Kingdom
high-demand events. What you’ll be doing Objectives of the role Maintain reliable production services across digital platforms Improve monitoring, alerting, and observability coverage Reduce operational toil through automation Support incident response and continuous improvement Contribute to performance and scaling of services Production operations Participate … Incident response & improvement Support incident triage and resolution Participate in post-incident reviews and implement remediation actions Maintain and improve runbooks and operational documentation Observability Implement and maintain monitoring using: + Splunk + CloudWatch + Grafana Improve: + Logging quality + Metrics coverage + Alerting accuracy Contribute to linking system ...

Senior Platform Engineer - Developer Experience

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
paved roads that other engineers use. You will work across the software development lifecycle, from creating a new service through to testing, deployment, observability and operating it in production. You will join an established Platform team and work alongside our existing Developer Experience Engineer. You will speak directly with engineers … reliability and usability of our CI/CD systems. Developing reusable platform capabilities that product engineers can consume through self-service. Helping engineers use observability effectively, with good defaults for logs, metrics, traces and service‐level indicators. Working directly with engineers to understand friction, test ideas and support adoption. Using ...

Network Automation Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Jinja2 Integrating network automation into CI/CD pipelines for reliable, repeatable deployments Creating APIs and self‐service tooling for engineering teams Implementing observability and telemetry solutions for performance and reliability Partnering with network, platform and security teams to deliver resilient, scalable systems Contributing to incident response and production reliability … Ansible, Terraform and Jinja2; also must have experience leveraging AI tools, such as Claude Code Familiarity with Docker and Kubernetes Exposure to monitoring, observability or telemetry in distributed systems Pragmatic problem solver who can operate in ambiguity and take ownership Comfortable working in collaborative, fast‐paced engineering teams Deep understanding ...

Lead Software Engineer - Java / Python - Equity Derivatives - Front Office Quant Developer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
test strategy, defect triage, evidence collection, and sign-offs Maintain strong release discipline, including regression assessment, rollback/fallback planning, post-deployment verification, and observability improvements Gather and synthesize data/telemetry to develop reporting and metrics that improve stability, quality, and delivery predictability Identify hidden failure patterns in production … Agile environment, managing multiple priorities/projects, and producing clear documentation Operational excellence mindset, including incident leadership, postmortems, and continuous improvement in observability and production reliability Knowledge of equity derivatives products and workflows, including options pricing concepts and understanding of how system issues can impact quoting, booking, and hedging Preferred ...

Vice President Software Engineering

Hiring Organisation
Jobleads-UK
Location
City of Edinburgh, Scotland, United Kingdom
where 80–90% of code is AI‐generated, with a roadmap to 95%+. Embed modern engineering excellence (CI/CD, trunk‐based development, observability, and automated testing). Partner cross‐functionally across Product, DevOps, Security, and Platform teams. Build a high‐performance culture grounded in accountability, innovation, and continuous … where software is shipped to production frequently or daily. Expertise in modern practices including CI/CD pipelines, trunk‐based development, automated testing strategies, observability and system reliability. Proven ability to use engineering metrics to drive performance and continuous improvement. Organisational Design & Methodologies Experience designing and evolving engineering organisations using ...

Software Engineering Manager - Tooling and Optimisations

Hiring Organisation
Jobleads-UK
Location
Windsor, England, United Kingdom
practice, reduce duplication, and support maintainable, secure and high-performing systems. Improve delivery capability through platform reliability and DevOps maturity Continuously strengthen deployment pipelines, observability, alerting, incident response, recovery procedures and operational readiness across Field Ops engineering teams. Manage stakeholders and maintain clear communication Build trusted relationships across product, operations … data modelling and data quality controls. Ability to produce both high‐level and detailed design specifications. Experience leading DevOps practices, including CI/CD, observability, monitoring and incident management. Demonstrated capability leading multi‐squad engineering delivery in a product‐led organisation. Mindset & Ways of Working Comfortable working in iterative, outcome ...

Backend Engineer - Platform - Stacks | UK | Remote

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand ...

Team Lead, Integrations

Hiring Organisation
Genesis10
Location
Irving, Texas, United States
Employment Type
Permanent
Salary
USD Annual
Code & DevOps: IaC with Terraform; able to review and approve infrastructure changes Azure DevOps pipeline design, branch policies, and mandatory-review gate configuration Observability & Monitoring: Application Monitoring using Application Insights Log Management and Analytics Platform Monitoring using Azure Monitor Query & Analysis using KQL (Kusto Query Language) Alerting, dashboards, and validating … solution observability Application Development: Minimum 8 years of solid .NET coding experience (C#, ASP.NET Core, background/worker services) Expert knowledge of common design and patterns including .NET patterns, libraries and Azure services Documentation & Communication: Strong technical writing skills Visual Modeling with Lucid chart/Azure architecture diagramming Comfortable delivering ...