1,951 to 1,975 of 4,483 Observability Jobs

Data Reliability Engineer

Hiring Organisation
Ashdown Group
Location
London, UK
Employment Type
Full-time
work from home 2 days per week. This is a high-impact role focused on improving data quality, reducing incidents, and building scalable observability across a modern enterprise data platform. You'll help ensure data across the organisation is accurate, reliable, and trusted for critical business decision-making. … style roles, with strong SQL and Python skills and experience working in modern cloud-based data environments. Hands-on experience with data observability tools such as Grafana, Monte Carlo, or Acceldata, and data governance/quality platforms like Informatica, Collibra or Microsoft Purview is highly desirable. Experience within the Azure ...

Senior Software Engineer – Agentic Development Enablement

Location
Greater London, England, United Kingdom
such as Claude Code and GitHub Copilot Design and implement guardrails, controls, and engineering patterns for AI-assisted development Contribute to endpoint and platform observability, telemetry, and policy enforcement Define how controls work consistently across local development environments and CI/CD pipelines Explore changes to development environments, including containerised … production Requirements Strong hands-on background as a software engineer Broad technical understanding across developer tooling, cloud platforms, operating systems, desktop environments, security controls, observability, telemetry, and CI/CD Experience working on developer workflows and engineering ways of working, not only end-user application delivery Ability to work ...

Site Reliability Engineer with Python

Hiring Organisation
BC Forward
Location
Pennington, New Jersey, United States
Employment Type
Permanent
Salary
USD Hourly
Global Markets. The ideal candidate will have strong experience in Python, Django, REST APIs, MySQL, Linux, Infrastructure-as-Code, CI/CD, and observability and a proven ability to increase platform reliability, improve observability, automate operations, and lead incident response at scale. Responsibilities: Own reliability and operational health of core … Build, enhance, and maintain internal tools that support platform operations. Lead incident response, root-cause analysis, and post-incident improvements. Improve application and platform observability using Dynatrace and OpenTelemetry. Partner with Quartz core teams on platform modernization and cloud migrations. Harden production systems and evolve support tooling for large-scale ...

Front End developer

Location
Greater London, England, United Kingdom
applications. You’ll work across Python services and React/Vite frontends , collaborating closely with engineering teams to improve application reliability, CI/CD, observability, security and developer tooling. Key experience: Docker and CI/CD pipelines Azure cloud experience is essential; AWS is a bonus Good knowledge of Linux … networking, security and observability Experience supporting scalable application platforms and cloud deployments Able to work independently and establish reusable engineering patterns and best practices A great opportunity for someone who enjoys working across development, cloud and platform engineering. #J-18808-Ljbffr ...

Director, Site Reliability Engineering

Location
Manchester, England, United Kingdom
engineered into every service throughout its lifecycle. Working alongside the Director of Site Reliability Operations, this leader will define the engineering standards, automation, observability, production readiness, and resilience capabilities that enable world‐class operational performance. While Site Reliability Operations owns the day‐to‐day operation of production services, the Site … global Site Reliability Engineering organization. This role owns the engineering strategy, governance, architecture, and technical practices that improve service reliability through software engineering, observability, automation, resilience engineering, production engineering, and operational readiness. Rather than operating production systems on a day‐to‐day basis, the Site Reliability Engineering organization develops ...

Cloud / Infrastructure Engineer

Location
Whiteley, England, United Kingdom
Partner with application engineering teams to improve deployment workflows and reduce operational risk Operational Readiness and Collaboration Integrate infrastructure with monitoring, logging, alerting, and observability systems Support production operations and participate in incident response activities related to infrastructure and deployments Collaborate with backend and data engineering teams to enable scalable … reliable service operation Develop and maintain operational documentation, procedures, and support runbooks Utilize observability data to identify issues and drive corrective actions Required Qualifications 5+ years of experience in cloud, infrastructure, or DevOps engineering roles Hands-on ownership of AWS infrastructure in production environments Strong experience with Infrastructure-as-Code ...

Agent Engineer

Location
Greater London, England, United Kingdom
users to create user-friendly solutions for complex business processes Participate in technical design reviews, planning sessions, and code reviews Contribute to infrastructure and observability practices with the Engineering team Continuously improve quality, reliability, and usability across internal platforms Support Elliptic's mission to make crypto markets safer, more transparent … framework experience such as NestJS or Express is nice to have Terraform or infrastructure-as-code experience is nice to have Datadog or similar observability platform experience is nice to have DynamoDB or other NoSQL database experience at scale is nice to have Distributed or event-driven architecture experience, including ...

Principal Software Engineer-AI

Location
Greater London, England, United Kingdom
production or in platforming (LLM or MCP gateway, agentic runtime, auth, data retrieval, eval tooling) Experience running AI systems in production at scale, including observability, cost and capacity planning, regression detection, and incident response for AI-powered applications Experience operating production distributed systems on AWS/Azure, with a strong … grasp of reliability, observability, and incident response at scale Deep knowledge of cloud-native technologies, serverless applications, event-driven architectures, data and inference pipelines, relational, NoSQL, and vector databases, and modern software architecture patterns Proven track record of owning multi-year technical strategy and architectural roadmaps, guiding teams from ...

Principal Software Engineer - Full Stack - AI

Location
York and North Yorkshire, England, United Kingdom
full stack, guiding the development of responsive frontend applications (React/TypeScript) and robust, scalable backend services (Python, Java, Kotlin, or Node.js). LLM Observability & Reliability: Establish robust LLM observability, evaluations, and caching, implementing latency optimisations and comprehensive monitoring (logging, usage tracking, agent behaviour). Operational Excellence: Champion high availability … performance optimisation, and observability across frontends and backend microservices, focusing on practices that maintain platform reliability and optimise MTTD and MTTR. Mentorship & Collaboration: Elevate the engineering organisation by mentoring senior and junior engineers, conducting rigorous code and system design reviews, and partnering with product managers to translate product visions into ...

Azure Systems Engineer

Location
Greater London, England, United Kingdom
integrations and production environments Implementing Infrastructure as Code (IaC) and automation solutions Supporting identity, access management, security and governance across Azure Improving platform monitoring, observability, resilience and performance Supporting SQL Server and Azure SQL environments, including access, backups and troubleshooting Contributing to vulnerability management, incident response and operational security Strong … Entra ID, IAM, RBAC and privileged access Understanding of secure application delivery, configuration and secrets management Experience with Azure Monitor, Application Insights or equivalent observability tooling Knowledge of SQL Server/Azure SQL , including access management, backups and troubleshooting Strong understanding of TCP/IP, DNS, routing, firewalls and private ...

Senior DevOps Engineer - Cloud Automation & AI Tools

Location
Greater London, England, United Kingdom
operate RX Cloud Platforms across AWS and Azure from Richmond, London. You will partner with software engineers to improve cloud infrastructure, deployment automation, observability and platform reliability. The role emphasizes automation-first priorities, AI-assisted tooling, and collaboration with agile teams to drive incident resolution and continuous improvement. #J ...

Senior Cloud DevOps Engineer — AWS/Azure, IaC & CI/CD

Location
Richmond, England, United Kingdom
from Richmond, London. You’ll collaborate with software teams to improve infrastructure, deployment automation and platform reliability. You will implement security best practices, maintain observability, and drive continuous improvement using AI-assisted tooling and modern CI/CD processes. Flexible working patterns are available. #J-18808-Ljbffr ...

Cloud DevOps Engineer — AI Tools, Flexible Hours

Location
Richmond, England, United Kingdom
Cloud Platforms across AWS and Microsoft Azure. You will work closely with software engineering and cloud teams to improve cloud infrastructure, deployment automation, observability and platform reliability. The role emphasizes Infrastructure as Code, Terraform/CloudFormation, GitHub Actions, and cloud security, with scripting #J-18808-Ljbffr ...

Senior Cloud SRE: Azure, Terraform & Kubernetes

Location
Newcastle upon Tyne, England, United Kingdom
Trimble is seeking a Site Reliability Engineer to own and scale production infrastructure on cloud environments. You will implement IaC with Terraform, enhance observability using multiple monitoring tools, and drive CI/CD pipelines with Jenkins and GitHub. The role involves leading incident response and collaborating across teams to ensure ...

Platform Engineer: Kubernetes, Cloud & Automation

Location
Greater London, England, United Kingdom
harden infrastructure, apply IaC with Terraform/Ansible/Helm, and ensure secure, scalable operations across GCP, AWS and Azure, while contributing to observability and performance tuning. #J-18808-Ljbffr ...

AI Platform Engineer - Azure, Kubernetes & CI/CD

Location
Manchester, England, United Kingdom
responsibilities include building AKS-based deployments, IaC with Terraform, and robust CI/CD pipelines using Azure DevOps, with a focus on reliability, observability, and secure practices. #J-18808-Ljbffr ...

VP DevOps/SRE: AI/ML Infra, CI/CD & Reliability

Location
Glasgow, Scotland, United Kingdom
DevOps/SRE Engineer - Vice President to own automation, reliability, and production operations for AI/ML services. You will build CI/CD, observability, and incident-management practices across international markets. Leverage Terraform, Kubernetes, and cloud-native tooling to scale release automation and reliability, while mentoring engineers and enforcing ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
performance and resilience of the systems that support our global product. This role combines software engineering, automation and incident response to reduce toil, sharpen observability and strengthen service health across a complex technical estate. You will work with Open Telemetry, logging, telemetry and automation to surface issues faster and improve … including testing, source control and delivery lifecycles. An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. Proficiency in shell scripting for automation and system management. Experience with Infrastructure as Code, including ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Stoke-On-Trent, Staffordshire, West Midlands, United Kingdom
Employment Type
Permanent, Work From Home
performance and resilience of the systems that support our global product. This role combines software engineering, automation and incident response to reduce toil, sharpen observability and strengthen service health across a complex technical estate. You will work with Open Telemetry, logging, telemetry and automation to surface issues faster and improve … including testing, source control and delivery lifecycles. An understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. Proficiency in shell scripting for automation and system management. Experience with Infrastructure as Code, including ...

DevOps and Machine Learning Operations Engineer

Location
Manchester, England, United Kingdom
runtime platform the project depends on, as well as the model-serving path. The role covers infrastructure as code, continuous integration and deployment, observability, cost control, and production support. Applicants must have experience running systems in production and being accountable for their reliability. Main Responsibilities Infrastructure and Environments: Define … gates that fail closed. Automate database migration and rollback for reversible releases. Support progressive delivery, including staged rollout and fast rollback, with deployment tracking. Observability and Operations: Instrument services with structured logging, metrics, and tracing, and define user-focused alerts. Establish service level objectives and report against them. Run incident ...

Lead SRE - AWS Platform

Location
Glasgow, Scotland, United Kingdom
your team to identify comprehensive service level indicators and partner with stakeholders to establish reasonable service level objectives and error budgets Design and implement observability frameworks and alerting strategies, including white and black box monitoring, service level objective-based alerting, and telemetry collection to ensure proactive detection and response Serve … resiliency best practices Fluency in at least one programming language such as Python, Java/Spring Boot, or .NET Proficient knowledge and experience in observability, including white and black box monitoring, service level objective alerting, and telemetry collection across large-scale production environments Proficiency with continuous integration and continuous delivery ...

Cloud Solution Architect — Java/Spring, GCP & IAM

Location
Leeds, England, United Kingdom
software development and 3+ years as an architect, you’ll lead architectural reviews, code reviews, and performance tuning, leveraging Kubernetes, Docker, Terraform, and modern observability tools. #J-18808-Ljbffr ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
Drive automation across infrastructure, application delivery, operational processes, and platform management Establish platform standards, engineering patterns, and best practices Improve platform reliability, scalability, performance, observability, and operational efficiency Reduce engineering friction and accelerate software delivery Establish engineering principles and guardrails for security, reliability, and governance Lead complex platform initiatives from … automated software delivery Experience automating operational processes and platform lifecycle management Experience establishing repeatable, standardised engineering workflows Experience designing for resilience, fault tolerance, observability, and operational excellence Experience applying SRE principles and practices Experience in performance analysis, capacity planning, scalability engineering, and proactive reliability improvement Experience establishing service-level objectives ...

Senior AI Platform Engineer - GenAI & LLM Architect

Location
Sheffield, England, United Kingdom
oversee LLM, RAG, and Agentic AI integrations across Azure, AWS, and GCP, collaborate with architects to ensure secure, maintainable solutions, and set standards for observability, governance, and #J-18808-Ljbffr ...