151 to 173 of 173 Observability Jobs in the South West

Lead Site Reliability Engineer (Dynatrace)

Hiring Organisation
SF Partners
Location
South West England, United Kingdom
Employment Type
Full-Time
Salary
£80,000 - £100,000 per annum
looking for an experienced Site Reliability Engineer/Observability Engineer with deep Dynatrace expertise to join a major technology and platform engineering programme. This is not a role for someone who has simply used Dynatrace dashboards. We're looking for an engineer who has been involved in the implementation, configuration … technical SME within complex production environments. What we're looking for Strong hands-on Dynatrace implementation and administration experience Experience designing and implementing observability/monitoring solutions end-to-end Strong SRE and production engineering background Experience configuring instrumentation, metrics, alerting and monitoring Understanding of technologies such as OneAgent, ActiveGate ...

Senior Software Engineer-AI

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
operation with limited supervision Hands-on experience with cloud-native technologies, serverless applications, event-driven architectures, data pipelines, relational and NoSQL databases, vector databases, observability tooling, and automated deployment pipelines Solid understanding of algorithms, data structures, scalability, reliability, performance optimization, security best practices, and engineering trade-offs Experience mentoring engineers … technical designs, participate in design reviews, and identify risks, constraints, trade-offs, and alternative approaches Maintain engineering excellence through automated testing, code reviews, observability, monitoring, alerting, operational readiness, and participation in on-call support Apply machine learning operations practices, including prompt versioning, automated evaluation, deployment pipelines, monitoring, and production issue ...

Senior Associate, Full-Stack Engineer Opportunities

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
ways: Design, build, and maintain backend services, batches and APIs, contributing to UI components as needed. Own end-to-end delivery: implementation, testing, deployment, observability, and reliability. Write clean, well-tested code; participate in code reviews and continuous improvement. Collaborate with product, design, and operations to translate business needs into … microservices Proficiency in Java with Spring. Experience with CI/CD, automated testing (JUnit/Spock), and containers (Docker). Familiarity with microservices, observability/telemetry (e.g., Splunk, AppDynamics), and cloud deployments. Curiosity to understand the business domain and translate product strategy into technical solutions. Working knowledge of Groovy with ...

Principal Platform Engineer (12 Month FTC)

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Data Architect

Location
Bristol, England, United Kingdom
We believe in the power of ingenuity to build a positive human future.We challenge where it matters and own the outcome.As strategies, technologies, and innovation collide, we create opportunity from complexity. Our teams of interdisciplinary ...

Data Consultant - DV CLEAR

Hiring Organisation
Hays Specialist Recruitment Limited
Location
Corsham, Wiltshire, United Kingdom
Employment Type
Full-Time
Salary
£518.88 per day
Your new company You will be working for a secure client delivering data engineering capability in a highly regulated environment. The client's identity will remain confidential at this stage. The position requires previous experience ...

Principal Software Engineer-AI

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
production or in platforming (LLM or MCP gateway, agentic runtime, auth, data retrieval, eval tooling) Experience running AI systems in production at scale, including observability, cost and capacity planning, regression detection, and incident response for AI-powered applications Experience operating production distributed systems on AWS/Azure, with a strong … grasp of reliability, observability, and incident response at scale Deep knowledge of cloud-native technologies, serverless applications, event-driven architectures, data and inference pipelines, relational, NoSQL, and vector databases, and modern software architecture patterns Proven track record of owning multi-year technical strategy and architectural roadmaps, guiding teams from ...

Product Associate - SRE Team - Chase UK

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … services are designed, delivered, and operated with reliability in mind. Job responsibilities Support the product strategy and delivery of reliability capabilities, including standards, observability, incident practices, automation, and developer experience improvements. Partner with engineers, site reliability engineers, and cross-functional teams to understand problems, gather requirements, and translate ideas into ...

Engineering Architect

Location
Bath, England, United Kingdom
will own and evolve the technical foundations that underpin product and service delivery, including engineering standards, development frameworks, CI/CD pipelines, testing harnesses, observability, and operational practices. Your focus will be on enabling a fast-moving, AI-first delivery team to consistently deliver high-quality, scalable, and supportable solutions … varied role, you will: own and version technical foundations across all seven harness dimensions (Security, performance, coding style, UI/UX baselines, logging and observability, testing) so that every deliverable inherits secure, observable and tested defaults build and maintain the shared ADO CI/CD pipelines and release governance ...

Senior Site Reliability Engineer

Location
Swindon, England, United Kingdom
drive automation, improve reliability, and champion DevOps and Site Reliability Engineering (SRE) best practices. You'll play a pivotal role in enhancing service resilience, observability, and operational efficiency while helping shape the future of our cloud capabilities through continuous improvement and innovation. Responsibilities Design, implement, and maintain cloud infrastructure, deployment … establish and promote DevOps best practices, standards, and ways of working. Contribute to the adoption of Site Reliability Engineering (SRE) practices, supporting improvements in observability, monitoring, service reliability, incident management, and operational excellence. Implement and support infrastructure provisioning, configuration management, and governance solutions using technologies such as Terraform, OpenTofu, Scalr ...

Senior DevOps Engineer

Hiring Organisation
Sanderson Recruitment
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent
Salary
£550 - £600 per day
configuration management Administer and optimise Linux environments, primarily CentOS/RHEL Improve automation across infrastructure and operational processes Implement and support monitoring, alerting and observability solutions Work closely with engineering teams to deliver reliable and secure platforms Support incident management, troubleshooting and root cause analysis Contribute to security, governance …/CD pipelines Strong Linux administration experience ( CentOS, RHEL or similar ) Experience implementing Infrastructure as Code and automation solutions Knowledge of monitoring and observability tools Scripting experience with Bash and/or Python Experience working within a regulated environment , such as Financial Services, Banking, FinTech, Insurance or another highly governed ...

Senior Infrastructure Platform Engineer – Veeam / VMware / Hyper-V

Hiring Organisation
100 Percent
Location
Bristol, City of Bristol, United Kingdom
Employment Type
Permanent
Salary
£55000 - £60000/annum Bonus + Benefits
major incident recovery, particularly around backup, restore and platform availability. Drive automation to reduce manual processes and improve operational efficiency. Develop and maintain monitoring, observability and alerting using tools such as Grafana. Coordinate the day-to-day priorities of the Platform Engineering team. Maintain engineering standards, documentation and operational procedures. … Azure DevOps-focused role. Experience supporting highly available production infrastructure. Strong troubleshooting and problem-solving skills across enterprise infrastructure. Experience with monitoring and observability platforms such as Grafana. Experience automating operational tasks using PowerShell, scripting or similar technologies. Excellent understanding of backup, disaster recovery and platform resilience. Ability to coordinate ...

Senior Manager, Platform Software Engineering

Hiring Organisation
Oracle Corporation
Location
Bristol, Gloucestershire, United Kingdom
Salary
£ 80 K
serving multiple consumers. Accountable for performance, cost control, and reliability outcomes; escalations for distributed debugging and upgrade safety. Drives significant cross-team improvements to observability, SLOs, and resilience patterns. Engages stakeholders inside and outside the group; socializes migration plans and adoption playbooks.Only Oracle brings together the data, infrastructure, applications … client libraries used by several consumer teams.Establish group-level guardrails for APIs/SDKs, deprecation strategies, and upgrade playbooks; track adoption and risk.Sponsor shared observability frameworks and capacity models across teams; review error-budget posture and corrective actions.Software Development and Coding - Design, Testing, and Optimization:Leads teams and provides technical ...

Junior Platform Engineer Telemetry, SIEM, Observability

Hiring Organisation
apto solutions
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent
Salary
£40,000
engineering seat with real customers, real supervision and a clear path up. ABOUT APTO SOLUTIONS Apto Solutions is a Bristol business that runs telemetry, observability and SIEM platforms for other organisations . Banks, airports, energy companies, engineering firms and government bodies depend on those platforms to tell them whether their ...

Vice President, Production Services Application Support

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
risk while improving platform stability. Drive initiatives through to completion with strong ownership, accountability, urgency, and quality. Champion an automation-first mindset, leveraging AI, observability, and tooling to reduce manual effort and improve service quality. Identify systemic issues and drive sustainable remediation through process simplification, platform improvements, and close partnership … incident, problem, and change management with measurable improvements in stability and service recovery. Demonstrated automation-first and AI-enabled mindset, with experience driving tooling, observability, and process automation. Strong ownership mentality and execution focus, with the ability to take initiatives from concept through delivery and embed sustainable outcomes. Deep technical ...

Lead Software Engineer - Proxy/SSE Network Security

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
resilience outcomes. Drive operational excellence at scale for perimeter, proxy, and SSE services in the US, including incident, change, and problem management rigor, observability and resiliency validation practices, automation to improve repeatability and evidence quality, reduction of client and partner impact, and execution of Technology Lifecycle Management (TLM) and modernization … design, exception frameworks, audit-ready traceability, and measurable risk reduction reporting. Experience with large-scale operations for externally facing or security enforcement services, including observability strategy, resilience testing, incident response alignment, and reduction of repeat incidents and client-impacting events. Experience designing and operating hybrid edge architectures and cloud interconnect ...

Senior Software Engineer

Hiring Organisation
IO Associates
Location
Bath, Somerset, South West, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
£500 - 550 per day
environment Work with AWS serverless technologies , including Lambda, SQS, EventBridge and API Gateway Work with MongoDB and document databases Contribute to CI/CD, observability, incident management and production support Work closely with Product, Engineering and Operations to turn business requirements into scalable solutions Mentor engineers and contribute to engineering … standards and best practice Key Skills Strong Node.js & TypeScript Strong system design and architecture AWS & Serverless MongoDB/document databases Distributed systems, observability and incident management CI/CD and production operations Experience working in a regulated environment , ideally FinTech or financial services Comfortable owning services within a build ...

Technical Lead (Portal & Applications)

Hiring Organisation
Lightfoot
Location
Exeter, Devon, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
built into delivery. Leading our move to agentic software development, turning a proven individual approach into a repeatable team capability. Building security and observability into the platform from the start. Developing the technical capability of the team through design reviews, mentoring and hands-on leadership. You’ll also line-manage … choosing a stack. The interesting decisions are the ones underneath it: service boundaries, API design, migration sequencing, testing strategy, engineering standards, secure-development practices, observability and how agentic development works across a team. One important constraint: our database is shared across several engineering teams and is mission-critical beyond Applications. ...

Front Office Equities Trading Technology Support

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
firm's systems to ensure operational stability and availability Assist in the monitoring of production environments for anomalies and address issues utilizing standard observability tools Identity issues for escalation and communication, and provide solutions to the business and technology stakeholders Analyze complex situations and trends to anticipate and solve incident … achieve common goals Demonstrates knowledge of applications or infrastructure in a large-scale technology environment both on premises and public cloud Experience in observability and monitoring tools and techniques Exposure to processes in scope of the Information Technology Infrastructure Library (ITIL) framework Preferred qualifications, capabilities, and skills Experience with ...

Hybrid SAP Basis Engineer — Cloud ALM & Observability

Location
Swindon, England, United Kingdom
estate. You’ll work across strategic programmes to enhance SAP Solution Manager capabilities and support the transition to SAP Cloud ALM, advancing observability and governance across SAP platforms. #J-18808-Ljbffr ...

SAP Basis Engineer – Cloud ALM & Observability

Location
Swindon, England, United Kingdom
capabilities across Technical Monitoring and Change Request Management. As Nationwide transitions to SAP Cloud ALM, you’ll contribute to enabling new capabilities and modern observability practices. #J-18808-Ljbffr ...

Sr Director, Platform Engineering – Data Platform & Agentic Platform

Hiring Organisation
RELX Group
Location
Exeter, Devon, United Kingdom
Salary
£ 100 K
mechanismsImplement and operate agent workflow platform capabilities aligned to product-defined standards and interfaces, including traceability, state handling, and convergence patternsImplement production-grade evaluation, observability, auditability, and guardrail mechanisms required for safe AI workflowsRequirements12+ years leading platform engineering teams delivering shared services or large-scale SaaS systems with production operations … including ingestion and integration patterns, data quality enforcement, metadata and lineage, and secure access controlsExperience enabling AI-powered systems in production, including evaluation, monitoring, observability, auditability, and guardrailsStrong ability to partner with product leaders on contract-driven platforms and deliver outcomes without creating bottlenecksProven capability to hire and develop high ...

Enterprise AI Architect - Tech for Good - Remote - Greenfield

Hiring Organisation
Sanderson Recruitment
Location
South West, United Kingdom
Employment Type
Contract
data platforms. A scalable agent architecture covering identity, authentication, delegated actions, human-in-the-loop controls, context and memory, registration, ownership, versioning and retirement. Observability requirements including minimum telemetry and audit evidence showing which agent acted, for whom, and using which model, tools and data. AI strategy and five-year … there is one proportionate review route rather than several. Priority deliverables: are the enterprise AI target architecture, a scalable agent architecture framework, AI observability and operational architecture, the five-year AI roadmap, AI design principles, AI and agent design standards, and an integrated governance and technical assurance framework. Knowledge transfer ...