76 to 100 of 140 Observability Jobs in the South West

Senior Platform Engineer Cheltenham

Location
Cheltenham, England, United Kingdom
delivery pipelines. Work across greenfield and brownfield platform engineering projects. Develop infrastructure-as-code solutions using modern tooling and engineering practices. Improve platform reliability, observability and operational maturity through automation and engineering excellence. Work closely with developers, architects and security teams to understand challenges and deliver pragmatic solutions. Champion platform … Experience with container technologies such as Docker and Kubernetes. Experience with cloud platforms such as AWS, Azure or GCP. Experience implementing monitoring, logging and observability solutions. Ability to work effectively across engineering, security and customer teams. Strong problem-solving skills with the ability to operate in complex and ambiguous environments. ...

Lead Java Developer (London)

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
automation, security and resilience. Champion effective CI/CD practices and continuous improvement across the development lifecycle. Implement and promote effective monitoring, logging and observability practices. Investigate and resolve production issues, taking ownership through to resolution and ensuring lessons are incorporated into future development. Collaborate with engineers, architects, product teams … environments. AWS cloud services and cloud-native application development. Terraform and Infrastructure as Code. CI/CD tooling and modern DevOps practices. Monitoring and observability tools, such as Grafana. Automated testing and engineering quality practices. Why join AND Digital? We have three values: wonder, share, and delight. These values inform ...

Software Engineer III - AI/ML Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Bournemouth, Dorset, South West, United Kingdom
Employment Type
Permanent
enhance the reliability and scalability of AI/ML platforms and applications to accommodate fast-growing demands. Own NFRs and develop tooling for observability, security, resilience, infrastructure management and operations excellence. Build and maintain scalable infrastructure to support the deployment and operation of large-scale AI platforms and apps. Build … architecture. Experience building large scale infrastructure and and cloud-native delivery practice in Google Cloud, AWS, or Azure and Terraform. Extensive experience implementing advanced observability using tools like Open Telemetry, Dynatrace, Grafana, and/or cloud-native services. Systematic problem-solving and troubleshooting skills in a complex system. Hands ...

Fullstack Engineer (DV Clearance) - Cloud Native, 37.5h

Location
Cheltenham, England, United Kingdom
Java applications, with cloud native deployments in OpenShift and Kubernetes. You will work across distributed architectures, implement Web API integrations, and contribute to observability and security practices. A DV clearance or recent status is required. #J-18808-Ljbffr ...

Lead Java/Go Engineer - Enterprise Automation Platform

Location
Bournemouth, England, United Kingdom
across AaaS, Ansible Automation Platform, AutoM8, and AI-enabled service management. You will work within an agile team, shaping platform tooling, event-driven automation, observability, and cloud readiness to improve reliability and scale across #J-18808-Ljbffr ...

Principal Machine Learning Engineer – Production Systems

Location
Bristol, England, United Kingdom
profiling. APIs : Proficiency in gRPC/Protobuf and REST for cross-language integration. Performance Optimization : GPU acceleration (CUDA/cuDNN), mixed precision, XLA, profiling. Observability : Metrics, tracing, structured logging, dashboards. Security : SBOM, image signing, role-based access, vulnerability scanning. Preferred Qualifications Experience with ONNX Runtime Training, PyTorch, or hybrid ...

AI Technical Lead

Hiring Organisation
F5 consultants
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent
Salary
£80,000
delivery of LLM, RAG and agentic AI solutions Take ideas from discovery and proof of concept through to production Establish standards for evaluation, security, observability, governance and cost control Work directly with product and operational stakeholders to turn high-level problems into deliverable solutions Mentor engineers, review designs and code ...

Developer Experience Engineer New Bristol

Location
Bristol, England, United Kingdom
tools that customers use to run, optimise, and monitor large language models on Fractile hardware. Build the interfaces to existing diagnostic, profiling, and observability tools that let developers understand performance profiles, debug issues, and optimise their deployments. Write the documentation, quickstarts, reference examples, and tutorials that define a developer ...

Principal Java Engineer - Backend Platform

Location
Bristol, England, United Kingdom
REST and WebSocket APIs that serve both a live data product for immediate consumption and an insights portal for historical analysis and performance comparison. Observability and reliability - instrumentation, logging, and tracing are built in from the start, not bolted on. The platform needs to run 24/7 without ...

Lead Software Developer

Hiring Organisation
ALFA TECHNOLOGY RECRUITMENT LTD
Location
South West London, London, United Kingdom
Employment Type
Permanent
it. Speed and reliability, designed in. Live voice needs sub-second reads and a platform that doesn't need babysitting. Latency budgets per domain, observability shipped with the feature, rollback rehearsed every release, small changes to production daily. You'll move live systems onto a modern stack without dropping anything. ...

Lead Backend Engineer

Hiring Organisation
ALFA TECHNOLOGY RECRUITMENT LTD
Location
Bournemouth, Dorset, South West, United Kingdom
Employment Type
Permanent
it. Speed and reliability, designed in. Live voice needs sub-second reads and a platform that doesn't need babysitting. Latency budgets per domain, observability shipped with the feature, rollback rehearsed every release, small changes to production daily. You'll move live systems onto a modern stack without dropping anything. ...

DevOps Engineer - SC Cleared

Hiring Organisation
Opus Recruitment Solutions Ltd
Location
Martock, Somerset, United Kingdom
Employment Type
Full-Time
Salary
£350.00 - £500.00 per day
CI. Automate deployment, configuration and operational processes to improve efficiency and reliability. Support containerised workloads and cloud-native deployment patterns. Implement monitoring, logging and observability solutions to improve platform visibility and operational performance. Drive infrastructure resilience, scalability, security and cost optimisation initiatives. Support incident response, troubleshooting and root cause analysis … Code expertise. Experience building and maintaining CI/CD pipelines, ideally using GitLab CI. Strong Docker and containerisation experience. Experience with monitoring and observability tools such as Prometheus, Grafana and CloudWatch. Knowledge of AWS security best practices, IAM and Well-Architected principles. Experience implementing automation and reducing manual operational overhead. ...

Staff Software Engineer - AI

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency Establish engineering best practices through hands-on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement Partner with ...

Lead .Net Software Engineer

Location
Cheltenham, England, United Kingdom
technical risks, bottlenecks, dependencies, and opportunities for improvement Contribute to AWS‐based architecture and engineering practices, including environments using Lambda, ECS, and EC2 Improve observability, reliability, performance, security, and operational readiness across the platform Contribute to Jenkins pipelines and CI/CD practices to improve consistency and delivery efficiency Help … with DevOps practices and infrastructure‐aware development Experience with messaging, event streaming, or related event‐driven technologies Experience with Blazor and MudBlazor Familiarity with observability and monitoring tooling in distributed systems Experience defining governance, controls, or operating models for AI agents or AI‐enabled internal tools Experience working ...

HPC Infrastructure Site Reliability Engineer

Location
Gloucester, England, United Kingdom
role in continuous service improvement (CSI)—reducing operational toil, increasing automation, and improving reliability, consistency, and operational efficiency across the platform. This includes strengthening observability, refining operational workflows, and eliminating repetitive or failure‐prone processes. Over time, you will help shape future infrastructure design and deployment approaches, feeding operational insight ...

Operations Team Lead (Production & Reliability)

Location
Bath, England, United Kingdom
Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under pressure, clear in communication Systems thinker, fixes root causes, not symptoms How We Think Production is sacred. Clear ownership beats ambiguity. ...

Operations Team Lead (Production & Reliability)

Location
Bournemouth, England, United Kingdom
Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under pressure, clear in communication Systems thinker, fixes root causes, not symptoms How We Think Production is sacred. Clear ownership beats ambiguity. ...

Staff Software Engineer-AI

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency Establish engineering best practices through hands-on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement Partner with ...

Engineering Manager: Platform & Growth Lead (Hybrid)

Location
Bristol, England, United Kingdom
client journeys on React and React Native. You'll coach engineers, own delivery, guide architectural decisions, partner with product, and drive observability and high-velocity delivery in a regulated financial services environment. #J-18808-Ljbffr ...

Senior Backend Engineer: Own architecture & resilient systems

Location
Bath, England, United Kingdom
production systems, aiming for scalable, resilient services and robust incident handling. Collaborating with Engineering, Product and Operations, you will drive CI/CD, observability and reliability improvements across a suite of business-critical platforms. #J-18808-Ljbffr ...

24/7 HPC Infra SRE for AI & GPU Compute

Location
Gloucester, England, United Kingdom
work across network, storage, virtualization and orchestration with hands‐on Linux expertise, NVIDIA GPU ecosystems, RoCE/InfiniBand, and performance benchmarking. This role champions observability, automation and on‐call reliability, shaping next‐gen HPC platforms within a globally distributed team. #J-18808-Ljbffr ...

Production Reliability Lead

Location
Bournemouth, England, United Kingdom
Complexio is seeking an Operations Team Lead to own production and build a scalable reliability platform. You will lead live customer-facing systems, drive observability, and push for proactive improvements across incident management, on-call rotations, and delivery readiness. This hands-on role demands strong SRE/DevOps experience, proven ...

AI Product Engineer – Shape Next-Gen National Security

Location
Cheltenham, England, United Kingdom
driven team building AI capabilities for national security and government customers, shaping product direction, and applying modern engineering practices including testing, CI/CD, observability and robust #J-18808-Ljbffr ...

Data Engineer

Hiring Organisation
Brio Digital
Location
London, Baker Street, United Kingdom
Employment Type
Contract
Contract Rate
£500/hour
with Microsoft Fabric and/or Azure Experience with AWS services such as S3, Glue, Athena, Lambda and Firehose Experience working with operational/observability data from CloudWatch Understanding of GDPR, NHS information governance and data minimisation principles Apply now or email for more information ...

Data Architect - DV CLEAR

Hiring Organisation
Hays Specialist Recruitment Limited
Location
South West England, United Kingdom
Employment Type
Full-Time
Salary
£800.00 - £871.55 per day
quality requirements Producing High-Level Designs, design principles, Architecture Decision Records and assurance documentation Defining non-functional requirements covering security, performance, scalability, resilience and observability Supporting accreditation, assurance and security reviews Providing technical governance throughout build, test, assurance and handover What you'll need to succeed You will need: Active ...