176 to 200 of 4,022 Observability Jobs

Senior AWS Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£60000 - £70000/annum Bonus, Medical Care
Have: Practical experience managing large-scale Kubernetes clusters; certifications in Kubernetes are a strong bonus Hands-on familiarity with the Grafana Observability Suite, including tools like Loki, Mimir, and Tempo Background in administering or developing with popular monitoring and automation tools such as Splunk, Datadog, PagerDuty, or Rundeck Experience using … with tools such as Jenkins, GitLab CI/CD, or CircleCI Strong understanding of containerization (e.g., Docker, Kubernetes) and microservices architecture Skilled in using observability and monitoring tools such as Prometheus, Grafana, ELK stack, or AWS CloudWatch Excellent analytical and troubleshooting abilities, especially within complex distributed systems Proven experience handling ...

Senior Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment
Location
Southampton, Hampshire, United Kingdom
Employment Type
Permanent
Salary
£60000 - £70000/annum
Have: Practical experience managing large-scale Kubernetes clusters; certifications in Kubernetes are a strong bonus Hands-on familiarity with the Grafana Observability Suite, including tools like Loki, Mimir, and Tempo Background in administering or developing with popular monitoring and automation tools such as Splunk, Datadog, PagerDuty, or Rundeck Experience using … with tools such as Jenkins, GitLab CI/CD, or CircleCI Strong understanding of containerization (e.g., Docker, Kubernetes) and microservices architecture Skilled in using observability and monitoring tools such as Prometheus, Grafana, ELK stack, or AWS CloudWatch Excellent analytical and troubleshooting abilities, especially within complex distributed systems Proven experience handling ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
people freedom while keeping us in control. What you’ll be doing from day one: Owning our infrastructure as code in Terraform, plus alerting, observability (Prometheus, Grafana) and reliability, including load testing and disaster recovery exercises. Making CI/CD faster (GitHub Actions, ArgoCD) and taking obstacles out of engineers … ideally multi-region. Hands-on with GCP, with Azure a bonus. You understand how model consumption works through each cloud. Deep experience with Terraform, observability and CI/CD, and you write solid Python (Elixir is a bonus, or you’re keen to learn). A clear communicator ...

Lead Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment Limited
Location
Southampton, Hampshire, South East, United Kingdom
Employment Type
Permanent
Salary
£85,000
that business-critical cloud platforms remain observable, measurable, secure, scalable and reliable. The successful candidate will combine strong Azure engineering expertise with platform automation, observability and operational leadership. Candidates are likely to come from a Site Reliability Engineering, DevOps, Cloud Engineering, Platform Engineering or Cloud Development background. Security requirement: Applicants … service level agreements, service level indicators and error budgets. Design and implement monitoring, alerting and dashboarding across cloud platforms and microservices. Deploy and configure observability technologies including Grafana, Prometheus, Azure Monitor and OpenTelemetry. Develop custom application and platform metrics to improve operational visibility. Create advanced queries, dashboards and alerts ...

Lead Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment Limited
Location
Northam, Devon, UK
that business-critical cloud platforms remain observable, measurable, secure, scalable and reliable. The successful candidate will combine strong Azure engineering expertise with platform automation, observability and operational leadership. Candidates are likely to come from a Site Reliability Engineering, DevOps, Cloud Engineering, Platform Engineering or Cloud Development background. Security requirement: Applicants … service level agreements, service level indicators and error budgets. Design and implement monitoring, alerting and dashboarding across cloud platforms and microservices. Deploy and configure observability technologies including Grafana, Prometheus, Azure Monitor and OpenTelemetry. Develop custom application and platform metrics to improve operational visibility. Create advanced queries, dashboards and alerts ...

VodafoneThree - SRE III

Location
Greater London, England, United Kingdom
enable teams to validate scalability, reliability, and operational readiness from the earliest stages of design and development. Your expertise in software engineering, performance testing, observability, and automation will help drive the adoption of engineering best practices across the organisation. You will lead initiatives to integrate performance and resilience validation into … service capabilities, providing tooling, frameworks, standards, and guidance for performance, resilience, and chaos testing. Support in defining and championing engineering best practices, including SLOs, observability, capacity planning, and operational readiness to improve service reliability and customer experience. You will collaborate closely Product, Engineering, Platform, Architecture, and SRE teams to design ...

Platform Engineer, SDO London, United Kingdom

Location
Greater London, England, United Kingdom
optimise cloud systems for performance, reliability, and cost efficiency. Assist in managing containerised workloads and orchestration platforms (e.g. Docker, Kubernetes). Implement and maintain observability tools (logging, metrics, alerting) to ensure system health and rapid incident response. Work with engineering teams to ensure infrastructure meets application requirements and supports scalable … cloud architecture, and system reliability. Strong troubleshooting and problem‐solving skills. Desirable: Experience with containerisation and orchestration (Docker, Kubernetes). Familiarity with monitoring and observability tools (e.g. Prometheus, Grafana, ELK). Experience working with Linux systems and shell scripting. Programming or scripting experience (e.g. Python, Bash, Go, or similar). ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
integrations, and human‐in‐the‐loop controls evolve across the stack. Continuously identify and exploit opportunities to improve performance, reliability, and user experience, using observability and analysis to find signals in noisy systems. Navigate confidently across legacy and greenfield contexts, applying AI tooling pragmatically to modernise where it matters most. … across both human-written and AI-generated code, embedding security validation into the development pipeline rather than treating it as an afterthought. Own system observability and reliability, moving from reactive alerting to proactive, model-assisted incident prevention. Participate in on-call rotation, bringing the same rigour to incident response that ...

Platform Engineer, SDO

Hiring Organisation
wayve
Location
London, United Kingdom
Salary
£ 70 K
troubleshoot, and optimise cloud systems for performance, reliability, and cost efficiency.Assist in managing containerised workloads and orchestration platforms (e.g. Docker, Kubernetes).Implement and maintain observability tools (logging, metrics, alerting) to ensure system health and rapid incident response.Work with engineering teams to ensure infrastructure meets application requirements and supports scalable software … understanding of networking, cloud architecture, and system reliability.Strong troubleshooting and problem-solving skills.Desirable:Experience with containerisation and orchestration (Docker, Kubernetes).Familiarity with monitoring and observability tools (e.g. Prometheus, Grafana, ELK).Experience working with Linux systems and shell scripting.Programming or scripting experience (e.g. Python, Bash, Go, or similar).Understanding of security ...

Senior Backend Engineer (.NET & Python)

Location
Greater London, England, United Kingdom
ship, contributing to testing, troubleshooting and continuous improvement. Collaborate with Product, Design and Engineering teams to deliver customer-focused solutions. Improve platform reliability, observability, performance and security. Help evolve Benifex's AI‐assisted development practices and support other engineers across the team. What are we looking for? Commercial experience developing … backend applications using C#, .NET/.NET Core, Python and SQL. Production GenAI experience, including technologies such as MCP, RAG, agent orchestration, evaluations and observability tooling. An understanding of responsible AI, software quality, performance and security best practices. Experience delivering features end-to-end within modern engineering environments. Exposure ...

Platform Engineer

Location
Leeds, England, United Kingdom
services from a global network of offices. The Platform Engineering team builds and operates Waystone’s engineering platform, spanning cloud infrastructure, CI/CD, observability, on-call and our internal developer portal. We provide the secure, automated foundations that the firm’s regulated services run on. We are looking … architectural best practice, and help maintain continuous compliance evidence against the firm’s obligations (SOC 2, ISO 27001, GDPR, DORA). Contribute to observability and monitoring across the estate, and to incident response and on-call, including troubleshooting and root cause analysis that drives systemic improvement. Support disaster recovery testing ...

Senior Network Site Reliability Engineer

Location
Greater London, England, United Kingdom
Build and lead all aspects of our CI/CD pipelines (GitLab/GitHub) and provide automated solutions for IaC deployment. Implement automation and observability across infrastructure resources. Bring your own ideas for improving our infrastructure stack and implement them. Take part in our operational rotation to react and resolve … teams Nice to have Experience with containers and Kubernetes Familiarity with AWS governance and security controls, including SCPs and IAM policies Experience improving reliability, observability, performance, or incident response processes Exposure to large-scale CDN, edge, or traffic-routing environments Experience working in globally distributed infrastructure or platform teams Basic ...

Principal AI Platform Engineer (Python)

Location
Greater London, England, United Kingdom
pipelines (GitLab CI, GitHub Actions, Jenkins) that test, package, and release software, and the practices around them, such as automated testing and staged rollouts. Observability: experience instrumenting systems with Prometheus, Grafana, OpenTelemetry, or an equivalent stack, and using that data to diagnose failures in distributed systems. Security (critical): a working … points into platform features. Replacing manual infrastructure operations with codified, self-service workflows, so that environments are reproducible, reviewable, and safe to change. Building observability into the platform through metrics, structured logging, and alerting, so failures surface before consuming teams report them. Owning production issues in the platform ...

DataOps Engineer

Location
United Kingdom
data platforms so they are reliable, secure, and easy to change. You will work across engineering and operations to automate delivery, improve data pipeline observability, and embed good governance so analytics and AI workloads can run at scale.You will be part of the Data Platforms team that sits within … deliver DataOps capabilities, with a focus on delivering data applications in an automated approach. You will also implement and manage comprehensive monitoring and observability solutions to ensure data quality across the entire data flow, including supporting delivery in air-gapped and other restricted environments.* Delivering data applications in an automated ...

Global Banking & Markets - Software Engineer - Vice President - London London · United Kingdom [...]

Location
Greater London, England, United Kingdom
standard for years to come. What You Will Do Design, build, and operate high‐availability, multi‐region, cloud‐native services with security and comprehensive observability (metrics, distributed tracing, structured logging) built in at every layer. Develop event‐driven architectures, multi‐stage processing pipelines, and optimized data paths for high‐throughput … patterns (retry, dead‐letter queues, error isolation). Cloud & Infrastructure : Cloud platforms (GCP, AWS), container orchestration (Kubernetes, Docker), and JVM tuning for containerized workloads. Observability & Operations : Application instrumentation (metrics, distributed tracing, structured logging) and production support in high‐availability environments. Data & Performance : Data modeling, SQL/NoSQL databases, caching strategies ...

Lead Software Developer

Location
Greater London, England, United Kingdom
discussions. Participate in release planning, deployment activities, and production support. Take responsibility for the production services developed by the team. Improve system robustness, resilience, observability, performance, and operational stability. Conduct code reviews and provide quality assurance for work in progress. Troubleshoot complex technical issues and drive root cause analysis. Identify … environments. Familiarity with Government Technology Code of Practice and Service Standards. Experience with CI/CD pipelines and automated testing frameworks. Knowledge of monitoring, observability, and operational support practices in production environments. Experience leading distributed or multi-supplier teams. Pension Scheme - contributions matched up to 10% Private medical cover Income ...

Cloud Support Engineer

Location
Greater London, England, United Kingdom
providing technical depth, structure and calm during high-pressure client situations Develop and refine tools, dashboards and automation to improve support delivery, observability and onboarding Identify recurring issues, propose and lead solutions that improve platform stability, reduce effort and enhance client experience Provide mentoring, training and technical oversight … production troubleshooting Strong Linux systems knowledge, including filesystems, networking and system internals Programming skills in Golang and Python and experience with infrastructure tools or observability stacks (e.g. Grafana, Prometheus, EFK) Confidence in working with cloud-native platforms and tools (e.g. Kubernetes, Terraform, AWS/GCP, Docker) Excellent communication skills under ...

Site Reliability Engineer

Location
United Kingdom
through engineering standard process. Support incident management, root cause analysis and continuous improvement activities, helping to strengthen service reliability over time. Contribute to monitoring, observability and service performance capabilities, using operational insights to improve customer outcomes. Why Lloyds Banking Group? Like the modern Britain we serve, we're evolving. Investing … identified as a key personal skill within the SBO Skills Library. Experience supporting production services, operational platforms or live environments. Familiarity with monitoring and observability platforms such as Dynatrace. Experience with Google SecOps. Experience working with GitHub, Terraform, CI/CD pipelines and cloud-native engineering practices. Experience using Jira ...

Secure Data Engineer

Hiring Organisation
Capgemini
Location
Birmingham, UK
Employment Type
Full-time
components that process, transform and expose data for analytical, operational or AI driven use cases. You will follow strong engineering practices, with testing and observability built in from the start. Streaming and Real Time Architecture Designing and implementing data ingestion and event driven patterns that support real time or near … using CI/CD principles adapted for Defence delivery. You will own your applications in production and contribute to secure patterns for deployment. Resilience, Observability and Compliance Implementing health monitoring, structured logging, metrics and lineage to meet Defence requirements for auditability, security and operational assurance. You will design systems that ...

Secure Data Engineer for Mission-Critical Analytics

Location
Gloucester, England, United Kingdom
support high‐performance analytical environments. Contribute to DevSecOps and cloud‐native delivery, including automated deployment, infrastructure provisioning and containerised environments. Support the ongoing operation, observability and resilience of production data platforms. Skills Required Experience in data engineering using modern programming languages such as Python, Java or Scala, and SQL. Hands … Familiarity with Infrastructure as Code and modern cloud‐native data architectures. Experience working in DevSecOps environments, including Docker, Kubernetes, CI/CD pipelines, and observability/monitoring tooling. Ability to work directly with stakeholders to understand data requirements and translate them into robust, scalable data engineering solutions. Comfortable working ...

Secure Data Engineer

Location
Newcastle upon Tyne, England, United Kingdom
components that process, transform and expose data for analytical, operational or AI driven use cases. You will follow strong engineering practices, with testing and observability built in from the start.**Streaming and Real Time Architecture** Designing and implementing data ingestion and event driven patterns that support real time or near … using CI/CD principles adapted for Defence delivery. You will own your applications in production and contribute to secure patterns for deployment.**Resilience, Observability and Compliance** Implementing health monitoring, structured logging, metrics and lineage to meet Defence requirements for auditability, security and operational assurance. You will design systems that ...

Network Automation Engineer

Location
Greater Manchester, England, United Kingdom
network provisioning and operational tasks Building and maintaining CI/CD pipelines Supporting AWS, Azure and hybrid cloud networking Improving platform reliability and observability Working within Agile engineering teams Collaborating with Architects, Platform Engineers and Security teams Driving continuous improvement and automation across network services Essential Skills Strong Infrastructure ...

Data Engineer - Security & Intelligence

Hiring Organisation
Hackajob Ltd
Location
Gloucester, Gloucestershire, South West, United Kingdom
Employment Type
Permanent
support high-performance analytical environments. Contribute to DevSecOps and cloud-native delivery, including automated deployment, infrastructure provisioning and containerised environments. Support the ongoing operation, observability and resilience of production data platforms. Skills Required Experience in data engineering using modern programming languages such as Python, Java or Scala, and SQL. Hands … Familiarity with Infrastructure as Code and modern cloud-native data architectures. Experience working in DevSecOps environments, including Docker, Kubernetes, CI/CD pipelines, and observability/monitoring tooling. Ability to work directly with stakeholders to understand data requirements and translate them into robust, scalable data engineering solutions. Comfortable working ...

Senior Fullstack Engineer (Python + React.js)

Location
Greater London, England, United Kingdom
minimal downtime. Write unit and integration tests to maintain code reliability and ensure high- quality releases. Continuously monitor and optimize backend performance using observability tools such as Datadog, Cloud Watch or similar. Participate in design discussions and decision-making to enhance system robustness and scalability. Maintain technical documentation to ensure … handling asynchronous communication. Experience with Infrastructure as Code (IaC) tools like Terraform or CloudFormation for managing cloud infrastructure. Knowledge of observability and monitoring tools, such as Cloud Watch or Datadog, to track and troubleshoot system performance. Familiarity with serverless architectures (e.g., AWS Lambda) and event‐-driven programming paradigms. Exposure ...

DevOps Team Manager

Hiring Organisation
Bromcom Computers Plc
Location
Bromley, London, United Kingdom
Employment Type
Permanent
technical quality while enabling engineers to own their work. Set and maintain engineering standards for Azure architecture, Azure DevOps, Bicep/ARM, deployment patterns, observability, resilience, security and operational support. Challenge designs and changes using risk, maintainability, failure-mode, rollback and supportability thinking; involve senior engineers and technical leadership where … access follows least-privilege principles, is reviewed regularly and is supported by effective joiner-mover-leaver, break-glass and segregation-of-duties controls. Own observability standards across Azure Monitor and Grafana, security and vulnerability follow-up, and cloud cost and FinOps accountability for the Azure estate. Stakeholder & Cross-Team Influence ...