251 to 275 of 926 Site Reliability Engineering Jobs in the UK

Monitoring & Observability Engineer (Dynatrace)

Hiring Organisation
Computacenter
Location
London, United Kingdom
Salary
£ 80 K
some of the world’s most well-known organisations. You’ll play a key role in helping our customers achieve greater visibility, performance, and reliability across their IT estates—contributing to their operational success through proactive insight and incident prevention.What you'll doDesign, implement, and manage observability solutions using … with a passion for continuous improvement and knowledge sharingCertificationsDynatrace Associate & ProSplunk Core Certified Power User Desirable ExperienceDevOps or Site Reliability Engineering (SRE) experienceAutomation with Terraform or similar toolsBuilding CI/CD pipelinesExperience with Docker and Kubernetes for packaging and deploymentAbility to adapt to new technologies in fast ...

Monitoring & Observability Engineer (Dynatrace)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
some of the world’s most well-known organisations. You’ll play a key role in helping our customers achieve greater visibility, performance, and reliability across their IT estates—contributing to their operational success through proactive insight and incident prevention. What you'll do Design, implement, and manage observability … passion for continuous improvement and knowledge sharing Certifications Dynatrace Associate & Pro Splunk Core Certified Power User DevOps or Site Reliability Engineering (SRE) experience Automation with Terraform or similar tools Experience with Docker and Kubernetes for packaging and deployment Ability to adapt to new technologies in fast-paced ...

Production Engineer

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
issues across our trading platform. You will leverage deep expertise in FIX, Linux, Windows Server, DevOps, databases, networking, and cloud technologies to ensure platform reliability and performance.This is a hands-on leadership role involving complex troubleshooting across cross-platform market-leading technologies, driving automation and tooling improvements, and acting … working hoursPositive approach to the day-to-day, with the resilience to handle high-pressure production incidentsDesiredExperience with Site Reliability Engineering (SRE) practices, including monitoring, incident response, and post-mortem analysisProven experience applying AI or machine-learning models to optimise workflows, identify patterns, and drive intelligent automation ...

Production Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
issues across our trading platform. You will leverage deep expertise in FIX, Linux, Windows Server, DevOps, databases, networking, and cloud technologies to ensure platform reliability and performance.This is a hands-on leadership role involving complex troubleshooting across cross-platform market-leading technologies, driving automation and tooling improvements, and acting … Positive approach to the day-to-day, with the resilience to handle high-pressure production incidentsDesired* Experience with Site Reliability Engineering (SRE) practices, including monitoring, incident response, and post-mortem analysis* Proven experience applying AI or machine-learning models to optimise workflows, identify patterns, and drive intelligent ...

Production Engineer

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
issues across our trading platform. You will leverage deep expertise in FIX, Linux, Windows Server, DevOps, databases, networking, and cloud technologies to ensure platform reliability and performance. This is a hands-on leadership role involving complex troubleshooting across cross-platform market-leading technologies, driving automation and tooling improvements … approach to the day-to-day, with the resilience to handle high-pressures production incidents Desired Experience with Site Reliability Engineering (SRE) practices, including monitoring, incident response, and post-mortem analysis Proven experience applying AI or machine-learning models to optimise workflows, identify patterns, and drive intelligent ...

Senior Lead SRE: Reliability, Observability & Resiliency

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
JPMorgan Chase & Co. is seeking a Senior Lead Site Reliability Engineer in Glasgow, Scotland. This role is pivotal for enhancing the reliability and observability of critical platforms. You will lead technical initiatives and contribute significantly to business impact through your expertise. The ideal candidate will have advanced … proficiency in software engineering, reliability best practices, and a strong background in programming languages. This position demands a blend of technical acumen and leadership in a collaborative environment focused on delivering trusted technology products. #J-18808-Ljbffr ...

Lead Software Engineer - Application Owner & Release Manager

Hiring Organisation
JP Morgan Chase
Location
Glasgow, Lanarkshire, United Kingdom
Salary
£ 80 K
Build solutions that matter in a highly regulated environment where resiliency and security are as important as innovation. You’ll partner across engineering, cyber, risk, and resiliency to keep critical analytics platforms healthy and compliant while enabling teams to deliver change safely. This role offers breadth across cloud, data … software engineering, plus the opportunity to lead through influence and strong execution. You’ll help reduce toil through automation and create space for engineers and data scientists to do their best work. Join a team that values inclusion, growth, and pragmatic problem-solving.As a Lead Software Engineer, Application Owner ...

Lead Data Platform Engineer - DataOps

Hiring Organisation
Jobleads-UK
Location
Nottingham, England, United Kingdom
Lead Data Platform Engineer to own the resilience, scalability, and security of our core financial data ecosystem. This is a senior, hands‐on engineering role for a technical expert who thrives on the unique challenge of designing, building, and running high‐throughput, mission‐critical batch and real‐time data … across a distributed team while contributing to the long‐term technical strategy and cloud modernization efforts. What we’re looking for A Platform or SRE Background: You have a solid background in platform engineering, site reliability (SRE), or data operations, ideally within a large‐scale or regulated ...

Lead Data Platform Engineer - DataOps

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Lead Data Platform Engineer to own the resilience, scalability, and security of our core financial data ecosystem. This is a senior, hands-on engineering role for a technical expert who thrives on the unique challenge of designing, building, and running high-throughput, mission-critical batch and real-time data … across a distributed team while contributing to the long-term technical strategy and cloud modernization efforts.****What we’re looking for***** A Platform or SRE Background: You have a solid background in platform engineering, site reliability (SRE), or data operations, ideally within a large-scale or regulated ...

Site Reliability Engineer, Infrastructure - ThousandEyes

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
deeply integrated across the Cisco technology portfolio, delivering AI‐powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios. Our distributed Site Reliability Engineering team of approximately nine engineers owns the availability, latency, performance, efficiency, monitoring, emergency response, and capacity planning of the platform while … call rotation. Hands‐on experience with infrastructure‐as‐code tooling and codebases, preferably Terraform. Hands‐on experience leveraging AI as a force multiplier of SRE activities, such as automating toil away and improving operational efficiency. Professional experience administering and troubleshooting GNU/Linux systems, including system libraries, file systems, networking ...

Site Reliability Engineer, Infrastructure - ThousandEyes

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
deeply integrated across the Cisco technology portfolio, delivering AI-powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios. Our distributed Site Reliability Engineering team of approximately nine engineers owns the availability, latency, performance, efficiency, monitoring, emergency response, and capacity planning of the platform while … operational on-call rotation. Hands-on experience with infrastructure-as-code tooling and codebases, preferably Terraform. Hands-on experienceleveraging AIas a force multiplier of SRE activities, such as automati ng toil away and improving operational efficiency. Professional experience administering and troubleshooting GNU/Linux systems, including system libraries, file systems ...

Cloud Operations Engineer (remote – London)

Hiring Organisation
Quant Capital
Location
London, United Kingdom
Salary
£ 80 K
remote) Cloud Operations Engineer/Site Reliability Engineer – Fintech80,000 Plus Bonus + 10% non-cont pension + 10-15k bonus and sharesQuant Capital is urgently looking for a Site Reliability Engineer to join or well-known Fintech50 client who produces software disrupting the wealth … TechniquesSolid understanding of the OSI ModelExperience in database technology and basic query writing MSSQL, Postgres.This role suits a senior Engineer from a DevOps or SRE background who is a real technologist and cloud specialist interested in the latest tooling and technologies that support software development and infrastructure. The firm ...

Site Reliability Engineer – NS London

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Location(s): [[mfield3]] The role of a Site Reliability Engineer (SRE) at BAE Systems Digital Intelligence involves combining operations and software engineering to automate system support and enhance reliability for a key national security customer. The SRE team works on continuous improvement of system health … deploy monitoring products, creating custom tools as needed to provide comprehensive, intelligent observations that demonstrate daily improvements. Participate in the wider DevOps/SRE community within the organization. Qualifications Experience in web development and object‐oriented programming. Knowledge of database technologies such as Oracle SQL, MongoDB, and PostgreSQL. Proficiency with ...

AMBG - Cloud Security & Exposure Management Architect

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
recovery sequencing. Identify risks, vulnerabilities, and single points of failure across workloads and operational processes. Recommend improvements aligned with Azure Well-Architected Framework, SRE principles, and ITIL practices. Engage customer stakeholders to understand RTO/RPO objectives and recovery workflows. Produce professional documentation outlining findings, risks, and recommended improvements. About … Azure architecture including availability zones, backup, recovery, and monitoring services. Familiarity with cloud-native resiliency patterns and site reliability engineering (SRE) practices. Experience designing and assessing Major Incident Response Plans (MIRPs). Experience in business continuity planning and operational resilience. Strong communication and documentation skills across technical ...

Senior DevOps / Platform Engineer (Google Cloud)

Hiring Organisation
Datatonic
Location
London, United Kingdom
Salary
£ 80 K
Cloud's premier partner in AI, driving transformation for world-class businesses. We push the boundaries of technology with expertise in machine learning, data engineering, and analytics on Google Cloud Platform. By partnering with us, clients future-proof their operations, unlock actionable insights, and stay ahead of the curve … scale-up environmentContainerisation/Virtualisation Expertise: Proficiency with technologies such as Terraform and KubernetesSRE Principles: Experience in implementing Site Reliability Engineering (SRE) principlesCloud Native Architecture: Hands-on experience with cloud-native architectures, ideally on Google CloudClient-Facing Role: Prior experience in a client-facing positionSDN Knowledge: Understanding ...

Technical Account Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Verda. The role works in two directions: outward, as the customer's trusted technical advisor, and inward, as their advocate inside Verda's engineering, infrastructure, and operations teams. Outward (customer-facing): Serve as the primary technical point of contact for assigned customers, building trusted relationships with their ML, research … facing role such as pre-sales solutions architecture, post-sales technical account management, solution architecture, or customer-oriented site reliability engineering (SRE). Practical understanding of the HPC/AI stack: GPU compute, job schedulers (e.g., Slurm, Kubernetes), high-performance networking (InfiniBand/RDMA), and parallel ...

CDSClear IT Site Reliability Engineer

Hiring Organisation
London Stock Exchange Group
Location
London, United Kingdom
Salary
£ 80 K
hours on‐call rota, ensuring continuity and resilience of critical clearing services.Support extended clearing operations as part of a late‐shift Site Reliability Engineering rota (up to 10:30 pm London time) to accommodate increased US trading hours, with scope evolving as the global support model matures.Be … stability and operational efficiency. Contribute to IT Change and Governance forums to prioritise enhancements supported by clear cases.Collaborate with development teams to improve release reliability and automation, using the approved LSEG DevOps toolset.Lead capacity and performance management, enhancing monitoring capabilities to ensure platforms scale in line with business growth ...

Data Platform Engineer

Hiring Organisation
MONY Group
Location
London, United Kingdom
Salary
£ 80 K
personalised customer experiences. We work closely with teams across the business to make data clean, reliable, secure and accessible for decision-making. Data & AI Engineering is a cross-functional team of engineers and scientists. We integrate with the group's operational data stores, maintain shared data models, build … ability to apply automation responsibly to real delivery and operational problems. You might come from data engineering, platform engineering, software engineering, SRE, analytics engineering, MLOps, or cloud infrastructure. What matters most is that you enjoy reducing toil, improving developer experience, and building secure, observable systems that ...

Network Site Reliability Engineer

Hiring Organisation
Quant Capital
Location
London, United Kingdom
Salary
£ 100 K
Network SRE – 250,000-350,000 total compensation – 4 days in officeQuant Capital is urgently looking Network SRE for our high profile client.Our client is a leading quantitative trading company and liquidity provider. Their focus on technology has allowed them to deeply penetrate the market and gain market share. … Shared Engineering team that focuses on designing, developing, and maintaining infrastructure and tools. The team requires a Network Site Reliability Engineer (SRE) with strong network fundamentals, problem-solving skills, and a keen interest in diverse tools and techniques. The role involves collaborative work across various teams, exploring ...

Systems Operations Lead

Hiring Organisation
Hays Technology
Location
City of London, London, United Kingdom
Employment Type
Contract
Contract Rate
£750 - £800/day Up to £800pd inside ir35 via umbrella
team of technical SMEs, ensuring workloads are prioritised and delivered effectively. Act as an escalation point for operational and infrastructure-related issues. Drive service reliability, operational excellence and continuous improvement across the environment. Required Experience Strong infrastructure background with … experience across Linux and Windows server environments. Good understanding of storage, backup and wider infrastructure technologies. Experience in Site Reliability Engineering (SRE), Infrastructure Operations, or Production Support environments. Proven experience leading and developing technical teams. Comfortable remaining hands-on and involved in technical delivery on a daily ...

Site Reliability Engineer

Hiring Organisation
UK Tote Group
Location
Wigan, Greater Manchester, United Kingdom
Salary
£ 55 K
mission to deliver a seamless and reliable digital experience for racing fans across the UK and beyond. As a Site Reliability Engineer (SRE), you’ll play a critical role in keeping our online platforms and infrastructure fast, stable, and scalable — especially during the most exciting moments … performance and stability. You’ll analyse telemetry data, identify bottlenecks, and drive improvements across our infrastructure and applications.You’ll lead the development of our SRE strategy, defining standards, best practices, and ways of working that embed reliability into everything we build. Working closely with engineering, operations, and product ...

Site Reliability Engineer - Banking & Finance

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Ready to take the next step in your career? Join a leading technology-driven trading firm where engineering, automation, and high-performance infrastructure are central to supporting global trading operations. The organisation invests heavily in modern platform engineering practices, enabling teams to build reliable, scalable, and highly automated … Have: Strong experience programming with Python, Go and/or C++ Strong Linux knowledge and understanding of distributed systems. Experience with monitoring, observability or SRE practices. Experience with CI/CD pipelines, Git and infrastructure automation. Familiarity with Kubernetes and containerised workloads. Strong analytical and troubleshooting skills. Benefits: Build ...

Pre-Sales Solutions Architect (Cloud / AI Managed Services)

Hiring Organisation
ThoughtWorks
Location
London, United Kingdom
Salary
£ 80 K
practices like XP and CI/CD to achieve "zero maintenance" products, revolutionizing how Run operates. By combining site reliability engineering (SRE), product evolution and data ops, DAMO managed services drive predictable cost reduction and future-proof operations.Principal solutions architects are a driving force in our transformative … best and that extends to empowering our employees in their career journeys.About ThoughtworksThoughtworks is a global technology consultancy that integrates strategy, design and engineering to drive digital innovation. For 30 years, our clients have trusted our autonomous teams to build solutions that look past the obvious. Here, computer science ...

Senior DevOps Engineer

Hiring Organisation
Anson Mccade
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£75,000
Engineer Location: Manchester (Client Travel Required) Why Join This Senior DevOps Engineer Opportunity? Join a leading technology consultancy delivering large-scale cloud and platform engineering solutions for some of the world's most recognisable organisations. You'll work with modern cloud technologies, influence engineering best practices and play … Engineer Design and support cloud-native platforms across AWS and Azure Implement DevSecOps and Infrastructure as Code best practices Drive platform reliability using SRE principles and observability tools Support incident management and continuous improvement initiatives Implement Terraform-based infrastructure solutions Leverage automation and AI-assisted engineering tools ...

IAM Secrets Management Engineering - SRE Platform Engineer - VP - London

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Role We are seeking a skilled and experienced lead Site Reliability Platform Engineer (SRE) to join our team. The ideal candidate will be responsible for ensuring the reliability, performance, and scalability of mission‐critical, high‐availability, high‐throughput systems and infrastructure. This role involves leading and collaborating … with cross‐functional teams, and implementing best practices in SRE, DevOps, and cyber security to enhance our operational efficiency and security posture. System Reliability, Performance, and Security Design, implement, and maintain highly available, scalable, and secure systems on AWS cloud services. Monitor system performance, reliability, and security, proactively ...