101 to 125 of 538 Site Reliability Engineering Jobs in London

Principal Site Reliability Engineer, Infrastructure Observability

Location
Greater London, England, United Kingdom
toolchain and systems, code build and deployment, incident response, and 24x7 monitoring and support. The candidate will also have extensive experience operating within a SRE function within a complex, distributed environment. They will have a demonstrated ability to work horizontally and vertically within an organization with diverse partners and sponsor … learning through blameless post-mortems to improve the shared goal of reliability across services Transform operations teams by facilitating internal change to adopt SRE standard methodologies across the organization and driving strategic growth in this area within Global Technology Analyzes incidents impacting technology availability for high-level trends across ...

Product Associate - SRE Team - Chase UK

Location
Greater London, England, United Kingdom
oriented and possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with … engineering teams to ensure services are designed, delivered, and operated with reliability in mind. Job responsibilities Support the product strategy and delivery of reliability capabilities, including standards, observability, incident practices, automation, and developer experience improvements. Partner with engineers, site reliability engineers, and cross‐functional teams ...

Product Associate - SRE Team - Chase UK

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
oriented and possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with … engineering teams to ensure services are designed, delivered, and operated with reliability in mind. Job responsibilities Support the product strategy and delivery of reliability capabilities, including standards, observability, incident practices, automation, and developer experience improvements. Partner with engineers, site reliability engineers, and cross-functional teams ...

Product Associate - SRE Team - Chase UK

Location
Westminster, West End, United Kingdom
oriented and possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with … engineering teams to ensure services are designed, delivered, and operated with reliability in mind. Job responsibilities Support the product strategy and delivery of reliability capabilities, including standards, observability, incident practices, automation, and developer experience improvements. Partner with engineers, site reliability engineers, and cross-functional teams ...

DevOps Engineer (Security Cleared)

Hiring Organisation
Solirius Consulting
Location
London, UK
Employment Type
Full-time
Solirius Reply, part of the Reply Group, is a technology consultancy and digital transformation partner that helps organisations solve complex challenges through strategy, design, engineering, and delivery. We work closely with our clients to deliver secure, accessible, user-focused services that evolve with their needs. By combining deep technical … Ministry of Housing, Communities and Local Government, UEFA, International Olympic Committee, and Mercedes-Benz. Our services span the full digital delivery lifecycle, including architecture, engineering, delivery management, user-centred design, business analysis, data, DevOps, and AI.We operate as a collaborative and inclusive organisation that empowers our people to take ...

DevOps Engineer (Security Cleared)

Location
Greater London, England, United Kingdom
Solirius Reply, part of the Reply Group, is a technology consultancy and digital transformation partner that helps organisations solve complex challenges through strategy, design, engineering, and delivery. We work closely with our clients to deliver secure, accessible, user-focused services that evolve with their needs. By combining deep technical … Ministry of Housing, Communities and Local Government, UEFA, International Olympic Committee, and Mercedes-Benz. Our services span the full digital delivery lifecycle, including architecture, engineering, delivery management, user-centred design, business analysis, data, DevOps, and AI. We operate as a collaborative and inclusive organisation that empowers our people ...

Senior DevOps / Platform Engineer (Google Cloud)

Location
Greater London, England, United Kingdom
Cloud's premier partner in AI, driving transformation for world-class businesses. We push the boundaries of technology with expertise in machine learning, data engineering, and analytics on Google Cloud Platform. By partnering with us, clients future-proof their operations, unlock actionable insights, and stay ahead of the curve … Experience: Previous experience working in a start-up or scale-up environment Containerisation/Virtualisation Expertise: Proficiency with technologies such as Terraform and Kubernetes SRE Principles: Experience in implementing Site Reliability Engineering (SRE) principles Cloud Native Architecture: Hands-on experience with cloud-native architectures, ideally on Google ...

Site Reliability Software Engineer (Hybrid)

Location
Greater London, England, United Kingdom
shapes without surgery. With over 600,000+ successful outcomes, EarWell® is a proven, non-invasive treatment option for families. We are looking for a Site Reliability Engineer to join our growing team. The ideal candidate has a strong technical background in software development and systems operations, with experience … security, privacy, and compliance requirements in mind. Document architecture, workflows, troubleshooting steps, deployment processes, and support procedures. Required Qualifications Experience as a Software Engineer, Site Reliability Engineer, DevOps Engineer, Systems Engineer, or similar technical role. Experience supporting production applications or infrastructure in a business-critical environment. Strong scripting ...

Senior Azure Platform Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City of London, London, United Kingdom
Employment Type
Permanent
Platform We're partnering with one of London's most exciting and rapidly growing fintech companies as they continue to invest in their Platform Engineering function. They're looking for a Senior Platform Engineer who wants more than just another engineering role. This is an opportunity to help … shape the platform, influence technical decisions, and build the cloud infrastructure that powers a rapidly scaling business. You'll join a collaborative engineering team where your ideas are encouraged, ownership is expected, and you'll play a key role in driving automation, reliability, and platform maturity. ...

Staff SRE, AI Infrastructure

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
each other to deliver impact. Make Wayve the experience that defines your career! The roleThis is a rare opportunity to be a founding Staff SRE shaping the reliability of large-scale AI systems and GPU compute infrastructure from the ground up. As a Staff Cloud Site Reliability … Compute platform (large-scale, multi-tenant GPU fleets and scheduling systems driving model training and inference at scale).This is a founding Cloud SRE role. You won't inherit a mature SRE function, you'll help create it. You will define the frameworks, automation, and operational standards that ensure ...

Chief Platform Architect - 801

Location
Greater London, England, United Kingdom
/classical application components Define standard interfaces for hosting application components and building applications by both internal and external developers Ensure the security and reliability of the platform as it scales up Deliver customizable cloud-based hybrid applications to end users Work across the business to help define … architecture of developer platforms Experience in building platform engineering teams from the ground up. Background in Site Reliability Engineering (SRE). Record of publication or patents in quantum computing. $208,000 - $260,000 a year Incentive Eligible – Range posted is inclusive of bonus target when applicable ...

DevOps Engineer

Location
City Of London, England, United Kingdom
London (4-5 days a week in office) This is an opportunity for a DevOps Engineer to work at the intersection of infrastructure engineering and AI technology within a high-performance environment. You will play a key role in building and scaling modern infrastructure platforms, with a particular focus … premise environments that support business-critical workloads. THE COMPANY They are a globally operating investment and technology-driven organisation with a strong engineering culture. Their teams work closely with technical and business stakeholders to deliver robust, scalable infrastructure across a complex environment. This role offers exposure to cutting-edge ...

DevOps Engineer

Hiring Organisation
Harnham - Data & Analytics Recruitment
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£100,000 - £120,000 per annum
London (4-5 days a week in office) This is an opportunity for a DevOps Engineer to work at the intersection of infrastructure engineering and AI technology within a high-performance environment. You will play a key role in building and scaling modern infrastructure platforms, with a particular focus … premise environments that support business-critical workloads. THE COMPANY They are a globally operating investment and technology-driven organisation with a strong engineering culture. Their teams work closely with technical and business stakeholders to deliver robust, scalable infrastructure across a complex environment. This role offers exposure to cutting-edge ...

Senior Linux DevOps Engineer

Hiring Organisation
RedTech Recruitment Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
annum + excellent benefits Requirements for Senior Linux DevOps Engineer: Strong commercial experience working as a Senior DevOps Engineer, Linux Engineer, Platform Engineer, Site Reliability Engineer or similar Excellent Linux systems administration and command line skills, with experience operating and troubleshooting large-scale production environments Strong scripting … Linux Engineer/Linux Systems Engineer/Linux Infrastructure Engineer/Senior Platform Engineer/Platform Engineer/Site Reliability Engineer/SRE/Infrastructure Engineer/DevSecOps Engineer/Linux/Bash/Shell Scripting/Python/Kubernetes/Docker/Terraform/Ansible/Microsoft ...

Core AI Engineer

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
Artificial Intelligence, Automation and Intelligent Engineering. We are building enterprise-scale AI capabilities that improve service resilience, automate operational workflows, accelerate engineering productivity and enhance customer outcomes. As a Core AI Engineer, you will play a leading technical role in the design, development and deployment of AI solutions across … systems design. AI observability, evaluation and governance frameworks. Desirable ExperienceExperience within Financial Services or highly regulated environments. Knowledge of Service Reliability Engineering (SRE) principles. Experience developing AI-powered operational tooling. Experience building internal AI platforms or developer enablement capabilities. Familiarity with Microsoft AI ecosystem, Copilot technologies and Azure ...

Monitoring & Observability Engineer (Dynatrace)

Hiring Organisation
Computacenter
Location
London, UK
Employment Type
Full-time
some of the world's most well-known organisations. You'll play a key role in helping our customers achieve greater visibility, performance, and reliability across their IT estates—contributing to their operational success through proactive insight and incident prevention. What you'll doDesign, implement, and manage observability solutions … with a passion for continuous improvement and knowledge sharingCertificationsDynatrace Associate & ProSplunk Core Certified Power User Desirable ExperienceDevOps or Site Reliability Engineering (SRE) experienceAutomation with Terraform or similar toolsBuilding CI/CD pipelinesExperience with Docker and Kubernetes for packaging and deploymentAbility to adapt to new technologies in fast ...

Monitoring & Observability Engineer (Dynatrace)

Location
Greater London, England, United Kingdom
some of the world’s most well-known organisations. You’ll play a key role in helping our customers achieve greater visibility, performance, and reliability across their IT estates—contributing to their operational success through proactive insight and incident prevention. What you'll do Design, implement, and manage observability … passion for continuous improvement and knowledge sharing Certifications Dynatrace Associate & Pro Splunk Core Certified Power User DevOps or Site Reliability Engineering (SRE) experience Automation with Terraform or similar tools Experience with Docker and Kubernetes for packaging and deployment Ability to adapt to new technologies in fast-paced ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
Overview CTI is seeking to appoint a head of engineering for a fintech scaleup. Requirements Proficient in AWS services such as ECS, EC2, Lambda, VPC, IAM, Route53, CloudFront, S3, and RDS. Solid … understanding of monitoring and logging tools, including Prometheus, AWS CloudWatch, Grafana, OpenTelemetry, Honeycomb, and ELK. Basic knowledge of Site Reliability Engineering (SRE) and experience with alerting and incident management systems like Opsgenie and PagerDuty. Demonstrated capability to develop and maintain robust and scalable Continuous Integration/Continuous ...

Production Engineer

Location
City Of London, England, United Kingdom
issues across our trading platform. You will leverage deep expertise in FIX, Linux, Windows Server, DevOps, databases, networking, and cloud technologies to ensure platform reliability and performance.This is a hands-on leadership role involving complex troubleshooting across cross-platform market-leading technologies, driving automation and tooling improvements, and acting … working hoursPositive approach to the day-to-day, with the resilience to handle high-pressure production incidentsDesiredExperience with Site Reliability Engineering (SRE) practices, including monitoring, incident response, and post-mortem analysisProven experience applying AI or machine-learning models to optimise workflows, identify patterns, and drive intelligent automation ...

Production Engineer

Location
Greater London, England, United Kingdom
issues across our trading platform. You will leverage deep expertise in FIX, Linux, Windows Server, DevOps, databases, networking, and cloud technologies to ensure platform reliability and performance.This is a hands-on leadership role involving complex troubleshooting across cross-platform market-leading technologies, driving automation and tooling improvements, and acting … Positive approach to the day-to-day, with the resilience to handle high-pressure production incidentsDesired* Experience with Site Reliability Engineering (SRE) practices, including monitoring, incident response, and post-mortem analysis* Proven experience applying AI or machine-learning models to optimise workflows, identify patterns, and drive intelligent ...

Production Engineer

Location
City Of London, England, United Kingdom
issues across our trading platform. You will leverage deep expertise in FIX, Linux, Windows Server, DevOps, databases, networking, and cloud technologies to ensure platform reliability and performance. This is a hands-on leadership role involving complex troubleshooting across cross-platform market-leading technologies, driving automation and tooling improvements … approach to the day-to-day, with the resilience to handle high-pressures production incidents Desired Experience with Site Reliability Engineering (SRE) practices, including monitoring, incident response, and post-mortem analysis Proven experience applying AI or machine-learning models to optimise workflows, identify patterns, and drive intelligent ...

AI Platform & SRE Leader — Scale & Govern AI

Location
Greater London, England, United Kingdom
Reliability Engineering Managing Consultant to help clients design, build and scale secure, reliable AI platforms. You will lead platform engineering, SRE, observability and guardrails, moving AI from experimentation to enterprise-scale delivery. You will partner with CIO/CTO stakeholders, shape platform strategies, and guide multidisciplinary … delivery teams, balancing reliability, innovation and cost in a global-capability context. #J-18808-Ljbffr ...

Site Reliability Engineer, Infrastructure - ThousandEyes

Location
City Of London, England, United Kingdom
deeply integrated across the Cisco technology portfolio, delivering AI-powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios. Our distributed Site Reliability Engineering team of approximately nine engineers owns the availability, latency, performance, efficiency, monitoring, emergency response, and capacity planning of the platform while … operational on-call rotation. Hands-on experience with infrastructure-as-code tooling and codebases, preferably Terraform. Hands-on experienceleveraging AIas a force multiplier of SRE activities, such as automati ng toil away and improving operational efficiency. Professional experience administering and troubleshooting GNU/Linux systems, including system libraries, file systems ...

Cloud Operations Engineer (remote - London)

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
remote) Cloud Operations Engineer/Site Reliability Engineer – Fintech80,000 Plus Bonus + 10% non-cont pension + 10-15k bonus and sharesQuant Capital is urgently looking for a Site Reliability Engineer to join or well-known Fintech50 client who produces software disrupting the wealth … understanding of the OSI ModelExperience in database technology and basic query writing MSSQL, Postgres. This role suits a senior Engineer from a DevOps or SRE background who is a real technologist and cloud specialist interested in the latest tooling and technologies that support software development and infrastructure. The firm ...

AMBG - Cloud Security & Exposure Management Architect

Location
Greater London, England, United Kingdom
recovery sequencing. Identify risks, vulnerabilities, and single points of failure across workloads and operational processes. Recommend improvements aligned with Azure Well-Architected Framework, SRE principles, and ITIL practices. Engage customer stakeholders to understand RTO/RPO objectives and recovery workflows. Produce professional documentation outlining findings, risks, and recommended improvements. About … Azure architecture including availability zones, backup, recovery, and monitoring services. Familiarity with cloud-native resiliency patterns and site reliability engineering (SRE) practices. Experience designing and assessing Major Incident Response Plans (MIRPs). Experience in business continuity planning and operational resilience. Strong communication and documentation skills across technical ...