226 to 250 of 597 Site Reliability Engineering Jobs in London

Cloud Operations Manager

Location
City Of London, England, United Kingdom
ensuring timely restoration, clear communication and effective root-cause resolution. Maintain operational controls, standards, runbooks and support models aligned with IT service management and engineering practices. Ensure cloud environments comply with security, risk, audit, data protection and regulatory requirements. Own and maintain vulnerability management processes to protect the estate. … Experience Relevant cloud certification at associate or professional level. ITIL qualification or equivalent service-management experience. Experience with infrastructure as code, DevOps practices, site reliability engineering or platform engineering. Knowledge of AWS, Azure, Salesforce & Oracle Cloud Experience operating services in a regulated or safety-critical environment. Knowledge ...

Principal Platform Engineer

Hiring Organisation
Sanderson Recruitment
Location
City of London, London, United Kingdom
Employment Type
Permanent
large-scale distributed systems and database platforms? We're looking for a hands-on technical leader to help shape the future of our platform engineering capability. This is an opportunity to lead complex engineering initiatives, define technical strategy, and act as a subject matter expert across AWS infrastructure … technical authority for distributed database and persistence technologies Required Experience 8+ years' experience in Platform Engineering, Infrastructure Engineering, DevOps, SRE or Software Engineering Expert-level AWS infrastructure experience Strong Infrastructure as Code expertise with Terraform Strong Linux systems administration and networking knowledge Experience designing and operating distributed ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
diversified range of trading strategies. We employ over 130 colleagues in Jersey, Geneva, London, Singapore, New York and Shanghai. This is a hands‐on engineering role within the Platform Engineering team, part of Technology Operations. Platform Engineering is responsible for building and operating the infrastructure, platforms … comply with all organisational, statutory and regulatory policies and procedures. Experience, Knowledge & Skills Five or more years of experience in platform engineering, DevOps, SRE, infrastructure engineering or a closely related role. Strong experience operating production or production‐like infrastructure, ideally across hybrid cloud and on‐premises environments. Hands ...

Trading Systems Engineer, Trading Platform

Location
Greater London, England, United Kingdom
communicate directly with exchanges, traders and developers to iron out any problems that arise. Qualifications & Skills: Minimum of 5 years working in trade support, site reliability engineering or related fields Bachelor’s degree in STEM or related field Familiarity with trading platforms and financial markets Thrives … high-pressure situations while working alongside traders, developers and other engineering teams Strong problem-solving skills and the ability to troubleshoot technical issues under pressure Excellent communication skills, both written and verbal Knowledge of Linux/Unix environments Experience with scripting languages such as Python and Bash for automation ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
iGaming company based in London. Estimated Benefits Health Insurance Pension Stock Options Benefits estimated based on industry standards We’re hiring a Site Reliability Engineer to join our London team This is a fantastic opportunity for someone passionate about reliability, scalability and automation. You’ll be pivotal ...

Cloud Operations Engineer

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Cloud Operations Engineer – FintechUp to 90,000 + 12% pension, bonuses and benefitsQuant Capital is urgently looking for a Site Reliability Engineer to join or well-known Fintech50 client who produces software disrupting the wealth management market. My client is a market leading SAAS provider to financial advisory … remediate with stable solutionsJenkinsLinux, CentOMS SQLFamiliarity with development languages, such as .NET, Java or PythonThis role suits a senior Engineer from a DevOps or SRE background who is real technologist and cloud specialist interested in the latest tooling and technologies that support software development and infrastructure. The firm ...

Solutions Architect, Studios

Location
Uxbridge, England, United Kingdom
procurement activity, with particular emphasis on pre sales engagement, proof of concept activity, and technical validation. This role will work across commercial, technical, operational, engineering, procurement, and transformation teams to translate business opportunities into clear solution options, technical direction, and delivery‐ready architecture. A core part of the role … conversations, and leading or coordinating proof of concepts that demonstrate feasibility, value, and operational fit. The role will also work closely with DevOps and Site Reliability Engineering teams to ensure proposed solutions are automatable, observable, supportable, resilient, and aligned to modern operational practices. The role will play ...

Solutions Architect, Studios

Location
Greater London, England, United Kingdom
procurement activity, with particular emphasis on pre sales engagement, proof of concept activity, and technical validation. This role will work across commercial, technical, operational, engineering, procurement, and transformation teams to translate business opportunities into clear solution options, technical direction, and delivery-ready architecture. A core part of the role … conversations, and leading or coordinating proof of concepts that demonstrate feasibility, value, and operational fit. The role will also work closely with DevOps and Site Reliability Engineering teams to ensure proposed solutions are automatable, observable, supportable, resilient, and aligned to modern operational practices. The role will play ...

Cloud Platform Engineer

Location
Greater London, England, United Kingdom
identity, network, workloads and data for both human and machine/agent identities; secure by default, least privilege, secrets management and continuous compliance. Apply SRE practices - SLOs/SLIs, observability, capacity planning, resilience and blameless incident management - to keep the platform reliable and cost-efficient. Partner with data engineering … ability (e.g. Python, Go) and strong observability, reliability and cost‐optimisation practices. Desirable requirements: Experience working as a Site Reliability Engineer (SRE) with SLOs/SLIs, error budgets and incident management. A third top‐tier cloud certification, or specialist security/Kubernetes certifications (e.g. CKA/ ...

Site Reliability Engineer - AI-Driven, Scalable Cloud Systems

Location
City Of London, England, United Kingdom
Cisco ThousandEyes is seeking an experienced Site Reliability Engineer to design and operate large-scale distributed systems that process growing telemetry data. You will apply AI tooling to write high-quality code and automated solutions to enable fast, reliable releases across regions. You will evaluate scalability, resiliency, performance ...

Trading Systems Engineer, Trading Platform

Hiring Organisation
DRW
Location
London, UK
Employment Type
Full-time
communicate directly with exchanges, traders and developers to iron out any problems that arise. Qualifications & Skills: Minimum of 5 years working in trade support, site reliability engineering or related fieldsBachelor's degree in STEM or related fieldFamiliarity with trading platforms and financial marketsThrives in high-pressure situations … while working alongside traders, developers and other engineering teamsStrong problem-solving skills and the ability to troubleshoot technical issues under pressureExcellent communication skills, both written and verbalKnowledge of Linux/Unix environmentsExperience with scripting languages such as Python and Bash for automation tasksAbility to devise complex SQL database queries ...

AI-Powered Production Reliability Engineer

Location
Greater London, England, United Kingdom
Cisco Systems, Inc is seeking a Software Engineer to design and deploy AI-powered production intelligence capabilities at scale. You will combine Site Reliability Engineering with agentic AI to monitor, diagnose, and auto-remediate global SaaS infrastructure. You will build AI agents, integrate MCP tooling, and implement … telemetry across distributed systems, driving reliability and rapid incident response. #J-18808-Ljbffr ...

DevOps Engineer – Security & Intelligence

Location
Greater London, England, United Kingdom
secure platforms that underpin mission-critical digital services. You'll operate within multi-disciplinary Agile teams, collaborating closely with software engineers, test engineers, architects, SRE specialists and mission stakeholders to ensure robust, production-grade delivery. This role is suited to someone who thrives in complex, secure environments and enjoys working … automated platform capabilities Supporting AWS-based environments, including Kubernetes, OpenShift, EKS and ECS Implementing observability, monitoring, logging and alerting for live services Supporting SRE practices, cloud migration activities and production platform operations Job Responsibilities Design, implement and maintain secure CI/CD pipelines to support efficient, automated software delivery Develop ...

Hybrid Linux Desktop Engineer (SRE) - London

Location
Greater London, England, United Kingdom
MLabs Ltd. in London is seeking a Linux Desktop Support Engineer to join the Site Reliability Engineering team. The role focuses on managing internal devices, supporting workplace technology, and maintaining office infrastructure with a strong emphasis on Ubuntu/Linux environments. You will own provisioning, onboarding ...

Junior Trading Support Engineer - Prop Trading

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
incidents. Required Skills & Experience: Experience supporting mission critical systems and high performance applications Minimum of 1 years working in trade support, app support or site reliability engineering Bachelor's degree in STEM or related field Have exposure to VCS, particularly Git/Github Have demonstrated scripting abilities … Comfortable with Linux and the command-line Financial Services experience Someone who thrives in high-pressure situations while working alongside traders, developers and other engineering teams Experience developing proprietary process automation and monitoring tools to streamline software configuration and rollout procedures The environment is that of Facebook or Google ...

Trading Systems Engineer - Prop Trading

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
bleeding edge. Required Skills & Experience: Experience supporting mission critical systems and high performance applications Minimum of 5 years working in trade support, site reliability engineering Bachelor's degree in STEM or related field Familiarity with trading platforms and financial markets Thrives in high-pressure situations while working … alongside traders, developers and other engineering teams Strong problem-solving skills and the ability to troubleshoot technical issues under pressure Knowledge of Linux/Unix environments Experience with scripting languages such as Python and Bash for automation tasks Ability to devise complex SQL database queries and updates Basic networking ...

Senior Network SRE: Cloud Reliability & IaC

Location
Greater London, England, United Kingdom
Miro is seeking a Senior Network Site Reliability Engineer to help strengthen reliability, availability, and scalability of our production environment. You will focus on cloud automation, IaC, and governance across our AWS infra, contributing to highly available services for millions of users. You will own automation, observability ...

Senior Engineering Manager (SRE)

Location
Greater London, England, United Kingdom
do. Imagine what you could do here! Join Apple, and help us leave the world better than we found it. The Apple Services Engineering (ASE) team builds and provides systems and infrastructure that power Apple’s services (such as iCloud, Apple Music, Apple TV, Apple Intelligence, and Maps). … response and operational excellence, while continuously identifying opportunities to make the on-call experience better for the team Minimum Qualifications Extensive Leadership in Kubernetes & SRE: In-depth experience building and leading high-performing engineering teams, with a deep focus on Kubernetes, hands-on experience writing and supporting production software ...

Application Support Engineer

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
hours on‐call rota, ensuring continuity and resilience of critical clearing services. Support extended clearing operations as part of a late‐shift Site Reliability Engineering rota (up to 10:30 pm London time) to accommodate increased US trading hours, with scope evolving as the global support model … stability and operational efficiency. Contribute to IT Change and Governance forums to prioritise enhancements supported by clear casesCollaborate with development teams to improve release reliability and automation, using the approved LSEG DevOps toolsetLead capacity and performance management, improving monitoring capabilities to ensure platforms scale in line with business growth ...

Technical Account Manager, Google Cloud Consulting

Location
City of Westminster, England, United Kingdom
Account Manager, Google Cloud Consulting Share Technical Account Manager, Google Cloud Consulting corporate_fare Google place London, UK Bachelor’s degree in Computer Science, Engineering, a related technical field, or equivalent practical experience. 5 years of experience in a customer-facing role working with stakeholders, driving customer technical implementations … better understand business and technical needs. Plan for customer events and launches, partnering with Support, Engineering and Site Reliability Enginee (SRE) to ensure customer success during critical moments, and work with customers and Support to guide issues and escalations to resolution. Develop best practices and assets based ...

AI-Powered Observability Tech Lead

Location
Greater London, England, United Kingdom
Collaboration Technology Group in London seeks a Technical Leader to drive architectural vision and implement an AI-powered Production Intelligence platform. You will blend Site Reliability Engineering with agentic AI to improve monitoring, incident response, and auto-remediation across global SaaS infrastructure. You will mentor engineers, shape ...

Research Engineer, Safety Oversight, DeepMind

Location
Greater London, England, United Kingdom
technical products. Experience in the domain area of generative AI and Large Language Models (LLMs). Preferred qualifications: Master’s degree or PhD in Engineering, Computer Science, or a related technical field. 3 years of experience developing code, running experiments and analyses collaboratively with coding agents. Experience building large … Software Engineer, Full Stack, Google AdsGoogle-2w agoLondon, UKFull-time14DetailsG### Software Engineer III, Full Stack, Publisher InventoryGoogle-2w agoLondon, UKFull-time15DetailsG### Software Engineer III, Site Reliability Engineering, Traffic Network Load BalancingGoogle-2w agoLondon, UKFull-time15Details## Explore related hubsCountry hubUnited Kingdom JobsCompany pageGoogle JobsSalary pageSoftware Engineer SalaryVisa pageSkilled ...

Senior Product Manager for AI Observability

Location
Greater London, England, United Kingdom
with AI Evaluation PM (previous role), Model Risk, GSSR, Legal and Compliance to align telemetry with governance frameworks. Work closely with Engineering and SRE teams to drive observability improvements and reliability engineering for AI systems. Optimisation & Insights Identify cost inefficiencies across model and MCP usage, and drive … support auditability, compliance and explainability requirements. Skills & Competencies Required Experience in product management with a strong foundation in observability, telemetry, data platforms, monitoring, or SRE/DevOps‐driven products. Understanding of LLMs, embeddings, vector search, MCP tools, and AI inference workflows. Deep familiarity with logging, tracing, metrics, and event‐based ...

Solutions Architect, Studios — Pre-Sales & Architecture

Location
Greater London, England, United Kingdom
solutions across client contracts, new business opportunities, and procurement activity. The role emphasizes pre-sales engagement, proofs of concept, and collaboration with DevOps and Site Reliability Engineering teams to ensure automatable, observable, and secure solutions. The position reports into IMG Studios and involves working across commercial, technical ...

Senior Product Manager for AI Observability

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
with AI Evaluation PM (previous role), Model Risk, GSSR, Legal and Compliance to align telemetry with governance frameworks. Work closely with Engineering and SRE teams to drive observability improvements and reliability engineering for AI systems. Optimisation & Insights Identify cost inefficiencies across model and MCP usage, and drive … support auditability, compliance and explainability requirements. Skills & Competencies Required Experience in product management with a strong foundation in observability, telemetry, data platforms, monitoring, or SRE/DevOps-driven products. Understanding of LLMs, embeddings, vector search, MCP tools, and AI inference workflows. Deep familiarity with logging, tracing, metrics, and event-based ...