126 to 150 of 592 Site Reliability Engineering Jobs in London

Site Reliability Engineer, Infrastructure - ThousandEyes

Location
City Of London, England, United Kingdom
deeply integrated across the Cisco technology portfolio, delivering AI-powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios. Our distributed Site Reliability Engineering team of approximately nine engineers owns the availability, latency, performance, efficiency, monitoring, emergency response, and capacity planning of the platform while … operational on-call rotation. Hands-on experience with infrastructure-as-code tooling and codebases, preferably Terraform. Hands-on experienceleveraging AIas a force multiplier of SRE activities, such as automati ng toil away and improving operational efficiency. Professional experience administering and troubleshooting GNU/Linux systems, including system libraries, file systems ...

Cloud Operations Engineer (remote - London)

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
remote) Cloud Operations Engineer/Site Reliability Engineer – Fintech80,000 Plus Bonus + 10% non-cont pension + 10-15k bonus and sharesQuant Capital is urgently looking for a Site Reliability Engineer to join or well-known Fintech50 client who produces software disrupting the wealth … understanding of the OSI ModelExperience in database technology and basic query writing MSSQL, Postgres. This role suits a senior Engineer from a DevOps or SRE background who is a real technologist and cloud specialist interested in the latest tooling and technologies that support software development and infrastructure. The firm ...

DBA Lead

Location
City Of London, England, United Kingdom
Server, PostgreSQL, Oracle, and cloud database technologies across hybrid environments while driving automation, observability, disaster recovery readiness, and continuous improvement initiatives. Working closely with engineering, infrastructure, security, and DevOps teams, DBAs play a key role in delivering resilient and scalable data services that power business-critical applications. About … Ansible Observability & Reliability Engineering Experience with enterprise monitoring and observability platforms including Grafana/Datadog/ELK Azure Monitor CloudWatch Understanding of SRE concepts including: Service Level Indicators (SLIs) and Service Level Objectives (SLOs) Error Budgets Incident Management Root Cause Analysis Blameless Post-mortems Database Reliability Engineering ...

Core AI Engineer

Location
Greater London, England, United Kingdom
Artificial Intelligence, Automation and Intelligent Engineering. We are building enterprise-scale AI capabilities that improve service resilience, automate operational workflows, accelerate engineering productivity and enhance customer outcomes. As a Core AI Engineer, you will play a leading technical role in the design, development and deployment of AI solutions across … direction across teams without formal management responsibility. Desirable Experience Experience within Financial Services or highly regulated environments. Knowledge of Service Reliability Engineering (SRE) principles. Experience developing AI-powered operational tooling. Experience building internal AI platforms or developer enablement capabilities. Familiarity with Microsoft AI ecosystem, Copilot technologies and Azure ...

Senior DevOps / Platform Engineer (Google Cloud)

Hiring Organisation
Datatonic
Location
London, UK
Employment Type
Full-time
Cloud's premier partner in AI, driving transformation for world-class businesses. We push the boundaries of technology with expertise in machine learning, data engineering, and analytics on Google Cloud Platform. By partnering with us, clients future-proof their operations, unlock actionable insights, and stay ahead of the curve … scale-up environmentContainerisation/Virtualisation Expertise: Proficiency with technologies such as Terraform and KubernetesSRE Principles: Experience in implementing Site Reliability Engineering (SRE) principlesCloud Native Architecture: Hands-on experience with cloud-native architectures, ideally on Google CloudClient-Facing Role: Prior experience in a client-facing positionSDN Knowledge: Understanding ...

AWS Engineer

Location
Greater London, England, United Kingdom
systems.Define data storage and persistence approaches that balance scalability, performance, and maintainability.Collaborate with business and technical stakeholders to translate requirements into effective solution designs.Promote engineering excellence through architecture governance, design best practices, and continuous improvement.Support delivery teams in resolving complex technical challenges and ensuring successful solution outcomes.Champion innovation, continuous …/CD pipelines (GitLab preferred).Demonstrated experience in automating build, test, deployment, and release management processes.Solid understanding of Site Reliability Engineering (SRE) principles, including observability, monitoring, resilience, availability, incident management, and operational support.Strong troubleshooting and production support skills across cloud infrastructure and application environments.Ability to independently deliver ...

Site Reliability Engineer (London) - Banking & Finance

Location
Greater London, England, United Kingdom
collaboration and technical excellence, the organisation continues to push the boundaries of low-latency infrastructure and reliable system design. The team is hiring a Site Reliability Engineer (London) to build, monitor, and optimise mission-critical trading systems. The role will focus on automation, system scalability, and incident response … improving the infrastructure. Drive automation and operational excellence by leveraging your Linux expertise, Kubernetes, and Python scripting skills. Monitor and ensure high availability and reliability of trading applications while being on top of system alerts and incidents. Key Requirements: 1-5 years working experience The right candidate will come ...

Site Reliability Engineer, Infrastructure - ThousandEyes

Location
City Of London, England, United Kingdom
deeply integrated across the Cisco technology portfolio, delivering AI-powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios. Our distributed Site Reliability Engineering team of approximately nine engineers owns the availability, latency, performance, efficiency, monitoring, emergency response, and capacity planning of the platform while … complex issues across infrastructure and platform services, participate in the on-call rotation and incident-management process, and turn root-cause findings into lasting reliability improvements. Collaborate with application development teams and other stakeholders to meet internal Service Level Objectives and customer-facing Service Level Agreements while improving ...

Senior Network Site Reliability Engineer

Location
Greater London, England, United Kingdom
About the Team Miro is a fast-growing engineering organization building a business-critical collaboration platform used by companies around the world. As our product and infrastructure scale, we are looking for a Senior Network Site Reliability Engineer to help strengthen the reliability, availability, and scalability … What you’ll need 8+ years of professional experience in infrastructure, reliability, networking, or software engineering 6+ years of experience as an SRE, DevOps Engineer, Network Engineer, Software Engineer, or similar Hands-on experience with AWS infrastructure, including EC2, VPC, ALB, S3, Route 53, and CloudFront Confident networking ...

Scala Engineer

Location
City Of London, England, United Kingdom
Experience working within Continuous Integration environments Strong understanding of Agile methodologies Experience with testing and automation Awareness of Site Reliability Engineering (SRE) principles and support Experience troubleshooting incidents and restoring services following outages Experience working in a you build it, you run it environment Strong collaborative ...

Head of Technology Resilience and Product Operations

Hiring Organisation
Deerfoot Recruitment Solutions
Location
London, UK
Employment Type
Full-time
Management, Release Management, Business Impact Analysis, Senior Leadership, ITIL 4, ISO22301, ISO/IEC 20000, CBCP, MBCI, Site Reliability Engineering (SRE), Vulnerability Management, PowerShell, Python, Splunk, CyberArk, GenAIDirector, Head of Technology Resilience and Production Operations London (Hybrid) | Banking Sector up to 140,000 + Bonus + Benefits … ability to lead confidently and calmly under pressureDesirable: ITIL 4, ITSM tooling (ServiceNow, Jira Service Management), ISO22301/ISO20000, CBCP/MBCI certification, SRE familiarity, or experience with tools such as Splunk, CyberArk PAM, or GenAI-driven service managementReady to take on a role where your leadership genuinely shapes resilience ...

Data Platform Engineer

Hiring Organisation
MONY Group
Location
London, UK
Employment Type
Full-time
personalised customer experiences. We work closely with teams across the business to make data clean, reliable, secure and accessible for decision-making. Data & AI Engineering is a cross-functional team of engineers and scientists. We integrate with the group's operational data stores, maintain shared data models, build … ability to apply automation responsibly to real delivery and operational problems. You might come from data engineering, platform engineering, software engineering, SRE, analytics engineering, MLOps, or cloud infrastructure. What matters most is that you enjoy reducing toil, improving developer experience, and building secure, observable systems that ...

Associate Site Reliability Engineer, SRE Platforms

Location
Greater London, England, United Kingdom
Goldman Sachs is seeking an Associate Site Reliability Engineer to help build, run and continuously improve the Consolidated Trade Ledger platform in London. This role combines software and systems engineering to boost reliability, observability and incident response for a cloud-native service. The candidate will implement ...

Scala Engineer

Location
City Of London, England, United Kingdom
services. Support incremental re-architecting initiatives to reduce technical complexity and improve maintainability. Develop clean, testable and maintainable code using Scala and modern engineering practices. Design, build and maintain secure APIs, databases and applications. Collaborate with Product Owners, Business Analysts, Data Engineers and wider technical teams to deliver effective … design and development experience. Experience working with databases and SQL. Hands-on AWS cloud experience. Understanding of Site Reliability Engineering (SRE) principles. Experience supporting and restoring production services during incidents. Strong appreciation of testing, automation and software quality practices. Experience working within Agile environments. Experience with Continuous ...

Linux Desktop Support Engineer

Location
Greater London, England, United Kingdom
client, a dynamic financial technology organization who is seeking a Linux Desktop Support Engineer to join its Site Reliability Engineering (SRE) team in London. This position focuses on managing internal devices, supporting workplace technology, and maintaining office infrastructure, with a strong emphasis on Ubuntu/Linux environments. … streamline support processes. This is an excellent opportunity for a hands-on, technical individual looking to own workplace technology within a high-performance engineering environment. Key Responsibilities Own the provisioning, configuration, and maintenance of Linux, macOS, and Windows devices. Manage employee onboarding technical setups and serve as the primary ...

Senior Business Analyst

Location
Greater London, England, United Kingdom
member of a multi-functional team involved in all phases of our product release lifecycle that adopts and promotes the DevOps & SRE (Site Reliability Engineering) methodologies, responsible for requirements, analysis, and delivery of our products.In this role you will have the opportunity to enhance your analytical ...

Network Site Reliability Engineer

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Network SRE – 250,000-350,000 total compensation – 4 days in officeQuant Capital is urgently looking Network SRE for our high profile client. Our client is a leading quantitative trading company and liquidity provider. Their focus on technology has allowed them to deeply penetrate the market and gain market share. … Shared Engineering team that focuses on designing, developing, and maintaining infrastructure and tools. The team requires a Network Site Reliability Engineer (SRE) with strong network fundamentals, problem-solving skills, and a keen interest in diverse tools and techniques. The role involves collaborative work across various teams, exploring ...

AI Native DevOps Platform Engineer

Hiring Organisation
Sanderson Recruitment
Location
London, United Kingdom
Employment Type
Permanent
Develop governance controls, guardrails and approval workflows for AI-driven infrastructure operations Implement monitoring, logging, tracing and alerting across cloud platforms and applications Establish SRE principles and improve platform reliability, resilience and operational performance Embed security, governance and compliance throughout the software delivery lifecycle Optimise Azure environments for performance … systems and microservices Strong understanding of cloud networking, security and production platform operations Experience implementing observability across monitoring, logging, tracing and alerting Experience applying SRE principles to improve platform reliability and operational performance Experience using AI tools, intelligent automation or AI agents to improve engineering productivity, infrastructure delivery ...

Site Reliability Engineer - Banking & Finance

Location
Greater London, England, United Kingdom
Ready to take the next step in your career? Join a leading technology-driven trading firm where engineering, automation, and high-performance infrastructure are central to supporting global trading operations. The organisation invests heavily in modern platform engineering practices, enabling teams to build reliable, scalable, and highly automated … Have: Strong experience programming with Python, Go and/or C++ Strong Linux knowledge and understanding of distributed systems. Experience with monitoring, observability or SRE practices. Experience with CI/CD pipelines, Git and infrastructure automation. Familiarity with Kubernetes and containerised workloads. Strong analytical and troubleshooting skills. Benefits: Build ...

Pre-Sales Solutions Architect (Cloud / AI Managed Services)

Location
Greater London, England, United Kingdom
practices like XP and CI/CD to achieve \"zero maintenance\" products, revolutionizing how Run operates. By combining site reliability engineering (SRE), product evolution and data ops, DAMO managed services drive predictable cost reduction and future-proof operations. Principal solutions architects are a driving force … that extends to empowering our employees in their career journeys. About Thoughtworks Thoughtworks is a global technology consultancy that integrates strategy, design and engineering to drive digital innovation. For 30 years, our clients have trusted our autonomous teams to build solutions that look past the obvious. Here, computer science ...

Nework Site Reliability Engineer - Algo trading

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Network SRE – 250,000-300,000 total compensation – 4 days in officeQuant Capital is urgently looking Network SRE for our high profile client. Our client is a leading quantitative trading company and liquidity provider. Their focus on technology has allowed them to deeply penetrate the market and gain market share. … Shared Engineering team that focuses on designing, developing, and maintaining infrastructure and tools. The team requires a Network Site Reliability Engineer (SRE) with strong network fundamentals, problem-solving skills, and a keen interest in diverse tools and techniques. The role involves collaborative work across various teams, exploring ...

Head of Cloud Platform Engineering

Location
Greater London, England, United Kingdom
exciting point in our journey. As we continue to evolve our platform and expand our SaaS offering, we're investing in the engineering foundations that will support the next phase of our growth. About the role This is a rare opportunity to shape the future of Totara's cloud … world. We're clear on where we're heading, but how we get there is still being built. As Head of Cloud Platform Engineering, you'll play a central role in defining that journey. Today our infrastructure capability spans multiple teams, regions, and platforms, including environments inherited through acquisition. ...

Sr. Network Site Reliability Engineer (SREs)

Location
Greater London, England, United Kingdom
/ML Technologies and Professional services in the UK and EU market. Job Description Overview We are seeking a highly experienced Senior Network SRE with deep expertise across multi-vendor network infrastructure, automation, and reliability engineering. The ideal candidate will possess strong technical leadership, hands‐on engineering capabilities … resilient, scalable, and observable network environments. Key Responsibilities Design, implement, and maintain highly available network solutions across routing, switching, firewalling, and wireless technologies. Apply SRE principles to improve network reliability, scalability, and performance. Develop and maintain automation workflows using Ansible, Salt, and related frameworks to reduce operational toil. Build ...

Senior Site Reliability Engineer

Hiring Organisation
Imanage
Location
London, UK
Employment Type
Full-time
SRE is part of a global organization that leverages the latest technology to communicate with our colleagues across the globe. We organize ourselves into distributed teams -- SRE teams are anchored to iManage offices across the globe. Tuesdays and Thursdays are dedicated to in-office collaboration, rapid innovation, and developing … engage in and often lead architectural discussions, reduce toil, and deliver scalable, resilient platforms that support our customers and organization. As a Senior SRE, you'll help scale our cloud platform, collaborate across teams to promote standardization and resiliency, and participate in on-call rotations. ...

Lead Site Reliability Engineer (Kubernetes Required) - Hybrid

Hiring Organisation
FactSet Research Systems
Location
London, UK
Employment Type
Full-time
curiosity is the key to anticipating our clients' needs and exceeding their expectations. About the RoleWe are looking for a skilled and motivated Lead Site Reliability Engineer to join our team. In this role, you will be responsible for ensuring the reliability, scalability, and performance … under pressure, particularly during incident response Commitment to a blameless culture and continuous learning Nice to HaveExperience contributing to open-source projects Familiarity with SRE principles as defined by the Google SRE handbook Previous experience in a DevOps or Platform Engineering role Company Overview: FactSet (NYSE:FDS | NASDAQ ...