101 to 125 of 572 Site Reliability Engineering Jobs in London

Site Reliabiity Engineer (DV cleared)

Hiring Organisation
Stott & May Professional Search Limited
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£700 - £755 per day
Site Reliability Engineer (SRE) - Contract *candidates must have an active DV clearance* Start: ASAP Duration: 6 months+ Location: 4-5 days on site in london Rate: up to £755 .day inside IR35 We're looking for an experienced Site Reliability Engineer (SRE) to join … London. This is a hands-on role covering application support, systems administration, production databases and backend infrastructure, working closely with an existing on-site engineer. Key Responsibilities - Manage and support multiple production databases, including MSSQL. - Provide 1st-3rd line support and troubleshoot application/system issues. - Administer Linux ...

Site Reliability Engineer with Python

Hiring Organisation
Nexus Jobs
Location
London, UK
Employment Type
Full-time
000Sector: I.T. & CommunicationsJob Type: PermanentWork Hours: Full TimeContact: Jas GujralEmail: cv@nexusjobs.comTelephone: 020 7488 6900Apply for this job nowJob DescriptionSite Reliability Engineer with PythonOur Client looking to bring on a site reliability engineer to help deploy, manage, troubleshoot, and enhance our complex cloud-based set of internal … variety of users across our wide-ranging organization. You will have at least 7 to 10 years hands-on expertise working as a Site Reliability Engineer. You will work closely with IT, product, and engineering to extend and maintain this set of tools and services ...

Principal Site Reliability Engineer, Infrastructure Observability

Location
Greater London, England, United Kingdom
toolchain and systems, code build and deployment, incident response, and 24x7 monitoring and support. The candidate will also have extensive experience operating within a SRE function within a complex, distributed environment. They will have a demonstrated ability to work horizontally and vertically within an organization with diverse partners and sponsor … learning through blameless post-mortems to improve the shared goal of reliability across services Transform operations teams by facilitating internal change to adopt SRE standard methodologies across the organization and driving strategic growth in this area within Global Technology Analyzes incidents impacting technology availability for high-level trends across ...

Systems Engineering Manager

Hiring Organisation
Hackajob Ltd
Location
City Of Westminster, London, United Kingdom
Employment Type
Permanent
Salary
GBP Annual
client is a global technology company. Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that our client's services-both our internally critical and our externally-visible systems-have reliability, uptime ...

Engineering Manager, Edge SRE

Location
Greater London, England, United Kingdom
coverage during daytime hours. SREs are supported by all engineering teams at Cloudflare who participate in on call schedules for their services. The SRE teams facilitate remediation and follow up of production issues and mature the tooling to enable all engineering teams to self-service on production. Incident … engineering teams is prioritized above product innovation and the impact of production incidents influences the priority. SREs support two main environments: Edge SRE are focused on edge distribution where most client traffic is served. Core SRE are focused on the core services like control plane, data pipeline and other ...

Staff Platform Site Reliability Engineer

Hiring Organisation
Index Exchange
Location
London, UK
Employment Type
Full-time
users—it just works. Mentor and raise the bar. Coach engineers, foster a culture of engineering excellence, and collaborate across Cloud Platform Operations, SRE, Network, Security, and Software Engineering teams. What You BringYou've built and shipped platform infrastructure at scale—not just operated someone else's. … care as much about the developer experience of your platform as you do about its architecture. Must Have8+ years in platform engineering, SRE, infrastructure engineering, or DevOps. Deep experience with Linux internals: kernel tuning, network stack, system observability, security. Strong Kubernetes expertise: cluster lifecycle, networking, storage, RBAC, multi ...

Site Reliability Engineer (SRE), London

Location
Greater London, England, United Kingdom
Selection changes the language of the page/content Site Reliability Engineer (SRE), London London, England, United Kingdom Software and Services People at Apple don’t just build products — they craft experiences our customers love and depend on. Apple Services Engineering (ASE) builds and supports the systems … groundbreaking approach to cloud intelligence, extending the security and privacy of Apple devices into the cloud to unlock even more intelligence for our users.This SRE team is responsible for the availability and automation of the critical systems and services that enable PCC to deliver cloud intelligence without compromising user privacy. ...

Senior DevOps / Platform Engineer (Google Cloud)

Location
Greater London, England, United Kingdom
Cloud's premier partner in AI, driving transformation for world-class businesses. We push the boundaries of technology with expertise in machine learning, data engineering, and analytics on Google Cloud Platform. By partnering with us, clients future-proof their operations, unlock actionable insights, and stay ahead of the curve … Experience: Previous experience working in a start-up or scale-up environment Containerisation/Virtualisation Expertise: Proficiency with technologies such as Terraform and Kubernetes SRE Principles: Experience in implementing Site Reliability Engineering (SRE) principles Cloud Native Architecture: Hands-on experience with cloud-native architectures, ideally on Google ...

DevOps Engineer (Security Cleared)

Location
Greater London, England, United Kingdom
Solirius Reply, part of the Reply Group, is a technology consultancy and digital transformation partner that helps organisations solve complex challenges through strategy, design, engineering, and delivery. We work closely with our clients to deliver secure, accessible, user-focused services that evolve with their needs. By combining deep technical … Ministry of Housing, Communities and Local Government, UEFA, International Olympic Committee, and Mercedes-Benz. Our services span the full digital delivery lifecycle, including architecture, engineering, delivery management, user-centred design, business analysis, data, DevOps, and AI. We operate as a collaborative and inclusive organisation that empowers our people ...

Cloud Network Engineer

Hiring Organisation
AMS CWS
Location
London, United Kingdom
Employment Type
Contract
services. Develop scripts and tools (e.g., Python, Go, Bash) to streamline network operations, ensure consistency, and improve efficiency Site Reliability Engineering (SRE) for Networks: Embrace a 'you build it, you run it' mindset for network services. Take ownership of the reliability, performance, and availability … Jenkins . Automated testing experience using Terratest, Cucumber, Pytest-BDD, AWS Fault Injection Simulator or Chaos Mesh . Experience applying DevOps, agile and SRE practices to cloud networks, including monitoring, logging, alerting, incident response and performance optimisation. Strong communication skills with strategic thinking and adaptability Next steps Next steps This ...

Staff SRE, AI Infrastructure

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
each other to deliver impact. Make Wayve the experience that defines your career! The roleThis is a rare opportunity to be a founding Staff SRE shaping the reliability of large-scale AI systems and GPU compute infrastructure from the ground up. As a Staff Cloud Site Reliability … Compute platform (large-scale, multi-tenant GPU fleets and scheduling systems driving model training and inference at scale).This is a founding Cloud SRE role. You won't inherit a mature SRE function, you'll help create it. You will define the frameworks, automation, and operational standards that ensure ...

Software Engineer

Location
Greater London, England, United Kingdom
years, or PhD + 3 years in Computer Science, Software Engineering, or a related technical field. Experience as a Senior/Lead SRE or Software Engineer delivering distributed, high-availability SaaS platforms at scale. Strong proficiency in Python, Go, Java, or C++ with experience designing microservices, APIs, and production … automation. Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments. Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated root cause analysis (RCA). Preferred Qualifications AI & Agentic Systems: E xperience building LLM pipelines ...

Platform Lead: Fintech Architecture & SRE

Location
Greater London, England, United Kingdom
supports rapid growth, and upholds the trust expected in financial services. They will design a resilient architecture that accelerates product development, and deliver exceptional reliability as we grow.Partnering closely with product and engineering teams, they will combine hands-on building with strategic technical leadership to ensure our platform … making process for platform technologies, balancing in-house development with best-in-class third-party solutions* Drive a Site Reliability Engineering (SRE) culture, ensuring high availability, low latency, and robust disaster recovery capabilities* Manage and optimize our cloud infrastructure, focusing on Infrastructure-as-Code (e.g., Terraform), containerization ...

Senior Customer Success Associate — Global Banking Experience & Intelligence

Location
Greater London, England, United Kingdom
help teams get more value from the tools they use every day. You combine hands-on troubleshooting with direct engagement, partnering with product and engineering to resolve incidents, reduce repeat issues, and improve the overall day-to-day experience. You turn what you learn from support into scalable improvements … analysis, and observability tools to isolate likely root causes Reproduce, validate, upscale.. Create reproducible tests, validate dependencies, and elevate effectively using runbooks Partner across engineering and product. Partner with engineering, site reliability engineering, and product teams to drive fixes through release and stabilization Monitor, triage ...

Senior Linux DevOps Engineer

Hiring Organisation
RedTech Recruitment Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
annum + excellent benefits Requirements for Senior Linux DevOps Engineer: Strong commercial experience working as a Senior DevOps Engineer, Linux Engineer, Platform Engineer, Site Reliability Engineer or similar Excellent Linux systems administration and command line skills, with experience operating and troubleshooting large-scale production environments Strong scripting … Linux Engineer/Linux Systems Engineer/Linux Infrastructure Engineer/Senior Platform Engineer/Platform Engineer/Site Reliability Engineer/SRE/Infrastructure Engineer/DevSecOps Engineer/Linux/Bash/Shell Scripting/Python/Kubernetes/Docker/Terraform/Ansible/Microsoft ...

Senior DevOps Engineer

Location
Greater London, England, United Kingdom
technical problems and developing software updates and 'fixes' Guide the team in the provision of best in class Site Reliability Engineering (SRE) capability and practices as it relates to the platform Help define and implement the onboarding process for new customers to the platform from the platform … intended Understand the needs of Engineers and Test Teams in relation to the platform roadmap priorities to ensure an infrastructure strategy that continuously compounds engineering productivity and effectiveness. Sizing, forecasting and Finops Define strong financial control capabilities for platform instances and contribute to price performance optimisation of infrastructure ...

Core AI Engineer

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
Artificial Intelligence, Automation and Intelligent Engineering. We are building enterprise-scale AI capabilities that improve service resilience, automate operational workflows, accelerate engineering productivity and enhance customer outcomes. As a Core AI Engineer, you will play a leading technical role in the design, development and deployment of AI solutions across … systems design. AI observability, evaluation and governance frameworks. Desirable ExperienceExperience within Financial Services or highly regulated environments. Knowledge of Service Reliability Engineering (SRE) principles. Experience developing AI-powered operational tooling. Experience building internal AI platforms or developer enablement capabilities. Familiarity with Microsoft AI ecosystem, Copilot technologies and Azure ...

Monitoring & Observability Engineer (Dynatrace)

Location
Greater London, England, United Kingdom
some of the world’s most well-known organisations. You’ll play a key role in helping our customers achieve greater visibility, performance, and reliability across their IT estates—contributing to their operational success through proactive insight and incident prevention. What you'll do Design, implement, and manage observability … passion for continuous improvement and knowledge sharing Certifications Dynatrace Associate & Pro Splunk Core Certified Power User DevOps or Site Reliability Engineering (SRE) experience Automation with Terraform or similar tools Experience with Docker and Kubernetes for packaging and deployment Ability to adapt to new technologies in fast-paced ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
Overview CTI is seeking to appoint a head of engineering for a fintech scaleup. Requirements Proficient in AWS services such as ECS, EC2, Lambda, VPC, IAM, Route53, CloudFront, S3, and RDS. Solid … understanding of monitoring and logging tools, including Prometheus, AWS CloudWatch, Grafana, OpenTelemetry, Honeycomb, and ELK. Basic knowledge of Site Reliability Engineering (SRE) and experience with alerting and incident management systems like Opsgenie and PagerDuty. Demonstrated capability to develop and maintain robust and scalable Continuous Integration/Continuous ...

Production Engineer

Hiring Organisation
Liquidnet
Location
London, UK
Employment Type
Full-time
issues across our trading platform. You will leverage deep expertise in FIX, Linux, Windows Server, DevOps, databases, networking, and cloud technologies to ensure platform reliability and performance. This is a hands-on leadership role involving complex troubleshooting across cross-platform market-leading technologies, driving automation and tooling improvements … working hoursPositive approach to the day-to-day, with the resilience to handle high-pressure production incidentsDesiredExperience with Site Reliability Engineering (SRE) practices, including monitoring, incident response, and post-mortem analysisProven experience applying AI or machine-learning models to optimise workflows, identify patterns, and drive intelligent automation ...

Production Engineer

Location
Greater London, England, United Kingdom
issues across our trading platform. You will leverage deep expertise in FIX, Linux, Windows Server, DevOps, databases, networking, and cloud technologies to ensure platform reliability and performance.This is a hands-on leadership role involving complex troubleshooting across cross-platform market-leading technologies, driving automation and tooling improvements, and acting … Positive approach to the day-to-day, with the resilience to handle high-pressure production incidentsDesired* Experience with Site Reliability Engineering (SRE) practices, including monitoring, incident response, and post-mortem analysis* Proven experience applying AI or machine-learning models to optimise workflows, identify patterns, and drive intelligent ...

Production Engineer

Location
City Of London, England, United Kingdom
issues across our trading platform. You will leverage deep expertise in FIX, Linux, Windows Server, DevOps, databases, networking, and cloud technologies to ensure platform reliability and performance. This is a hands-on leadership role involving complex troubleshooting across cross-platform market-leading technologies, driving automation and tooling improvements … approach to the day-to-day, with the resilience to handle high-pressures production incidents Desired Experience with Site Reliability Engineering (SRE) practices, including monitoring, incident response, and post-mortem analysis Proven experience applying AI or machine-learning models to optimise workflows, identify patterns, and drive intelligent ...

Site Reliability Engineer, Infrastructure - ThousandEyes

Location
City Of London, England, United Kingdom
deeply integrated across the Cisco technology portfolio, delivering AI-powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios. Our distributed Site Reliability Engineering team of approximately nine engineers owns the availability, latency, performance, efficiency, monitoring, emergency response, and capacity planning of the platform while … operational on-call rotation. Hands-on experience with infrastructure-as-code tooling and codebases, preferably Terraform. Hands-on experienceleveraging AIas a force multiplier of SRE activities, such as automati ng toil away and improving operational efficiency. Professional experience administering and troubleshooting GNU/Linux systems, including system libraries, file systems ...

Cloud Operations Engineer (remote - London)

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
remote) Cloud Operations Engineer/Site Reliability Engineer – Fintech80,000 Plus Bonus + 10% non-cont pension + 10-15k bonus and sharesQuant Capital is urgently looking for a Site Reliability Engineer to join or well-known Fintech50 client who produces software disrupting the wealth … understanding of the OSI ModelExperience in database technology and basic query writing MSSQL, Postgres. This role suits a senior Engineer from a DevOps or SRE background who is a real technologist and cloud specialist interested in the latest tooling and technologies that support software development and infrastructure. The firm ...

DBA Lead

Location
City Of London, England, United Kingdom
Server, PostgreSQL, Oracle, and cloud database technologies across hybrid environments while driving automation, observability, disaster recovery readiness, and continuous improvement initiatives. Working closely with engineering, infrastructure, security, and DevOps teams, DBAs play a key role in delivering resilient and scalable data services that power business-critical applications. About … Ansible Observability & Reliability Engineering Experience with enterprise monitoring and observability platforms including Grafana/Datadog/ELK Azure Monitor CloudWatch Understanding of SRE concepts including: Service Level Indicators (SLIs) and Service Level Objectives (SLOs) Error Budgets Incident Management Root Cause Analysis Blameless Post-mortems Database Reliability Engineering ...