1 to 25 of 105 Proactive Monitoring Jobs in London

Observability Manager

Location
Greater London, England, United Kingdom
roadmap across infrastructure, applications, and services. Deliver scalable observability solutions using metrics, logs, traces, and events to improve service visibility and performance. Enable proactive monitoring, predictive alerting, and faster incident detection, diagnosis, and resolution. Support Major Incident and Problem Management through real-time insights and evidence-based root … Operations, and third-party partners to maximise operational value. Ensure observability practices align with security, compliance, and regulatory requirements. Qualifications & Experience Deep understanding of monitoring frameworks, telemetry, and observability concepts. Hands-on experience with enterprise monitoring tools (e.g., LogicMonitor, ManageEngine, ServiceNow). Proven ability to deliver operational excellence ...

Senior Technical Account Manager, Strategic Industries - Global Financial Services

Location
Greater London, England, United Kingdom
technical guidance to help plan and build solutions using best practices, and proactively keep your customers' AWS environments operationally healthy through application-specific reviews, proactive monitoring, and custom runbooks. You will establish and evolve frameworks governing customer operations on AWS, including DevSecOps practices, Operational Excellence standards, AI strategy … Target Operating Models. You will use generative AI tools fluently to accelerate your own work and, more importantly, to unlock deeper insights and proactive recommendations for your customers. You will guide customers on adopting AWS generative AI services to modernise their operations and drive measurable business outcomes. You will ...

Subject Matter Expert (Support&Ops)

Location
Greater London, England, United Kingdom
Monthly/weekly OS and platform patching of Azure VMs and PaaS servicesPatch validation, compliance tracking, and post-patch verification2. Backup & Restore Operations Daily monitoring of Azure Backup jobs (VMs, PaaS-supported workloads)Backup validation, restore testing, and failure remediation3. Monitoring & Alert Management Proactive monitoring using … incidents and escalationsDeep troubleshooting of Azure IaaS, PaaS and DevopsPatch strategy planning and compliance reportingBackup architecture review, restore validation, and DR drillsAdvanced automation and monitoring improvementsProblem management and preventive action implementationLead MSR, QBR, and capacity planning discussionsMentor L2 teams and improve SOPs/KB articles Skills & Experience :7–10+ ...

Director, Credit Risk

Hiring Organisation
Airwallex
Location
London, UK
Employment Type
Full-time
management: Set and oversee counterparty risk standards for banking partners, payment partners, merchants, customers, and other material counterparties. Establish appropriate exposure limits, concentration thresholds, monitoring requirements, and escalation processes. Portfolio monitoring and early warning indicators: Lead proactive monitoring of credit, FX, and counterparty portfolios, including exposure … other relevant risk indicators. Identify emerging trends and escalate material issues before they become systemic. Loss mitigation and remediation: Review and challenge first-line monitoring, collections, collateral, prefunding, reserves, chargeback, and other mitigation strategies. Ensure deteriorating counterparties, customers, merchants, or exposures are identified and addressed promptly. New products ...

SysOps Team Lead

Location
Greater London, England, United Kingdom
platforms. Optimise toolset processes (e.g. n‐Able, Acronis, Mimecast, Intune, Microsoft MDE, BitDefender, Qualys, Meraki, etc.) for efficiency and scalability. Good scripting ability Drive proactive monitoring and self‐healing systems. Systems Administration: Manage servers, Active Directory (Entra), M365 and cloud platforms (Azure/AWS). Ensure system reliability ...

Azure Data Support Engineer

Hiring Organisation
Avanade
Location
London, UK
Employment Type
Full-time
Services environment. This role is responsible for ensuring the stability, performance, and reliability of enterprise-scale data and machine learning workloads on Azure through proactive monitoring, issue investigation, automation, and continuous improvement. The ideal candidate will also play a key role in supporting data warehouse systems by managing ...

Senior Reliability Engineer

Hiring Organisation
Fitch Ratings
Location
London, UK
Employment Type
Full-time
application deployments for reliability, security, and efficiencyIdentify, contain, and mitigate risk across all cloud environments, maintaining a robust security posture for infrastructure and applicationsImplement proactive monitoring and observability practices to detect and prevent issues before they impact usersDevelop and maintain automation and tooling solutions, including AI-assisted approaches ...

Senior Performance Support Specialist

Hiring Organisation
Unit4
Location
Greater London, United Kingdom
Employment Type
Full Time
Troubleshoot performance issues across new implementations (Project) and existing implementations (BAU) of Unit4 Financials by Coda, this includes Technical and Application aspects. Develop enhanced, proactive monitoring, detection and remediation of performance issues. Develop proactive, continuous performance tuning and processes to maintain optimal performance levels. Comply with … Team Context One of three in the U4F by Coda team, alongside a Senior Performance & Optimisation Consultant and a Database Administrator. Together you deliver proactive performance and technical support before and after go live. Location - this role is fully remote. Occasional customer visits may be required, so candidates should ...

Dynamics 365 F&O Platform Engineer

Location
Greater London, England, United Kingdom
technical bridge between the D365 functional/business teams and the underlying Azure infrastructure. You will be responsible for environment management, deployment governance, platform monitoring, release management, and integration health across F&O and connected systems, ensuring the platform is stable, secure, resilient, and ready to support the business. … versioning, and release processes Monitor and troubleshoot F&O performance issues, including batch jobs, SQL/Azure SQL performance, and overall environment health Develop proactive monitoring, alerting, and operational documentation, driving continuous improvement through automation Manage data integrations between F&O and downstream systems (e.g. Shopify Plus, Patchworks ...

Systems Engineer III

Hiring Organisation
Elsevier
Location
London, UK
Employment Type
Full-time
ensuring changes are properly documented, tested, and executedCollaborate closely with development, support, and platform teamsAssist in daily support of assigned systems and products through proactive monitoring and early detection of issuesConfigure, troubleshoot, and maintain cloud infrastructure components across multiple environmentsRequirementsCloud infrastructure experience with AWS, Azure, or GCPInfrastructure … Python, or similar technologiesFamiliarity with containerisation and orchestration technologies such as Docker and KubernetesUnderstanding of CI/CD pipelines and DevOps practicesFamiliarity with system monitoring tools such as New Relic or similarKnowledge of change and incident management processesGood oral and written communication skillsWork in a Way That Works ...

Systems Engineer III

Location
Greater London, England, United Kingdom
properly documented, tested, and executed Collaborate closely with development, support, and platform teams Assist in daily support of assigned systems and products through proactive monitoring and early detection of issues Configure, troubleshoot, and maintain cloud infrastructure components across multiple environments Cloud infrastructure experience with AWS, Azure … technologies Familiarity with containerisation and orchestration technologies such as Docker and Kubernetes Understanding of CI/CD pipelines and DevOps practices Familiarity with system monitoring tools such as New Relic or similar Knowledge of change and incident management processes Good oral and written communication skills Work ...

Systems Engineer III

Hiring Organisation
Elsevier
Location
Greater London, United Kingdom
Employment Type
Full Time
properly documented, tested, and executed Collaborate closely with development, support, and platform teams Assist in daily support of assigned systems and products through proactive monitoring and early detection of issues Configure, troubleshoot, and maintain cloud infrastructure components across multiple environments Requirements Cloud infrastructure experience with AWS, Azure … technologies Familiarity with containerisation and orchestration technologies such as Docker and Kubernetes Understanding of CI/CD pipelines and DevOps practices Familiarity with system monitoring tools such as New Relic or similar Knowledge of change and incident management processes Good oral and written communication skills Work ...

Dynamics 365 F&O Platform Engineer

Hiring Organisation
End Clothing
Location
London, UK
Employment Type
Full-time
technical bridge between the D365 functional/business teams and the underlying Azure infrastructure. You will be responsible for environment management, deployment governance, platform monitoring, release management, and integration health across F&O and connected systems, ensuring the platform is stable, secure, resilient, and ready to support the business. … control, branching, versioning, and release processesMonitor and troubleshoot F&O performance issues, including batch jobs, SQL/Azure SQL performance, and overall environment healthDevelop proactive monitoring, alerting, and operational documentation, driving continuous improvement through automationManage data integrations between F&O and downstream systems (e.g. Shopify Plus, Patchworks ...

Platform Engineering Manager (Cloud Foundations)

Location
Greater London, England, United Kingdom
psychologically safeenvironment;with leadership tailored to individual strengths and motivations. Platform Engineering, Operations&Reliability Cloudplatform and keyservices arereliable,incident detection and resolutionaresmooth due to proactive monitoring, well‐maintainedalerts/logs, and complete observability coverage. Platform resilience and DR planning/testing strategyisdefined and operational,working across the business … network and key components aremaintainedfor audit, security, knowledgesharing purposes. Whatyou’llbring Proven experience in cloud operations and platform engineeringmanagement, overseeing cloud infrastructureand platforms, monitoring,reliability, and service delivery in production environments. Previoushands‐on experience running large cloud‐based website environmentswith GKE, service mesh, load balancers, CDN/WAF, Kafka ...

Platform Engineering Manager (Cloud Foundations)

Hiring Organisation
Rightmove
Location
London, UK
Employment Type
Full-time
individual strengths and motivations. Platform Engineering, Operations & Reliability Cloud platform and key services are reliable, incident detection and resolution are smooth due to proactive monitoring, well‐maintained alerts/logs, and complete observability coverage. Platform resilience and DR planning/testing strategy is defined and operational, working across … audit, security, knowledge sharing purposes. What you'll bring Proven experience in cloud operations and platform engineering management, overseeing cloud infrastructure and platforms, monitoring, reliability, and service delivery in production environments. Previous hands‐on experience running large cloud-based website environments with GKE, service mesh, load balancers, CDN/ ...

Platform Engineering Manager (Cloud Foundations) London, UK

Location
Greater London, England, United Kingdom
resilience mechanisms are regularly tested, ensuring the platform can recover predictably. The platform is reliable, incident detection and resolution are smooth due to proactive monitoring, well‐maintained alerts/logs, and complete observability coverage. Platform resilience and DR planning/testing strategy is defined and operational, working across … leadership to individual strengths and motivations. What you’ll bring Proven experience in cloud operations and platform engineering management, overseeing cloud infrastructure and platforms, monitoring, reliability, and service delivery in production environments. Hands‐on experience running large cloud‐based website environments with GKE, service mesh, load balancers, CDN/ ...

Production Support Engineer

Hiring Organisation
Balyasny Asset Management
Location
London, UK
Employment Type
Full-time
data related issues that impact trade flow. · Familiarity with enterprise scheduling tools such as Active Batch, Airflow, Control-M, AutoSys · Experience using/building monitoring on platforms like ITRS Geneos, Zabbix ...

Service Engineer

Hiring Organisation
Microsoft
Location
London, UK
Employment Type
Full-time
improve the customer experience. You will partner closely with engineering and business stakeholders to turn customer signals into product improvements, scalable troubleshooting guidance, proactive monitoring, and readiness for new features. This role is ideal for someone who combines technical curiosity, strong judgement, customer empathy, and a bias … manual work, recurring issue patterns, and high-volume support scenarios that can be simplified, standardized, scripted, or automated. Design, improve, and maintain automation, dashboards, proactive checks, alerting signals, diagnostic workflows, and AI-assisted support tools that reduce manual troubleshooting, improve consistency, and accelerate issue resolution. Own complex customer, partner ...

Head of Technology Operations & Service Delivery

Location
City of Westminster, England, United Kingdom
provide transparent reporting on vulnerabilities and remediation. Infrastructure, cloud & end-user technology Lead cloud infrastructure, networks, end-user computing, identity services, collaboration platforms, monitoring, backup and recovery. Maintain appropriate availability, capacity, performance, patching, lifecycle and asset management disciplines. Ensure production services have current documentation, monitoring and tested operating … utilisation, demand, optimisation opportunities and cost drivers. Work with engineering, architecture, finance and suppliers to improve efficiency without weakening control. Observability & operational intelligence Develop proactive monitoring and end-to-end observability across critical services. Use operational data to identify service degradation, capacity constraints and control weaknesses before material ...

Senior DevOps Engineer (Java)

Hiring Organisation
Netcompany
Location
London, UK
Employment Type
Full-time
/CD pipeline creation and maintenancePractical experience using Terraform(IaC), Kubernetes and HelmExperienced with multiple ecosystem databases, architecture and integrationExperience with monitoring tools to ensure system reliability and performance. Practical experience using AWS CloudCI/CD: Proficiency in setting up and maintaining pipelines, particularly with Jenkins. Database Management: Experience … Python)GitOps: Understanding of GitOps principles. Java & Maven: Must have experience with Java programming and Maven for buildsCore Principles: A mindset focused on automation, proactive monitoring, continuous improvement, and defining best practices. EssentialsMust have willingness to travelMust have the right to work in the UKMust be able ...

Senior DevOps Engineer (Python)

Hiring Organisation
Netcompany
Location
London, UK
Employment Type
Full-time
Kubernetes and HelmExperience in PythonExtensive experience of cloud computing, and best practices/emerging technologiesExperienced with multiple eco-system databases, architecture and integrationExperience with monitoring tools to ensure system reliability and performance. Kubernetes & Helm: Strong understanding of container orchestration with Kubernetes and package management using Helm. Cloud Services … e.g., Python).GitOps: Understanding of GitOps principles. Java & Maven: Experience with Java programming and Maven for builds. Core Principles: A mindset focused on automation, proactive monitoring, continuous improvement, and defining best practices. EssentialsMust have willingness to travelMust have the right to work in the UKMust be able ...

BizOps Manager (sales focused)

Location
Greater London, England, United Kingdom
sales-led revenue ops end to end: compensation, forecasting, growth reviews, and the metrics layer everyone trusts. and you'll build the proactive systems that surface problems before anyone has to ask. It's a blank slate. You're building the systems, definitions and source of truth where none … numbers reach people without being pulled. Treat the review as a decision making tool: leading indicators, levers and gaps, not just numbers. Triggers & proactive monitoring: build alerting on the things that actually move sales led revenue: deals slipping stages, forecast diverging from pacing mid quarter, reps falling behind ...

backend developer for retail POS platforms

Location
Greater London, England, United Kingdom
Contribute to cloud-native, event-driven architecture for global retail operations Build scalable integration patterns using Kafka, RabbitMQ, and serverless workflows Support observability through proactive monitoring, alerting, and performance tuning with tools such as NewRelic Ensure reliability through contract tests, load tests, and test data pipelines Collaborate with … Kafka and asynchronous data processing Solid understanding of databases such as PostgreSQL and infrastructure-as-code tools such as Terraform Experience designing, scaling, and monitoring cloud-native services on AWS, GCP, or Azure Familiarity with OAuth2, token-based authentication, and secure service design Strong debugging and telemetry skills across ...

backend developer for POS platforms

Location
Greater London, England, United Kingdom
Contribute to cloud-native, event-driven architecture for global retail operations; Build scalable integration patterns using Kafka, RabbitMQ, and serverless workflows; Support observability through proactive monitoring, alerting, and performance tuning with tools such as NewRelic; Ensure reliability through contract tests, load tests, and test data pipelines; Collaborate with … asynchronous data processing; Solid understanding of databases such as PostgreSQL and infrastructure-as-code tools such as Terraform; Skilled in designing, scaling, and monitoring cloud-native services on AWS, GCP, or Azure; Familiarity with OAuth2, token-based authentication, and secure service design; Strong debugging and telemetry skills across logs ...

Network Engineer (Cisco/Meraki)

Hiring Organisation
Taylor James Resourcing
Location
London, UK
Employment Type
Full-time
others. Create and maintain site documentations and technical documentations. Develop knowledge in a specific technology area, to become a trusted SME in that field. Proactive monitoring of IT systems and preventative measures taken to reduce system downtime. Write post incident review documents and suggest/implement recommendation plans … provide performance statistics and reports, develop strategies for maintaining and improving core infrastructure. Adherence to all IT security policies and assistance in enforcing and monitoring of IT security policies. Be able to spot potential vulnerabilities and suggest resolution. Evaluate, design, maintain, infrastructure systems, including LANs, WANs, Internet, intranet, security ...