51 to 75 of 98 Observability Jobs in the City of London

Lead Network Engineer

Hiring Organisation
DGH Recruitment
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£95,000
Microsoft Azure networking including virtual networks, subnets, NSGs, route tables and hybrid connectivity. - Experience supporting network resilience, high availability, disaster recovery testing, monitoring and observability across enterprise LAN, WAN, wireless, data centre and cloud services. - Experience producing and maintaining HLDs, LLDs, standards, implementation plans and operational documentation. Lead Network Engineer ...

Networks Technical Lead (SD-WAN / SASE)

Hiring Organisation
DGH Recruitment
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£95,000
Microsoft Azure networking including virtual networks, subnets, NSGs, route tables and hybrid connectivity. - Experience supporting network resilience, high availability, disaster recovery testing, monitoring and observability across enterprise LAN, WAN, wireless, data centre and cloud services. - Experience producing and maintaining HLDs, LLDs, standards, implementation plans and operational documentation. Networks Technical Lead ...

Product Lead - AI

Hiring Organisation
Trismik
Location
City of London, London, United Kingdom
discovery or user research experience. Comfortable taking ambiguous problems to evidence-backed decisions without needing consensus on every call. Enterprise AI deployment, evaluation or observability experience a plus. How you work You listen before proposing solutions and change your mind when the evidence changes. You want real decision-making authority ...

Software Engineer, Trading - Cumberland Systematic

Location
City Of London, England, United Kingdom
operation with high availability requirements. You will be expected to design and develop trading systems, market data connectivity, execution algorithms, research infrastructure, monitoring and observability tooling, and integrations with DRW’s core services. The team’s existing systems are written in C++ and Python. Candidates should have strong initiative ...

Software Engineering Manager - ThousandEyes

Location
City Of London, England, United Kingdom
technology portfolio and beyond, helping customers deploy at scale while also delivering AI-powered assurance insights within Cisco's leading Networking, Security, Collaboration, and Observability portfolios. Your Impact You will be the Engineer Manager on the Endpoint team. Our goal is to ensure our customers have an unparalleled ability ...

AI Product Analyst (AI Metrics & Model Evaluation)

Hiring Organisation
The Portfolio Group
Location
City of London, London, Castle Baynard, United Kingdom
Employment Type
Permanent
Salary
£80000 - £85000/annum
metrics that matter, retrieval quality, correctness of output, and what users go on to do with what they're given Owning production quality observability, from thumbs-down and regeneration rates through to abandoned tasks and drift in how the product is being used Building the analysis and dashboards the team ...

Principal Engineer, Mobile Platform & Backend Systems

Location
City Of London, England, United Kingdom
security implications of client‐service interactions. Experience improving customer‐visible reliability and performance across client and backend systems, with practical expertise in observability and production validation. A track record of leading consequential technical change across teams through influence rather than reporting authority. Product and stakeholder leadership that connects technical decisions ...

Site Reliability Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
play a key role in building highly reliable, scalable, and observable infrastructure. This is a hands-on role focused on AWS, Kubernetes, Terraform, observability, monitoring, and automation, working closely with software engineering teams to improve platform reliability and developer experience. You'll have genuine ownership and the opportunity to influence … infrastructure using Terraform and Infrastructure as Code principles Develop and optimise CI/CD pipelines using GitHub Actions Build and improve comprehensive monitoring and observability across the platform Implement and maintain effective logging, metrics, tracing, alerting, and dashboards Define and improve SLIs, SLOs, and reliability metrics Proactively identify and resolve ...

Senior Forward Deployment Engineer

Hiring Organisation
Luxoft
Location
City of London, London, United Kingdom
analysis, upgrading Java and NPM runtimes, modernizing Spring and legacy middleware applications, improving CI/CD pipelines, containerizing applications, automating deployments, and introducing standard observability and resilience patterns. The Expert FDE is expected to lead complex engagements, work directly with development and client stakeholders, define the technical remediation approach, implement … testing, release, resilience, and legacy technology challenges with development teams. Assess application code, dependencies, runtime environment, test coverage, deployment architecture, CI/CD pipelines, observability, and operational risks. Write, debug, review, and enhance production-quality code and configuration throughout engagements. Define and implement practical modernization and remediation plans with clear ...

Senior UI Platform Engineer – React/TypeScript Foundations

Location
City Of London, England, United Kingdom
packages, components, and tooling. You will create starter templates, define standards for structure, testing, deployment, and supportability, and partner with other teams on authentication, observability, and frontend/backend integration patterns. This is a platform engineering role spanning teams and regions. #J-18808-Ljbffr ...

Senior Platform Engineer: Backend & Developer Tools

Location
City Of London, England, United Kingdom
platform capabilities that improve how teams build, deploy, and operate software. The ideal candidate will create reusable services, frameworks, and tooling, spanning API enablement, observability, and governance. This is a high-ownership position within a fast-moving engineering environment. #J-18808-Ljbffr ...

Site Reliability Engineer – Scalable Trading Infra

Location
City Of London, England, United Kingdom
class hedge fund in London to hire a Site Reliability Engineer who bridges software engineering and infrastructure. You will design scalable, automated platforms, improve observability, and ensure production systems achieve top-tier availability for trading and research environments. Working with software engineers, researchers, and traders, you will build tooling, tackle ...

Data Architect

Hiring Organisation
Experis
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
£800 - £860 per day
quality requirements. Produce High Level Designs, Low Level Design principles, Architecture Decision Records and design assurance material. Define non-functional requirements for performance, resilience, observability, scalability, security and maintainability. Support assurance, accreditation and security review activities with architecture evidence and design rationale. Provide technical governance and design oversight during build ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
City of London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Senior Frontend Engineer — Form-Heavy React/TS

Location
City Of London, England, United Kingdom
onboarding clients and accounts with KYC/KYB workflows. You will design, implement, and own features end-to-end, ensuring security, reliability, and observability from day one. Strong collaboration and clear communication with engineers and non-engineers are essential. #J-18808-Ljbffr ...

Data Platform Engineer

Location
City Of London, England, United Kingdom
consistency and reusability across environments. Build and optimize CI/CD pipelines using Azure DevOps and GitHub Actions to support rapid, reliable deployments. Implement observability practices including logging, metrics, and alerting using observability tools. Collaborate with the Lead Engineer and Architects to align implementation with platform standards and patterns. Provide … Fabric. Proven experience with infrastructure-as-code using Terraform and building CI/CD pipelines via Azure DevOps and GitHub Actions. Strong grasp of observability practices, including logging, metrics, alerting, and performance optimization. Deep understanding of cloud security, with experience applying secure-by-design principles in Azure and/ ...

Principal Platform Engineer

Hiring Organisation
Sanderson Recruitment
Location
City of London, London, United Kingdom
Employment Type
Permanent
persistence platforms Provide technical leadership and architectural guidance across multiple engineering teams Define engineering standards, platform roadmaps and best practices Drive automation, resilience, observability and operational excellence initiatives Support and mentor engineers through code reviews, coaching and technical leadership Collaborate with architects and stakeholders to translate business requirements into technical … automation and DevOps practices Experience mentoring engineers and providing technical leadership Key Technologies AWS Terraform Linux Cassandra Couchbase ScyllaDB Kafka CI/CD Pipelines Observability & Monitoring Platforms Distributed Database Technologies Nice to Have Experience with additional distributed persistence technologies Background in large-scale cloud-native environments Experience defining enterprise platform ...

Site Reliability Engineer

Location
City Of London, England, United Kingdom
development, validation, and optimization of configuration-as-code, improving delivery speed and reducing deployment risk. Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality … applications without these: Hands-on with Helm or Kustomize Experience with GitOps (e.g., Argo CD) Knowledge of secrets management (e.g., HashiCorp Vault) Experience with observability (metrics/logs/tracing) Why Cisco? At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations ...

Contract - Senior CXE Engineer - Amazon Connect

Hiring Organisation
INNOVATIVE TECH PEOPLE LTD
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
hands-on: model tier selection (Haiku vs. Sonnet vs. Opus), prompt caching, and token budgeting against containment-rate targets Diagnose AI agent performance using observability tooling (agent spans, and CloudWatch) that correlates contact flow logs, conversation transcripts, AI agent spans, tool executions, and token usage to isolate latency, cost … Functions, Kinesis) Infrastructure as code proficiency with AWS CDK or Terraform, including multi-account deployment patterns Experience shipping and supporting production systems, testing discipline, observability instrumentation, and incident debugging. ...

Principal Engineer I, Prepurchase Platform (Remote)

Location
City Of London, England, United Kingdom
demand on-sales. You will write production code daily, influence technical direction, and collaborate across multiple teams within the Prepurchase domain. You will drive observability, resilience patterns, and AI-assisted enhancements while embedding across services or working horizontally. #J-18808-Ljbffr ...

Site Reliability Engineer

Location
City Of London, England, United Kingdom
global trading and investment operations. Working closely with software engineers, quantitative researchers, traders, and infrastructure teams, you will be responsible for building automation, improving observability, and ensuring critical production systems operate at the highest levels of availability and efficiency. Key Responsibilities Design, build, and maintain highly reliable, scalable, and automated … infrastructure platforms. Drive improvements in system performance, monitoring, observability, and operational efficiency. Troubleshoot and resolve complex production incidents across distributed systems. Develop tools and automation to reduce operational overhead and improve platform resilience. Partner with engineering teams to improve system design, deployment processes, and operational readiness. Participate in incident management ...

Core Platform Developer

Location
City Of London, England, United Kingdom
reliability of internal systems. This person should be comfortable working across multiple areas of the stack, from service frameworks and API enablement to observability, governance, and developer workflows. This is a high-ownership role within a global, fast-moving engineering environment. Key Responsibilities Design and build shared backend services, frameworks … developer tooling that support internal application and service development. Develop common platform capabilities such as service templates, authentication and authorization patterns, API standards, observability integrations, error handling, and shared runtime utilities. Improve the developer experience through better tooling, automation, documentation, onboarding patterns, and paved-road workflows for engineering teams. Help ...

MLOps Engineer

Location
City Of London, England, United Kingdom
platform reliability. Key Responsibilities Design, deploy, and manage AI platforms and agent infrastructure Build and maintain CI/CD pipelines and DevOps workflows Implement observability, monitoring, and logging solutions Optimise performance, scalability, and cost efficiency Support AI teams with infrastructure, deployment, and integration Ensure platform security, compliance, and high availability …/CD, automation, and DevOps best practices Experience with Kubernetes/containerisation technologies Strong programming skills (e.g. Python, Go, Node.js) Experience with observability tools (e.g. OpenTelemetry, Datadog) Understanding of security, performance optimisation, and scalability Desirable Skills Experience working on AI/ML platforms or deployments Exposure to large-scale distributed ...

MLOps Engineer

Hiring Organisation
DGH Recruitment
Location
City of London, London, United Kingdom
Employment Type
Permanent
platform reliability. Key Responsibilities - Design, deploy, and manage AI platforms and agent infrastructure - Build and maintain CI/CD pipelines and DevOps workflows - Implement observability, monitoring, and logging solutions - Optimise performance, scalability, and cost efficiency - Support AI teams with infrastructure, deployment, and integration - Ensure platform security, compliance, and high availability …/CD, automation, and DevOps best practices - Experience with Kubernetes/containerisation technologies - Strong programming skills (e.g. Python, Go, Node.js) - Experience with observability tools (e.g. OpenTelemetry, Datadog) - Understanding of security, performance optimisation, and scalability Desirable Skills - Experience working on AI/ML platforms or deployments - Exposure to large-scale distributed ...