1,376 to 1,400 of 2,378 Observability Jobs in London

Senior Software Developer

Location
Greater London, England, United Kingdom
includes:* Improve how we connect to an Exchange or Broker (eg. CME, ICE, EEX, etc)* Implement customer facing features via SDKs and APIs* Improve observability in hot paths* Implement a new Automated Trading feature* Benchmarking code to reduce latency and improve throughput* Mentoring and coaching junior engineers* Design, plan … RabbitMQ* We run some services on dedicated hardware and on AWS* We use Azure DevOps for build/deploy pipelines and story management* Our observability stack is Splunk, Grafana, Pyroscope and Prometheus **You**For this position, you will likely be a very successful team member if:* You like working ...

Data Solutions Architect

Location
Greater London, England, United Kingdom
using modern data integration technologies. Establish standards for analytics engineering, dbt, the modern data stack, DataOps, testing, documentation, CI/CD, Infrastructure as Code, observability and automated deployment. Embed metadata, lineage, privacy, security, regulatory compliance, monitoring and data quality controls into platform and data product designs. Lead technical workstreams, mentor … scale modern data stack architectures, including framework design, governance, testing, documentation and deployment standards. Ability to establish DataOps practices that improve automation, quality and observability across data pipelines. Software Engineering & DevOps Strong software engineering background, with proficiency in Python and SQL and knowledge of architectural principles, design patterns, testing frameworks ...

Senior Software Developer

Location
City Of London, England, United Kingdom
you. Improve how we connect to an Exchange or Broker (eg. CME, ICE, EEX, etc) Implement customer facing features via SDKs and APIs Improve observability in hot paths Implement a new Automated Trading feature Benchmarking code to reduce latency and improve throughput Mentoring and coaching junior engineers Design, plan … RabbitMQ We run some services on dedicated hardware and on AWS We use Azure DevOps for build/deploy pipelines and story management Our observability stack is Splunk, Grafana, Pyroscope and Prometheus For this position You like working in a cross-functional team with Product Managers, Testers and DevOps engineers ...

Staff Software Engineer-AI

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency Establish engineering best practices through hands-on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement Partner with ...

Staff Software Engineer-AI

Hiring Organisation
Hackajob Ltd
Location
Westminster, Greater London, UK
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency Establish engineering best practices through hands-on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement Partner with ...

Lead SRE - Chase UK

Location
Greater London, England, United Kingdom
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … knowledge of microservice infrastructure components, including service discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes ...

Lead GCP Engineer

Location
City Of London, England, United Kingdom
practices across the platform. Drive adoption and maturity of CI/CD pipelines, Infrastructure as Code and automated deployment processes. Implement and improve platform observability, monitoring, alerting and operational support capabilities. Develop and maintain reusable automation patterns and engineering standards. Leverage Terraform and Infrastructure as Code principles to ensure repeatable … Infrastructure as Code expertise. Strong experience with CI/CD tooling and deployment automation. Experience building and operating resilient cloud platforms. Strong understanding of observability, monitoring, incident management and operational excellence. Data Engineering & Integration Experience designing secure data integration solutions. Knowledge of APIs, databases, data pipelines and cloud‐based data ...

Technology Lead (Remote - UK)

Location
Greater London, England, United Kingdom
term sustainability Manage and reduce technical debt strategically Design cost‐efficient, cloud‐native solutions Engineering Excellence Set and uphold standards for code quality, testing, observability, security, and documentation Champion automated testing and “shift‐left” quality practices Promote DevSecOps principles to embed security and compliance into development Lead through thorough code … understanding of system design, distributed systems, scalability, APIs, and data modelling Experience with Infrastructure as Code (Terraform or similar) Strong CI/CD and observability experience Experience implementing robust testing strategies Comfortable working cross‐functionally in Agile environments Strong communication skills and ability to operate effectively in ambiguous situations What ...

Senior Manager- Software Engineering

Location
Greater London, England, United Kingdom
microservices, including REST and event‐driven patterns, and enforce best practices for versioning, contracts, and backward compatibility. Advance operational excellence by defining SLOs, improving observability and alerting, hardening on‐call procedures and runbooks, and leading incident response and post‐mortems. Solve complex distributed system challenges (such as throughput, latency, consistency … with microservices, API design, and event‐driven systems, as well as experience with containerization and orchestration (Docker, Kubernetes). Strong operational mindset: expert in observability, monitoring, incident response, performance engineering, and adherence to security best practices. Familiarity with CI/CD pipelines, automated testing strategies (unit, integration, e2e), and modern ...

Safety Engineer - Free Tier Abuse

Location
Greater London, England, United Kingdom
models into production systems Architect robust APIs, data pipelines, and service architectures supporting real‐time and batch moderation workflows Implement comprehensive monitoring, alerting, and observability systems; establish SLIs, SLOs, and performance benchmarks Partner with ML engineers to translate research models into production‐ready systems and integrate them across our product … pipelines, and Python expertise (asynchronous Python, backend frameworks) Infrastructure & DevOps proficiency: cloud platforms (AWS/GCP), containerization (Docker/K8s), CI/CD pipelines Observability mindset with experience in monitoring tools (Prometheus, Grafana) and building observable systems Track record of taking products or systems from 0→1 with measurable impact ...

SRE Engineer

Location
Greater London, England, United Kingdom
engineering standards Driving infrastructure‐as‐code and automation across Azure and on‐prem Improving image bakery pipelines for secure, repeatable server builds Embedding observability using metrics, logs, traces, and effective alerting Ensuring all practices align with ISO 27001 and internal security frameworks Managing automated patching, vulnerability remediation and configuration compliance … code (Terraform, ARM/Bicep), configuration management (Ansible, PowerShell DSC), and CI/CD tooling (Azure DevOps, GitHub Actions) Experience with monitoring and observability stacks Solid understanding of OS fundamentals (Windows/Linux), security, networking Background in scripting or software development (PowerShell, Python, Go) Experience with containers and orchestration (Docker ...

Principal Engineer

Location
Greater London, England, United Kingdom
their full potential.* Modern Software Practices: Champion SOLID principles, automated testing and CI/CD best practices, ensuring deployments are fast, safe and seamless.* Observability & Performance: Help teams build highly observable systems using DataDog and other tools, making sure we can see and fix issues before they impact users.* Collaboration …/CD & Infrastructure as Code: Experience deploying production systems using Terraform and Helm, and experience with GitHub Actions as a CI/CD platform.* Observability Mindset: Believes in measuring everything, and has worked with DataDog (or similar) to ensure teams have visibility into system health.* Software Craftsmanship: Writes high-quality ...

Platform / DevOps Engineer

Location
City of Westminster, England, United Kingdom
Manager, EventBridge provisioned as code, sized sensibly, and cost‐aware (you'll make the calls on things like NAT vs VPC endpoints). Own observability and reliability. CloudWatch alarms and dashboards, SNS alerting, data freshness and quality signals, and automated recovery for the pipelines that need it. When something breaks … deliberate applies, and easy rollbacks. Everything is code. Infrastructure, pipelines, access, and policy all live in version control and ship through review. Observability first. If we run it, we can see it - and we get told before our users do. You own what you ship. Strong ownership, low ceremony. ...

Sr. Software Engineer

Hiring Organisation
Meltwater Group
Location
London, UK
Employment Type
Full-time
while architecting systems that handle high-throughput data pipelines. Participate in building robust export pipelines, streaming architectures, webhook integrations and MCP servers. Maintain high observability and reliability standards using tools like Coralogix, CloudWatch, and Grafana. Participate in on-call rotation and incident response for owned services. What You'll Bring4+ … Swagger, static site generators).Familiarity with authentication, API gateways, and rate limiting strategies. Experience in compliance standards for APIs and data handling. Experience with observability tools and practices. Our Tech Stack: Languages: Golang (primary) with some TypeScriptInfrastructure: AWS, S3, Lambda, SQS, SNS, CloudFront, Kubernetes (Helm), Kong API GatewayDatabases: Postgres, Redis ...

Staff Cloud SRE – AI/ML Platform & GPU Compute London, United Kingdom on-site

Location
Greater London, England, United Kingdom
escalation, communications, and root cause analysis. Translate post-incident learning into durable architectural or automation improvements. Continuously reduce alert noise and recurring operational burden. Observability & Operational Excellence Design and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. Build dashboards that reflect real user-centric platform … Python, Go, C++) with a bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g. Datadog, Prometheus, Grafana, OpenTelemetry). Clear communication skills, including leading incidents, writing postmortems, and influencing teams to prioritise reliability improvements. Desirable skills Familiarity ...

Lead SRE - Chase UK

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … knowledge of microservice infrastructure components, including service discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes ...

Staff SRE, AI Infrastructure

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
escalation, communications, and root cause analysis. Translate post-incident learning into durable architectural or automation improvements. Continuously reduce alert noise and recurring operational burden. Observability & Operational ExcellenceDesign and operate monitoring, logging, tracing, and alerting systems that enable rapid detection and recovery. Build dashboards that reflect real user-centric platform health … Python, Go, C++) with a bias toward automation. Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale. Experience designing and operating observability stacks (e.g. Datadog, Prometheus, Grafana, OpenTelemetry).Clear communication skills, including leading incidents, writing postmortems, and influencing teams to prioritise reliability improvements. Desirable skillsFamiliarity with infrastructure ...

Senior Full Stack Engineer

Location
Greater London, England, United Kingdom
technical designs and code produced by delivery partners, identifying risks and opportunities for improvement Champion engineering excellence across code quality, testing, continuous delivery, security, observability and platform reliability Collaborate closely with AI, Machine Learning, Product, Infrastructure and DevOps teams to deliver integrated platform capabilities and share knowledge across the engineering … Familiarity with AI orchestration frameworks, intelligent workflows and emerging AI technologies Experience with data platforms and infrastructure that support AI workloads Knowledge of monitoring, observability and continuous delivery practices Experience within telecommunications or other large scale customer facing digital platforms What's in it for you? Competitive salary and bonus ...

Devops Engineer

Location
Greater London, England, United Kingdom
deployment, software repositories, databases, and web servers Own the patching and update lifecycle for managed systems Monitoring & Reliability Implement and maintain monitoring, alerting, and observability across both the existing VM estate and the new container environment Proactively identify risks, bottlenecks, and failure patterns before they impact users Define and track … e.g. Terraform, Ansible, Puppet, or similar) Solid scripting ability in Bash and at least one higher‐level language (Python preferred) Experience with monitoring and observability tooling (e.g. Prometheus, Grafana, Datadog, or similar) Strong incident diagnosis skills—able to work from vague symptoms to root cause using logs, metrics, and reasoning ...

Staff Platform Site Reliability Engineer

Hiring Organisation
Index Exchange
Location
London, UK
Employment Type
Full-time
about its architecture. Must Have8+ years in platform engineering, SRE, infrastructure engineering, or DevOps. Deep experience with Linux internals: kernel tuning, network stack, system observability, security. Strong Kubernetes expertise: cluster lifecycle, networking, storage, RBAC, multi-cluster—across bare-metal and cloud (EKS, GKE).Infrastructure-as-code at scale: Terraform, Ansible … driving technical strategy across teams—not just executing within one. Valuable ExperienceDistributed storage systems (e.g. Ceph)Big data infrastructure: Hadoop, Spark, HBase, Kafka. Observability stack design: Prometheus, Grafana, ELK, Mimir, Loki, Tempo. Secrets management (Vault), certificate management, access control at scale. Hybrid cloud architectures: federating public cloud (AWS, GCP) with ...

Lead Data Engineer

Hiring Organisation
Quilter
Location
London, UK
Employment Type
Full-time
Owners, Architects, and stakeholders to prioritise and deliver work effectively. Drive adoption of modern engineering practices including automation, testing, CI/CD, source control, observability, and infrastructure-as-code. Contribute to roadmap planning by identifying technical opportunities, risks, dependencies, and improvement initiatives. Mentoring & Capability DevelopmentProvide coaching, mentoring, and technical guidance … products and platform capabilities across one or more business domains. Experience implementing and promoting data quality frameworks, service level objectives (SLOs), data contracts, and observability tooling. Strong understanding of data governance principles, including data lineage, metadata management, cataloguing, security, and access controls. Experience ensuring compliance with regulatory and organisational requirements ...

Senior Data Engineer

Hiring Organisation
TripAdvisor
Location
London, UK
Employment Type
Full-time
About TripadvisorThe Tripadvisor Group connects people to experiences worth sharing, and aims to be the world's most trusted source for travel and experiences. We leverage our brands, technology, and capabilities to connect our global ...

Senior Backend Engineer - Java

Hiring Organisation
Capco
Location
London, UK
Employment Type
Full-time
Senior Backend Engineer – JavaLocation: London (Hybrid) | Practice Area: Technology & Engineering | Type: PermanentPower the platforms behind modern finance. Shape cloud-native solutions with Java at their core. The RoleWe're seeking a Senior Backend Engineer (Java ...

Lead Data Engineer

Hiring Organisation
Source Group International
Location
London, UK
Employment Type
Full-time
Reference: 56862Location: London, United KingdomSalary: 130,000 per annumIndustry: EngineeringOverview We are seeking a Lead Data Engineer to design, build and own scalable data platforms that enable high-quality analytics, reporting and data-driven products. ...

Senior Software Engineer - iCloud Platform - Observability

Location
Greater London, England, United Kingdom
services platform and infrastructure. You will be working on foundational systems that power iCloud services, including distributed data platforms, storage systems, and a unified observability infrastructure that standardizes monitoring and telemetry approaches across all iCloud services serving billions of customers.iCloud manages data and services at massive scale! Our unified observability … health of all iCloud services, providing comprehensive telemetry collection, real-time processing, and sophisticated analysis capabilities that span billions of active Apple customers. This observability ecosystem is purpose-built to deliver highly scalable and performant solutions, that maintain the highest standards of user privacy and security.We are a world-class ...