876 to 900 of 2,373 Observability Jobs in London

Staff / Senior Staff Security Engineer, Vinted Pay

Location
Greater London, England, United Kingdom
leader. Excellent written and spoken English. Advantage: experience building and running systems at massive scale (2+ million requests per minute), deep knowledge of observability tooling (Kibana, Grafana, Prometheus), and a passion for introducing new practices. Advantage: AWS security depth (IAM, KMS, multi‐account and multi‐region architecture), or hands ...

Senior Data Engineer

Location
Greater London, England, United Kingdom
data moves between our application, warehouse and product surfaces. Build clean, well-tested and reusable solutions using Python, SQL and dbt. Improve data quality, observability and reliability so teams can trust the data they use. Work with product, engineering, commercial teams and educational experts to understand problems, shape solutions ...

Tech Lead Manager - AI Squad

Location
Greater London, England, United Kingdom
wider strategy and commitments. Operational Excellence: Setting high standards and driving engineering best practice across the AI platform stack. This includes focusing on observability (monitoring, logging, tracing) for the cloud‐native solution, participating in peer reviews, and facilitating the maintenance of engineering standards, reliability, and performance. Squad Empowerment: Contributing ...

Staff / Senior Staff Security Engineer, Vinted Pay

Hiring Organisation
Vinted
Location
London, UK
Employment Type
Full-time
leader. Excellent written and spoken English. Advantage: experience building and running systems at massive scale (2+ million requests per minute), deep knowledge of observability tooling (Kibana, Grafana, Prometheus), and a passion for introducing new practices. Advantage: AWS security depth (IAM, KMS, multi-account and multi-region architecture), or hands ...

Software Engineer, Cyber and Autonomous Systems Team

Location
Greater London, England, United Kingdom
shared components from design through implementation, iteration, and maintenance. Write production-quality code: build scalable, robust, maintainable Python software with strong testing, documentation, and observability practices. Onboard a cyber range: deploy a new cyber range on AISI's evaluation infrastructure (for example, Proxmox) and verify it meets the internal Evaluation ...

Head of AI Solutions, COO Technology - MD (C16)

Hiring Organisation
Citigroup
Location
London, UK
Employment Type
Full-time
agentic safety at enterprise scaleAI/ML Engineering & MLOps: Full AI/ML lifecycle ownership: training pipelines, model deployment, versioning, monitoring, drift detection, observability (e.g., Weights & Biases, Arize), and lifecycle management using platforms such as MLflow, Vertex AI, or SageMakerCloud AI Platforms: Demonstrated deployment of AI workloads ...

Senior Solution Architect - Email & Collaboration Security

Location
Greater London, England, United Kingdom
Domain ECS covers the following product and capability areas: Email threat detection and prevention Email security efficacy Collaboration security DMARC analyzer Data platform and observability Analysis and Response End user application integration Internal operations enablement What You Will Do Lead from Problem Space, Not Solution Space The role engages ...

Staff Software Engineer, Runtime Systems

Location
Greater London, England, United Kingdom
control-plane components. Build adapters and integrations for heterogeneous execution environments. Diagnose behaviour across application, orchestration, cluster, and infrastructure boundaries. Improve the reliability, observability, and debuggability of distributed workload execution. Work closely with Go, Kubernetes, and infrastructure engineers to turn architecture into production systems. Performance & Experimentation Develop rigorous ways ...

Product Engineer (all levels)

Location
Greater London, England, United Kingdom
help shape that), here’s a current snapshot: Full‐stack TypeScript, React, Postgres, and Temporal for long‐running orchestration. Infra: AWS, Terraform, strong observability via Sentry and Datadog (full‐stack, not just logs). Data: we lean heavily on Snowflake, Omni, dbt, Fivetran, Amplitude and Segment — and we’ve even ...

Senior DevOps Engineer – Build Pipelines

Location
Greater London, England, United Kingdom
Application Engineering Teams To Design, Implement And Roll Out These Standards, With a Strong Focus On Repeatable CI/CD Automation Security Quality gates Observability Deployment consistency Developer experience Platform engineering principles This is a hands-on engineering role , rather than a traditional operational DevOps position. The Role … signing Security and quality thresholds Automated quality gates The objective is for security and quality controls to enabled by default across the application estate. Observability Integrate pipeline and deployment processes with appropriate monitoring and observability tooling. Help teams understand build and deployment health. Work with technologies such as: Prometheus Grafana ...

Infrastructure Software Engineer — Cloud & Observability

Location
Greater London, England, United Kingdom
A leading fintech company in the United Kingdom is seeking a Software Engineer specialized in Infrastructure to support deployment and maintenance of cutting-edge software. The role involves building reliable and scalable applications, developing tools ...

Site Reliability Engineer I: Cloud-Native & Observability

Location
Greater London, England, United Kingdom
Axon is seeking an experienced Site Reliability Engineer for its Real Time Operations in London. You will contribute to building reliable cloud-native services, collaborate with RTO engineering teams, and enable product teams to scale ...

Test Environment Manager

Location
Greater London, England, United Kingdom
configuration, and teardown of test environments. Integrate environment automation seamlessly into CI/CD pipelines to enable on-demand, self-service environment delivery. Reliability & Observability Define and maintain Service Level Objectives (SLOs) and key Service Level Indicators (SLIs), such as environment availability, provisioning time, and stability metrics. Monitor environment health … using observability tools (Prometheus, Grafana, Splunk, etc.) and proactively identify and resolve performance issues or bottlenecks. Incident & Problem Management Lead incident response for environment-related issues, driving quick resolution and facilitating blameless post-mortems. Implement permanent fixes based on root cause analysis and reduce repeat incidents. Automation & Toil Reduction Identify ...

Platform Engineer (DevOps / MLOps Focus)

Hiring Organisation
The Portfolio Group
Location
London, United Kingdom
Employment Type
Permanent
Salary
£100000/annum
environments for production workloads. Developing and managing Infrastructure as Code using Terraform. Supporting CI/CD pipelines and platform automation initiatives. Improving platform reliability, observability and scalability. Collaborating closely with Software Engineers, Data Engineers and ML teams to optimise deployment workflows and infrastructure performance. Essential experience: Strong commercial experience with … available, scalable production environments. Nice to have: Experience with Kubeflow and ML platform tooling. Exposure to AI, machine learning or GenAI projects. Experience with observability tooling such as Prometheus, Grafana or OpenTelemetry. Experience working within regulated or enterprise environments. This is a fantastic opportunity to join a high-profile ...

Senior Platform Engineer

Hiring Organisation
Workable Software Limited
Location
South East London, London, United Kingdom
Employment Type
Permanent, Work From Home
Agents: Develop the infrastructure required to run increasingly agent-driven AI workflows reliably in production, including state management, task execution, model routing, retries and observability Support AI Infrastructure: Work closely with our AI engineers supporting production NLP, LLM and quantitative model workloads, including our in-house GPU infrastructure. Automate Infrastructure … Build and manage Infrastructure-as-Code using Pulumi and automate deployments through GitHub Actions. Own Reliability & Observability: Build monitoring, logging, tracing and alerting across our data, model, service and agent infrastructure. Requirements Degree in a related technical field: Computer science, software engineering or similar. 7+ years of professional experience ...

Senior Data Engineer

Location
Greater London, England, United Kingdom
them into robust technical solutions. Contribute to solution architecture, platform design, and technology selection decisions. Implement software engineering and DataOps best practices including testing, observability, CI/CD, version control, and infrastructure automation. Develop and optimise data models, transformations, and storage layers to ensure performance, reliability, and scalability. Collaborate with … Solid understanding of data lakehouse, warehouse, and medallion architecture patterns. Experience applying DevOps and DataOps best practices including CI/CD, automated testing, monitoring, observability, and release management. Experience with Infrastructure as Code technologies such as Terraform, Pulumi, CloudFormation, or Bicep. Strong understanding of distributed systems, data modelling, data governance ...

DevOps Engineer – Security & Intelligence

Location
Greater London, England, United Kingdom
pipelines to support continuous delivery Developing Infrastructure as Code and automated platform capabilities Supporting AWS-based environments, including Kubernetes, OpenShift, EKS and ECS Implementing observability, monitoring, logging and alerting for live services Supporting SRE practices, cloud migration activities and production platform operations Job Responsibilities Design, implement and maintain secure … Ansible) Support containerised platforms and environments, including Kubernetes, OpenShift, EKS and ECS Embed security controls, quality gates and compliance checks into DevSecOps workflows Implement observability, monitoring, logging and alerting to support reliable live services Contribute to SRE practices, improving service reliability, performance and operational resilience Automate build, deployment and platform ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
overhead. Support and enhance CI/CD platforms and engineering workflows, including the safe promotion of infrastructure and platform changes. Improve platform reliability, security, observability, performance and operational excellence. Partner with software engineers, quantitative developers, research teams, Technology Operations and security colleagues to deliver secure, scalable and resilient platforms. Education … experience with AWS and cloud‐native infrastructure. Experience with Kubernetes, including Amazon EKS, and Docker or other container runtimes. Experience with monitoring and observability tools such as Grafana, Prometheus, Loki or Datadog. Experience with artifact‐management platforms such as JFrog Artifactory. Knowledge of AWS Batch, AWS Step Functions, AWS Identity ...

Senior AWS Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£60000 - £70000/annum Bonus, Medical Care
Have: Practical experience managing large-scale Kubernetes clusters; certifications in Kubernetes are a strong bonus Hands-on familiarity with the Grafana Observability Suite, including tools like Loki, Mimir, and Tempo Background in administering or developing with popular monitoring and automation tools such as Splunk, Datadog, PagerDuty, or Rundeck Experience using … with tools such as Jenkins, GitLab CI/CD, or CircleCI Strong understanding of containerization (e.g., Docker, Kubernetes) and microservices architecture Skilled in using observability and monitoring tools such as Prometheus, Grafana, ELK stack, or AWS CloudWatch Excellent analytical and troubleshooting abilities, especially within complex distributed systems Proven experience handling ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
people freedom while keeping us in control. What you’ll be doing from day one: Owning our infrastructure as code in Terraform, plus alerting, observability (Prometheus, Grafana) and reliability, including load testing and disaster recovery exercises. Making CI/CD faster (GitHub Actions, ArgoCD) and taking obstacles out of engineers … ideally multi-region. Hands-on with GCP, with Azure a bonus. You understand how model consumption works through each cloud. Deep experience with Terraform, observability and CI/CD, and you write solid Python (Elixir is a bonus, or you’re keen to learn). A clear communicator ...

Principal Site Reliability Engineer, Infrastructure Observability

Location
Greater London, England, United Kingdom
opportunity to grow and make a difference in ways that matter to you. Role Summary In this role as Principal Site Reliability Engineer, Infrastructure Observability you will help formulate, develop, and implement a team of Site Reliability Engineers (SREs) focused on the observability, sustainability, scalability, measurability and recoverability … Proficiency with understanding and explaining incident situations and their recovery plans to prevent recurrence Knowledge/experience driving dashboard standardization across the ecosystem for observability, APM and infrastructure monitoring, and application‐specific logging Knowledge/experience with observability tools such as New Relic, SolarWinds DPA, Elastic Stack, Prometheus, Grafana, Splunk ...

Senior Backend Engineer (.NET & Python)

Location
Greater London, England, United Kingdom
ship, contributing to testing, troubleshooting and continuous improvement. Collaborate with Product, Design and Engineering teams to deliver customer-focused solutions. Improve platform reliability, observability, performance and security. Help evolve Benifex's AI‐assisted development practices and support other engineers across the team. What are we looking for? Commercial experience developing … backend applications using C#, .NET/.NET Core, Python and SQL. Production GenAI experience, including technologies such as MCP, RAG, agent orchestration, evaluations and observability tooling. An understanding of responsible AI, software quality, performance and security best practices. Experience delivering features end-to-end within modern engineering environments. Exposure ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
integrations, and human‐in‐the‐loop controls evolve across the stack. Continuously identify and exploit opportunities to improve performance, reliability, and user experience, using observability and analysis to find signals in noisy systems. Navigate confidently across legacy and greenfield contexts, applying AI tooling pragmatically to modernise where it matters most. … across both human-written and AI-generated code, embedding security validation into the development pipeline rather than treating it as an afterthought. Own system observability and reliability, moving from reactive alerting to proactive, model-assisted incident prevention. Participate in on-call rotation, bringing the same rigour to incident response that ...

Platform Engineer, SDO London, United Kingdom

Location
Greater London, England, United Kingdom
optimise cloud systems for performance, reliability, and cost efficiency. Assist in managing containerised workloads and orchestration platforms (e.g. Docker, Kubernetes). Implement and maintain observability tools (logging, metrics, alerting) to ensure system health and rapid incident response. Work with engineering teams to ensure infrastructure meets application requirements and supports scalable … cloud architecture, and system reliability. Strong troubleshooting and problem‐solving skills. Desirable: Experience with containerisation and orchestration (Docker, Kubernetes). Familiarity with monitoring and observability tools (e.g. Prometheus, Grafana, ELK). Experience working with Linux systems and shell scripting. Programming or scripting experience (e.g. Python, Bash, Go, or similar). ...

Senior Network Site Reliability Engineer

Location
Greater London, England, United Kingdom
Build and lead all aspects of our CI/CD pipelines (GitLab/GitHub) and provide automated solutions for IaC deployment. Implement automation and observability across infrastructure resources. Bring your own ideas for improving our infrastructure stack and implement them. Take part in our operational rotation to react and resolve … teams Nice to have Experience with containers and Kubernetes Familiarity with AWS governance and security controls, including SCPs and IAM policies Experience improving reliability, observability, performance, or incident response processes Exposure to large-scale CDN, edge, or traffic-routing environments Experience working in globally distributed infrastructure or platform teams Basic ...