1,676 to 1,700 of 2,250 Observability Jobs in London

Software Engineer (EMS)

Location
Greater London, England, United Kingdom
improve performance. Distributed Systems: Maintain distributed, fault-tolerant systems with high availability across multiple regions, ensuring data consistency and low-latency communication. Monitoring & Observability: Implement real-time monitoring, telemetry and alerting systems for proactive issue detection and resolution. Collaboration: Work closely with product, QA and infrastructure teams to align technical … Azure) and tools for multi-region deployments . A working knowledge of databases and caches (PostgreSQL, SQLite, Redis, Zookeeper). Experience with observability tools (e.g. Prometheus, Grafana, ELK stack). Soft Skills: Strong analytical and problem-solving abilities. Excellent communication and collaboration skills, with a track record of working ...

Platform Engineer London, UK · Full time · Hybrid

Location
Greater London, England, United Kingdom
engineers. The role combines infrastructure and software development, with an approximate 65/35 split. You’ll work on cloud infrastructure, deployment automation, observability, CI/CD, and internal tooling. The goal is to make development and releases more reliable, efficient, and straightforward. What you’ll do Improve developer experience … Find the causes of slow builds, failed pipelines, and flaky tests. Develop internal tools, improve test infrastructure, and occasionally work on product features. Improve observability across our infrastructure and delivery workflows using OpenTelemetry, ELK. Leverage AI to improve engineering workflows by building cloud-based agent tools and helping teams adopt ...

Platform Technical Lead (Cloud Infrastructure & AI Enablement)

Hiring Organisation
Vodafone
Location
London, UK
Employment Type
Full-time
engineering technologies including coding assistants, agent frameworks, evaluation platforms and engineering automation tooling. Own engineering platform capabilities including CI/CD, deployment automation, observability, secrets management, platform security and self-service engineering tooling. Drive improvements in developer productivity, engineering effectiveness and delivery throughput through platform automation and AI enablement capabilities. … engineering tooling and AI-native software delivery approaches. Good understanding of LLM platforms, agents, evaluation frameworks and AI governance practices. Experience with observability, DevOps, CI/CD and engineering productivity platforms. Experience managing multi-region, multi-tenant and large-scale cloud platforms. ...

Principal Software Engineer (Hybrid)

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
other teams shift faster and more reliably Taken AI and agentic systems from prototype to production, with a strong understanding of infrastructure, authentication, and observability Delivered on high-stakes consulting engagements across multiple language paradigms, stacks, ecosystems, and client industries Built high-quality, maintainable software collaboratively, incrementally, and through … legacy systems with short and long-term business needs Led and delivered solutions to architecture-level problems including scalability, security, reliability, performance, maintainability, and observability Facilitated alignment across technical and non-technical stakeholders to move initiatives forward through ambiguity and complexity Provided mentorship and team support at scale, while sharing ...

Kubernetes Platform Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
pipelines (ArgoCD, FluxCD) for safe, auditable changes. Drive Infrastructure as Code practices with Terraform and Helm for reliable and repeatable builds. Heavily embed observability using Prometheus, Grafana, and OpenTelemetry to make systems measurable and reliable. Stay ahead of Kubernetes evolution by testing and adopting new versions and features early. Collaborate … including namespace isolation, network policies, and Open Policy Agent. Proven ability to troubleshoot complex performance and reliability issues across infrastructure and workloads. Experience with observability tools such as Prometheus, Grafana and OpenTelemetry to monitor cluster metrics and health. Great communication skills, with experience collaborating with internal platform users to gather ...

Engineering & Delivery Lead - Risk & Compliance

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
enabled engineering practices across design, development, testing, security, documentation and operations with governance and traceability Strengthen production stability and service resiliency by improving observability, operability, technical debt management and controlled release practices Establish a strong technology control environment aligned to our client standards and regulatory expectations and address gaps … systems including low-latency architectures and synchronous and asynchronous execution models Use strong knowledge of containers, Kubernetes, cloud, service mesh, event streaming and production observability to improve resilience and operability Show experience automating delivery and operations including CI/CD, deployment automation, performance testing, rollback and controlled production releases Apply ...

Senior Frontend Developer

Hiring Organisation
Rullion Managed Services
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£650 - £700/day
secure, performant, maintainable, and well-tested. Lead incident management and troubleshooting activities, implementing preventative measures to improve platform stability. Design systems with strong observability, scalability, and operational resilience. Create and maintain technical documentation, runbooks, and operational processes. Mentor and support engineers through coaching, code reviews, and technical guidance. Required Skills … testing frameworks. Knowledge of accessibility standards, including WCAG 2.1 AA . Experience implementing frontend security best practices Strong debugging and monitoring experience using observability tools. Experience designing and maintaining CI/CD pipelines and deployment processes. Understanding of build optimisation and module bundling techniques. Experience Minimum 5 years' commercial experience ...

Software Engineer, AI Libraries

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
delivery. Strong software architecture and system design skills. Experience building tools, platforms, or libraries for internal or external users. Strong understanding of testing, observability, maintainability, and engineering best practices. Experience working with cloud environments, ideally Azure. Experience with concurrent, parallel, or distributed computing. Familiarity with ML frameworks such as PyTorch … solutions. Desirable skills Experience working with large GPU clusters or distributed training environments. Familiarity with distributed training techniques such as DDP or FSDP.Experience with observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry. Experience with data pipeline orchestration tools such as Airflow, Flyte, Ray, Metaflow, or Argo Workflows. Experience with ...

Software Engineer London, United Kingdom

Location
Greater London, England, United Kingdom
concept through to delivery. Strong software architecture and system design skills.Experience building tools, platforms, or libraries for internal or external users.Strong understanding of testing, observability, maintainability, and engineering best practices. Experience working with cloud environments, ideally Azure. Experience with concurrent, parallel, or distributed computing. Familiarity with ML frameworks such … Desirable skills Experience working with large GPU clusters or distributed training environments. Familiarity with distributed training techniques such as DDP or FSDP. Experience with observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry. Experience with data pipeline orchestration tools such as Airflow, Flyte, Ray, Metaflow, or Argo Workflows. Experience with ...

Java Lead Software Engineer — Digital Markets Execution Technology, Execute

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
debugs code written by othersLeads technical analysis, estimation, planning, code reviews, architecture sessions, and retrospectives to drive delivery outcomesEstablishes reliability goals and implements observability, resilience patterns, and operational readiness practicesLeads incident response and post-incident reviews to improve production stability and performance; identifies recurring issues and drives automation/remediationUpholds … globally distributed teamsPreferred qualifications, capabilities, and skillsExposure to messaging systems and market protocols (e.g., MQ/Kafka; familiarity with FIX and Solace)Experience with observability stacks and resilience engineering for low-latency/latency-sensitive platformsFamiliarity with PythonExperience operating services in regulated environments with strong auditability and controlsJ.P. Morgan ...

Mid-Level DevOps Engineer

Location
Greater London, England, United Kingdom
willingness to learn the other – is fine. Solid Linux fundamentals, Bash scripting, MySQL/RDS basics (slow queries, EXPLAIN), and comfort reading observability dashboards and logs. Able to own routine tickets end-to-end, and to recognise when something needs a senior. Note: Magento or PHP application hosting experience … working in a PCI DSS environment (or similar compliance regime). Experience with Datadog, New Relic or Sumo Logic. They are our on-call observability stack. We will teach them otherwise. Familiarity with GitHub, Github Actions and Jira What we offer Out-of-hours on-call, paid per hour rota ...

Senior Platform Engineer (12 Month FTC)

Location
Greater London, England, United Kingdom
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Senior Platform Engineer (12 Month FTC)

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Advanced Solutions Architect

Location
Greater London, England, United Kingdom
proof-of-concept development when needed to de-risk architectural decisions Engineering Quality & Standards Define and champion non-functional standards: performance baselines, resilience patterns, observability requirements, and security controls Lead or contribute to performance and load testing design, interpreting results in the context of SLAs and platform growth targets Establish … delivery leads, product managers, and executive stakeholders Drive alignment between iGaming and wider L&W engineering teams on shared concerns: API strategy, data platforms, observability, and shared services Participate in hiring and technical assessment processes for engineering and architecture candidates Qualifications Degree in Computer Science, Software Engineering, or a related ...

Advanced Solutions Architect

Hiring Organisation
Light & Wonder
Location
London, UK
Employment Type
Full-time
prototyping and proof-of-concept development when needed to de-risk architectural decisionsEngineering Quality & StandardsDefine and champion non-functional standards: performance baselines, resilience patterns, observability requirements, and security controlsLead or contribute to performance and load testing design, interpreting results in the context of SLAs and platform growth targetsEstablish and evolve … delivery leads, product managers, and executive stakeholdersDrive alignment between iGaming and wider L&W engineering teams on shared concerns: API strategy, data platforms, observability, and shared servicesParticipate in hiring and technical assessment processes for engineering and architecture candidatesQualificationsDegree in Computer Science, Software Engineering, or a related technical discipline, or demonstrably ...

Software Engineering Manager

Location
Greater London, England, United Kingdom
appropriate technical documentation. Champion strong engineering practices, including collaborative programming, testing approaches such as TDD, CI/CD and production ownership. Ensure strong observability and operational health, with useful code‐quality and Production metrics, appropriate KPIs, well‐calibrated alarms and clear ownership of actions following incidents and PIRs. Build relationships … Data/Data Science, CRM, Security, Legal, Privacy and other engineering teams. Experience establishing strong operational ownership of production software, including CI/CD, observability, support practices and continuous improvement. Commitment to inclusive leadership, frequent feedback and creating an environment where engineers can grow, challenge ideas and take meaningful ownership. ...

Applied AI Engineering Lead - VP, Markets Operations

Location
Greater London, England, United Kingdom
information Build and enhance robust AI services and infrastructure using modern engineering practices, including APIs, event‐driven patterns, CI/CD, Infrastructure‐as‐Code, observability, and automated testing Partner with AI researchers, data scientists, and software engineers to translate emerging AI capabilities into practical, reliable, and compliant enterprise applications Establish … data engineering concepts, ETL and data pipelines, structured and unstructured data, and integration with enterprise data platforms Experience with CI/CD, automated testing, observability, production monitoring, and operational readiness practices Familiarity with Infrastructure‐as‐Code solutions such as Terraform and cloud or container‐based deployment patterns Working knowledge ...

SFRC Platform - Senior Delivery Lead

Hiring Organisation
Bank of America
Location
Bromley, Greater London, UK
Employment Type
Full-time
business analysis and production support groups, ensuring platform improvements translate into stronger production stability, better compute utilisation, more reliable lower-lane execution, improved batch observability, data quality controls, disciplined regression and release processes, and sustained adoption of agreed engineering changes. Role Description: This is a critical technical platform leadership role … calculation cost, guiding engineers towards software-led optimisation rather than assuming additional infrastructure capacity is available. Drive improvements in batch performance, resilience and observability by analysing execution patterns, bottlenecks, dependencies and failure modes, and converting findings into practical engineering actions. Lead execution from problem identification through to implementation, rollout ...

Senior Engineer

Location
Greater London, England, United Kingdom
turn ideas into working product quickly Contribute to product decisions and pragmatic technical trade-offs Diagnose and fix bugs quickly Improve testing, monitoring, and observability Maintain data integrity, system stability, and security Work closely with technical leadership Collaborate with our senior technical advisor on architecture and technical direction Implement technical … processes Design and maintain data-compliant systems with privacy, security, and regulatory standards embedded by default (e.g. UK GDPR, NHS DSPT, MHRA guidance) Implement observability and metrics to understand product performance, user behaviour, and system health Ensure our systems and features are audit-ready and aligned with healthcare regulatory requirements ...

Advanced Solutions Architect

Location
Chiswick, England, United Kingdom
concept development when needed to de-risk architectural decisions## ## **Engineering Quality & Standards*** Define and champion non-functional standards: performance baselines, resilience patterns, observability requirements, and security controls* Lead or contribute to performance and load testing design, interpreting results in the context of SLAs and platform growth targets* Establish … delivery leads, product managers, and executive stakeholders* Drive alignment between iGaming and wider L&W engineering teams on shared concerns: API strategy, data platforms, observability, and shared services* Participate in hiring and technical assessment processes for engineering and architecture candidates**Qualifications**Degree in Computer Science, Software Engineering, or a related ...

Software Development Engineer in Test

Location
Greater London, England, United Kingdom
quality gates, and make production failures easier to identify and diagnose. You will work hands-on across C#, TypeScript, APIs, data, pipelines, and observability rather than focusing solely on UI testing. Your work will help Athos reduce manual checking, catch failures earlier, and create stronger evidence behind every release. … quality gates into Azure DevOps CI/CD pipelines. Investigate failures across application logs, APIs, databases, telemetry, test environments, and production systems. Improve observability and production quality through monitoring, alerting, and better visibility into feed and syndication failures. Contribute to test strategy, including determining what should be tested, at which ...

Senior Product Engineer

Hiring Organisation
Zapp
Location
London, UK
Employment Type
Full-time
home in NestJS (or a similar Node.js framework) and React. Cloud and DevOps minded. Comfortable across GCP or a similar cloud, CI/CD, observability, and modern infrastructure tooling. Use AI daily. AI assistants are part of how you build. You bring back patterns that help others get more … code. You measure your work by what changed for the customer, not the lines of code shipped. Production minded. Solid grasp of API design, observability, scaling, reliability, and security. Care about craft. Clean APIs, attention to detail, and how the system feels to work in. Comfortable in uncertainty. You move ...

Applied AI Engineering Lead - VP, Markets Operations

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
contextual informationBuild and enhance robust AI services and infrastructure using modern engineering practices, including APIs, event-driven patterns, CI/CD, Infrastructure-as-Code, observability, and automated testingPartner with AI researchers, data scientists, and software engineers to translate emerging AI capabilities into practical, reliable, and compliant enterprise applicationsEstablish evaluation, monitoring … with data engineering concepts, ETL and data pipelines, structured and unstructured data, and integration with enterprise data platformsExperience with CI/CD, automated testing, observability, production monitoring, and operational readiness practicesFamiliarity with Infrastructure-as-Code solutions such as Terraform and cloud or container-based deployment patternsWorking knowledge of database design ...

Principal Backend Engineer

Hiring Organisation
Teya Solutions
Location
London, UK
Employment Type
Full-time
enforce standards for API-first development (REST/gRPC) and domain-driven design (DDD)Set the bar for availability, performance, fault tolerance, and observability across critical systemsPartner deeply with Product, Data, and leadership to translate business priorities into scalable technical strategyLead cross-organisational initiatives that are essential to company growth … environmentsStrong grounding in system design, data structures, and performance trade-offsHands-on experience with AWS, Docker, and Kubernetes in production environmentsProven ability to implement observability, resilience patterns, and CI/CD best practicesExcellent communication skills with the ability to influence and align stakeholdersA track record of mentoring senior engineers ...

Platform Technical Lead (Cloud Infrastructure & AI Enablement)

Location
Greater London, England, United Kingdom
engineering technologies including coding assistants, agent frameworks, evaluation platforms and engineering automation tooling. Own engineering platform capabilities including CI/CD, deployment automation, observability, secrets management, platform security and self-service engineering tooling. Drive improvements in developer productivity, engineering effectiveness and delivery throughput through platform automation and AI enablement capabilities. … engineering tooling and AI-native software delivery approaches. Good understanding of LLM platforms, agents, evaluation frameworks and AI governance practices. Experience with observability, DevOps, CI/CD and engineering productivity platforms. Experience managing multi-region, multi-tenant and large-scale cloud platforms. #IBIZAplatform Vodafone is committed to attracting, developing ...