1,501 to 1,525 of 2,378 Observability Jobs in London

Regional Vice President

Location
Greater London, England, United Kingdom
advantage of all structured and unstructured data — securing and protecting private information more effectively — Elastic’s complete, cloud-based solutions for search, security, and observability help organizations deliver on the promise of AI. What Is The Role Elastic, the Search AI company, is looking for a high-energy Regional Vice ...

Fullstack Engineer

Location
City Of London, England, United Kingdom
things You take ownership of your work and treat the company’s success as your own Nice-to-Haves Experience with analytics, monitoring or observability tooling Background in automotive, marketplaces, or high-SKU B2B platforms Exposure to AI/ML-powered products or startups Prior experience leading UI best practices ...

Director, Engineering - Partnerships

Location
Greater London, England, United Kingdom
definition, and performance optimization for high-scale enterprise platforms. Experience with API lifecycle management, OAuth/SSO, developer portals, certification, SDKs, event-driven systems, observability, or partner data access. Experience operating across globally distributed teams and time zones. Benefits and Perks Work from (almost) anywhere for up to 20 days ...

Payments Architect

Hiring Organisation
Endava
Location
London, UK
Employment Type
Full-time
scalable, resilient payment architectures capable of supporting high transaction volumes and demanding availability requirementsDefine and govern non-functional requirements including latency, throughput, availability, recoverability, observability and securityShape API strategies, integration models and event-driven or distributed system designs that support extensibility and regulatory complianceArchitect and implement Single Sign ...

ML Research Engineer - Member of Technical Staff

Location
Greater London, England, United Kingdom
context and memory management; agent steering, task decomposition and specialisation; continual learning during deployment; and inter‐agent communication Build agent analysis and evaluation infrastructure - observability, behaviour analysis, evaluation harnesses, auditing - as durable instruments the whole company works from, not one‐off scripts Publish. Evaluation results, methods and discoveries, as both ...

Technical Program Manager, Evaluation & Validation

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
tracking activity. DesirableExperience in autonomous vehicles, robotics, or another safety-critical AI/ML program. Familiarity with simulation, synthetic data, evaluation harnesses, or ML observability tooling at production scale. Working knowledge of SOTIF (ISO 21448), ISO 26262, or other safety standards relevant to validating learned systems. An engineering or computer ...

Research Engineer, Evals - Member of Technical Staff

Location
Greater London, England, United Kingdom
defining what each model is genuinely good at, what benchmarks really measure, and what capability profile a real compound task demands Build evaluation and observability infrastructure as durable instruments the whole company works from What You'll Bring Evidence that you can run research of your own: you take ...

Customer Support Engineering Lead - London

Location
Greater London, England, United Kingdom
share pain points for our customers Investigate and resolve hard technical issues Root-cause customer issues across the stack using logs, traces, and observability (e.g. Datadog), session data, and the codebase. Ship pragmatic data fixes and tightly-scoped code fixes where appropriate, and partner with Engineering to land deeper fixes. ...

Engineering Manager, Search

Location
Greater London, England, United Kingdom
your day-to-day work. Bonus points if: You are familiar with search engine technology such as OpenSearch, ElasticSearch, Vespa.. You are familiar with observability, tracking and data pipeline tools and methodologies. Additional Information Health + Mental Wellbeing PMI and cash plan healthcare access with Bupa Subsidised counselling and coaching ...

Senior Director - EMEA Enterprise Business Development

Location
Greater London, England, United Kingdom
direction and the requirements of large-scale Enterprise customers. Translate customer and market insight into clear recommendations for product investment, developer experience, security, governance, observability, integration and Enterprise platform capabilities. Work closely with Field Engineering to define Enterprise adoption strategies using Forward Deployed Engineers, specialist technical resources and services. Build ...

Principal Engineer

Location
Greater London, England, United Kingdom
front of the funnel through to production release at the back, and make sure we know in advance what would break first. Improve observability and operability across the flow from "buy" to "delivered", reducing "where is my order" contacts and manual interventions. AI-assisted development: Set the standard ...

Senior Forward Deployed ML Engineer, Agents

Location
Greater London, England, United Kingdom
agent development, MLOps pipeline implementation, and production optimization. You understand what makes agents perform well in production and how to systematically improve quality through observability and evaluation. Experience with voice AI platforms, RAG systems, and LLM orchestration frameworks is highly desirable. You bring exceptional communication skills, customer empathy … validate datasets for fine-tuning, evaluation, and synthetic data generation Work with other MLEs, MLOps, SREs to carry out model deployment and productionization Observability, Evaluation & Production Operations Implement LLM and agents observability and monitoring tracking token usage, latency, costs, and quality metrics across deployments on aion's infrastructure Instrument applications ...

Observability Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
high-impact research - designing systems that scale, accelerate discovery and support innovation across the firm. Take the next step in your career. The roleThe Observability Engineering Team manages the doors - both entry and exit - to the telemetry backends at G-Research, ensuring our engineers can effectively produce and consume telemetry … their services. As an Observability Engineer, you'll help make observability seamless for developers and platform teams by building pipes to ingest and route data in predictable, composable ways, as well as visualising that data after the fact. You'll have deep experience across observability stacks, a clear understanding ...

devops engineer for AI platforms

Location
Greater London, England, United Kingdom
secure supply chain configurations; Establish IaC and GitOps standards with automated testing for every infrastructure change; Prototype agentic infrastructure components, including deployment and observability platforms in service meshes; Contribute to the Kong AI Gateway, including Dataplane deployments, ACM/SSL integration, and DataDog observability; Champion DevSecOps maturity through SAST/… Data, and AI teams to shape DevOps and AI platform architectures with regulatory compliance; Stay current with CNCF and AI ecosystem innovations, including eBPF observability and agent‐aware orchestration; Lead and mentor a team of DevOps engineers while remaining hands‐on. Требования: Experience leading or mentoring engineering teams, setting direction ...

Senior or Staff Software Engineer, SRE/ Platform Team

Location
Greater London, England, United Kingdom
code with Kubernetes and Terraform. You'll be at the forefront of shaping our foundational architecture, ensuring it’s both resilient and scalable. Drive Observability and Monitoring: Establish and maintain a state‐of‐the‐art observability and monitoring stack. Your insights will enable us to stay ahead of potential issues ...

Expert Forward Deployment Engineer

Hiring Organisation
Luxoft
Location
London, UK
Employment Type
Full-time
analysis, upgrading Java and NPM runtimes, modernizing Spring and legacy middleware applications, improving CI/CD pipelines, containerizing applications, automating deployments, and introducing standard observability and resilience patterns. The Expert FDE is expected to lead complex engagements, work directly with development and client stakeholders, define the technical remediation approach, implement … testing, release, resilience, and legacy technology challenges with development teams. Assess application code, dependencies, runtime environment, test coverage, deployment architecture, CI/CD pipelines, observability, and operational risks. Write, debug, review, and enhance production-quality code and configuration throughout engagements. Define and implement practical modernization and remediation plans with clear ...

Infrastructure Software Engineer, Apps Platform

Location
Greater London, England, United Kingdom
Apps Platform, London As an Infrastructure Engineer on the Apps Platform Infrastructure team, you'll help build and evolve our platform’s deployment and observability layers across multiple cloud providers and on-premises, for both internal and customer-managed environments. This is a role for someone who cares about building … both on-premise, all major cloud platforms and beyond Ensure fast, secure, and reproducible deployments across all supported CSPs and on-prem Expand the observability of the platform so that forward deployed infrastructure teams can easily and efficiently operate customer deployments Partner closely with internal teams deploying the platform ...

Lead Cloud Engineer

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
architecture, and business partners. What You'll DoLead the design, build, and optimisation of AWS-based cloud platforms, including EKS, Kubernetes, networking, identity, and observability capabilities supporting enterprise-scale workloadsDrive cloud migration programmes end-to-end, including application assessment, readiness analysis, refactoring strategies, and transition planningShape and deliver platform roadmaps … point for complex cloud and Kubernetes challenges, making informed technical and operational decisionsWhat You BringDeep expertise in AWS services, including EKS, IAM, networking, compute, observability, and securityStrong hands-on experience with Kubernetes and container platforms, including cluster operations and workload optimisationProven experience with Infrastructure-as-Code, CI/CD pipelines ...

Principal Architect — Distributed Systems & Event Streaming (Director I - Product Architect)

Hiring Organisation
UST Global
Location
London, UK
Employment Type
Full-time
architecture and evolution of our event platform, including Kafka/Redpanda, event contracts, schema management, event sourcing, CQRS, distributed consistency, high-throughput processing, resilience, observability, and operational excellence. This is not an advisory or documentation-focused architecture role. We are looking for a practitioner who stays close to implementation … resilient systems using retries, circuit breakers, bulkheads, dead-letter queues, replay, rate limiting, and graceful degradation patterns. Strong debugging skills across code, messaging, infrastructure, observability, and production operations. Excellent communication and mentoring capability across globally distributed teams. PreferredRedpanda architecture and operations. Event sourcing and CQRS in production environments. Stream processing ...

DevOps Engineer

Hiring Organisation
Sanderson Recruitment
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
Up to £500 per day + Inside IR-35
Infrastructure as Code using Terraform Design, implement, and support CI/CD pipelines Manage containerised applications using Kubernetes and Docker Implement monitoring, logging, and observability solutions Automate operational tasks and infrastructure processes through scripting Manage environments across development, testing, and production Support cloud security initiatives and best practices Coordinate … Infrastructure as Code Proven experience designing and maintaining CI/CD pipelines Strong knowledge of Kubernetes, Docker, and container technologies Experience implementing monitoring and observability solutions Automation and scripting skills (e.g. Python, Bash, PowerShell) Good understanding of cloud security principles and controls Experience with environment management and release processes Strong ...

Senior ML Engineer

Hiring Organisation
Harnham - Data & Analytics Recruitment
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£75,000 - £85,000 per annum
maintain cloud-native infrastructure using Kubernetes and Infrastructure as Code technologies such as Terraform or Bicep. Develop robust CI/CD processes and observability frameworks to ensure reliable and secure ML operations. Collaborate closely with Data Scientists, software engineers, and client-facing teams to deliver scalable AI solutions. Influence technical … containerised workloads. Expertise in Infrastructure as Code using Terraform, Bicep, Pulumi, or comparable technologies. Experience building CI/CD pipelines and implementing monitoring and observability practices. Experience working with cloud platforms, ideally Azure, although other cloud backgrounds will be considered. Exposure to LLMs, NLP, text analytics, or generative AI applications. ...

Intelligent Automation Engineering Manager

Location
Greater London, England, United Kingdom
deployment and operational support of AI agents and AI-powered solutions. Establish engineering standards and best practices for AI architecture, orchestration, retrieval, tool invocation, observability, governance, privacy, security and cost management. Review technical designs and architecture documentation to ensure solutions align with engineering, security and governance standards. Translate emerging … with MCP (Model Context Protocol), MCP Servers, MCP Clients or enterprise AI integration frameworks. Knowledge of LLMOps, AI evaluation frameworks, model routing and AI observability tooling. Experience with enterprise integration technologies including APIs, Middleware, ESB or iPaaS platforms. Hands‐on experience integrating internal and third-party systems. Azure cloud experience ...

SC Cleared DevOps Engineer

Hiring Organisation
Sanderson Recruitment
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
Up to £570 per day + Inside IR-35
Develop and maintain Infrastructure as Code solutions. * Create and support CI/CD pipelines to enable secure, repeatable deployments. * Implement monitoring, logging, alerting and observability capabilities. * Ensure platform resilience, availability and operational readiness. * Manage configuration, secrets and environment controls. * Embed security best practice throughout the delivery lifecycle. * Support incident response … Code, ideally Terraform, CloudFormation or AWS CDK. * Strong CI/CD pipeline experience. * Experience with containerisation technologies, ideally Docker and Kubernetes. * Strong understanding of observability, monitoring and operational support. * Knowledge of cloud security, IAM, secrets management and governance controls. * Experience delivering and supporting production cloud environments within secure or regulated ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
infrastructure Troubleshoot and resolve infrastructure issues and fix root causes of issues at the source Implement monitoring, logging and alerting solutions to ensure system observability Collaborate with engineers and architects to improve platform documentation, standards and adoption Requirements Strong hands-on experience with cloud platforms (AWS & Azure ideally) Hands …/VNets, load balancers, DNS, security groups/NSGs) Experience with secrets management and identity/access control (IAM, OIDC, Azure AD) Familiarity with observability tooling (Prometheus, Grafana, CloudWatch, Azure Monitor) Risk Benefit Statement Learn more about the LexisNexis Risk team and how we work here We know your well ...

Technical Delivery Lead

Location
Greater London, England, United Kingdom
idempotency, and dead‐letter handling. Experience implementing data quality controls, including validation, reconciliation, and exception handling. Experience establishing engineering standards for coding, testing, documentation, observability, and security. Excellent written and verbal communication skills, with the ability to explain complex technical concepts to diverse audiences and align cross‐team efforts. Strong … onboarding vendor or third‐party source systems. Experience with event‐driven or streaming services such as Pub/Sub or Dataflow. Experience with observability tooling such as Cloud Monitoring, Cloud Logging, or Datadog. Experience in banking, insurance, or other regulated financial services environments. Bits In Glass (BIG) operates ...