1,551 to 1,575 of 2,452 Observability Jobs in London

Senior Field Enablement Manager

Location
Greater London, England, United Kingdom
enable Enterprise teams to succeed with customers across one of Datadog's fastest-growing and most complex regions. About Datadog: Datadog is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure ...

Head of Engineering

Location
Greater London, England, United Kingdom
infrastructure that enables our machine learning and research teams to succeed. Outcome – an elite engineering team. We've built solid foundations: continuous deployment, modern observability, high test coverage and a fear-free culture. You'll help the organisation continually improve by coaching engineers and team leads, strengthening engineering practices ...

Agent Runtime & Systems - Member of Technical Staff

Location
Greater London, England, United Kingdom
efficient as we expand our model and tool footprint. This is a high-leverage role tackling complex systems challenges across the entire stack. Runtime, observability tooling, and framework design are three separate specialties. We are glad to hire someone with real depth in a combination of them. What ...

Regional Vice President

Location
Greater London, England, United Kingdom
advantage of all structured and unstructured data - securing and protecting private information more effectively - Elastic's complete, cloud-based solutions for search, security, and observability help organizations deliver on the promise of AI. What is The Role: Elastic, the Search AI company, is looking for a NEW high-energy Regional ...

Fullstack Engineer

Location
City Of London, England, United Kingdom
things You take ownership of your work and treat the company’s success as your own Nice-to-Haves Experience with analytics, monitoring or observability tooling Background in automotive, marketplaces, or high-SKU B2B platforms Exposure to AI/ML-powered products or startups Prior experience leading UI best practices ...

Director, Engineering - Partnerships

Location
Greater London, England, United Kingdom
definition, and performance optimization for high-scale enterprise platforms. Experience with API lifecycle management, OAuth/SSO, developer portals, certification, SDKs, event-driven systems, observability, or partner data access. Experience operating across globally distributed teams and time zones. Benefits and Perks Work from (almost) anywhere for up to 20 days ...

Payments Architect

Hiring Organisation
Endava
Location
London, UK
Employment Type
Full-time
scalable, resilient payment architectures capable of supporting high transaction volumes and demanding availability requirementsDefine and govern non-functional requirements including latency, throughput, availability, recoverability, observability and securityShape API strategies, integration models and event-driven or distributed system designs that support extensibility and regulatory complianceArchitect and implement Single Sign ...

ML Research Engineer - Member of Technical Staff

Location
Greater London, England, United Kingdom
context and memory management; agent steering, task decomposition and specialisation; continual learning during deployment; and inter‐agent communication Build agent analysis and evaluation infrastructure - observability, behaviour analysis, evaluation harnesses, auditing - as durable instruments the whole company works from, not one‐off scripts Publish. Evaluation results, methods and discoveries, as both ...

Technical Program Manager, Evaluation & Validation

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
tracking activity. DesirableExperience in autonomous vehicles, robotics, or another safety-critical AI/ML program. Familiarity with simulation, synthetic data, evaluation harnesses, or ML observability tooling at production scale. Working knowledge of SOTIF (ISO 21448), ISO 26262, or other safety standards relevant to validating learned systems. An engineering or computer ...

Research Engineer, Evals - Member of Technical Staff

Location
Greater London, England, United Kingdom
defining what each model is genuinely good at, what benchmarks really measure, and what capability profile a real compound task demands Build evaluation and observability infrastructure as durable instruments the whole company works from What You'll Bring Evidence that you can run research of your own: you take ...

Customer Support Engineering Lead - London

Location
Greater London, England, United Kingdom
share pain points for our customers Investigate and resolve hard technical issues Root-cause customer issues across the stack using logs, traces, and observability (e.g. Datadog), session data, and the codebase. Ship pragmatic data fixes and tightly-scoped code fixes where appropriate, and partner with Engineering to land deeper fixes. ...

Engineering Manager, Search

Location
Greater London, England, United Kingdom
your day-to-day work. Bonus points if: You are familiar with search engine technology such as OpenSearch, ElasticSearch, Vespa.. You are familiar with observability, tracking and data pipeline tools and methodologies. Additional Information Health + Mental Wellbeing PMI and cash plan healthcare access with Bupa Subsidised counselling and coaching ...

Senior Director - EMEA Enterprise Business Development

Location
Greater London, England, United Kingdom
direction and the requirements of large-scale Enterprise customers. Translate customer and market insight into clear recommendations for product investment, developer experience, security, governance, observability, integration and Enterprise platform capabilities. Work closely with Field Engineering to define Enterprise adoption strategies using Forward Deployed Engineers, specialist technical resources and services. Build ...

Principal Engineer

Location
Greater London, England, United Kingdom
front of the funnel through to production release at the back, and make sure we know in advance what would break first. Improve observability and operability across the flow from "buy" to "delivered", reducing "where is my order" contacts and manual interventions. AI-assisted development: Set the standard ...

Principal Engineer

Hiring Organisation
MOO
Location
London, United Kingdom
Salary
£ 70 K
front of the funnel through to production release at the back, and make sure we know in advance what would break first. Improve observability and operability across the flow from "buy" to "delivered", reducing "where is my order" contacts and manual interventions.AI-assisted development: Set the standard for how engineers ...

Senior Forward Deployed ML Engineer, Agents

Location
Greater London, England, United Kingdom
agent development, MLOps pipeline implementation, and production optimization. You understand what makes agents perform well in production and how to systematically improve quality through observability and evaluation. Experience with voice AI platforms, RAG systems, and LLM orchestration frameworks is highly desirable. You bring exceptional communication skills, customer empathy … validate datasets for fine-tuning, evaluation, and synthetic data generation Work with other MLEs, MLOps, SREs to carry out model deployment and productionization Observability, Evaluation & Production Operations Implement LLM and agents observability and monitoring tracking token usage, latency, costs, and quality metrics across deployments on aion's infrastructure Instrument applications ...

Observability Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
high-impact research - designing systems that scale, accelerate discovery and support innovation across the firm. Take the next step in your career. The roleThe Observability Engineering Team manages the doors - both entry and exit - to the telemetry backends at G-Research, ensuring our engineers can effectively produce and consume telemetry … their services. As an Observability Engineer, you'll help make observability seamless for developers and platform teams by building pipes to ingest and route data in predictable, composable ways, as well as visualising that data after the fact. You'll have deep experience across observability stacks, a clear understanding ...

devops engineer for AI platforms

Location
Greater London, England, United Kingdom
secure supply chain configurations; Establish IaC and GitOps standards with automated testing for every infrastructure change; Prototype agentic infrastructure components, including deployment and observability platforms in service meshes; Contribute to the Kong AI Gateway, including Dataplane deployments, ACM/SSL integration, and DataDog observability; Champion DevSecOps maturity through SAST/… Data, and AI teams to shape DevOps and AI platform architectures with regulatory compliance; Stay current with CNCF and AI ecosystem innovations, including eBPF observability and agent‐aware orchestration; Lead and mentor a team of DevOps engineers while remaining hands‐on. Требования: Experience leading or mentoring engineering teams, setting direction ...

Senior or Staff Software Engineer, SRE/ Platform Team

Location
Greater London, England, United Kingdom
code with Kubernetes and Terraform. You'll be at the forefront of shaping our foundational architecture, ensuring it’s both resilient and scalable. Drive Observability and Monitoring: Establish and maintain a state‐of‐the‐art observability and monitoring stack. Your insights will enable us to stay ahead of potential issues ...

Expert Forward Deployment Engineer

Hiring Organisation
Luxoft
Location
London, UK
Employment Type
Full-time
analysis, upgrading Java and NPM runtimes, modernizing Spring and legacy middleware applications, improving CI/CD pipelines, containerizing applications, automating deployments, and introducing standard observability and resilience patterns. The Expert FDE is expected to lead complex engagements, work directly with development and client stakeholders, define the technical remediation approach, implement … testing, release, resilience, and legacy technology challenges with development teams. Assess application code, dependencies, runtime environment, test coverage, deployment architecture, CI/CD pipelines, observability, and operational risks. Write, debug, review, and enhance production-quality code and configuration throughout engagements. Define and implement practical modernization and remediation plans with clear ...

Infrastructure Software Engineer, Apps Platform

Location
Greater London, England, United Kingdom
Apps Platform, London As an Infrastructure Engineer on the Apps Platform Infrastructure team, you'll help build and evolve our platform’s deployment and observability layers across multiple cloud providers and on-premises, for both internal and customer-managed environments. This is a role for someone who cares about building … both on-premise, all major cloud platforms and beyond Ensure fast, secure, and reproducible deployments across all supported CSPs and on-prem Expand the observability of the platform so that forward deployed infrastructure teams can easily and efficiently operate customer deployments Partner closely with internal teams deploying the platform ...

Lead Cloud Engineer

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
architecture, and business partners. What You'll DoLead the design, build, and optimisation of AWS-based cloud platforms, including EKS, Kubernetes, networking, identity, and observability capabilities supporting enterprise-scale workloadsDrive cloud migration programmes end-to-end, including application assessment, readiness analysis, refactoring strategies, and transition planningShape and deliver platform roadmaps … point for complex cloud and Kubernetes challenges, making informed technical and operational decisionsWhat You BringDeep expertise in AWS services, including EKS, IAM, networking, compute, observability, and securityStrong hands-on experience with Kubernetes and container platforms, including cluster operations and workload optimisationProven experience with Infrastructure-as-Code, CI/CD pipelines ...

Principal Architect — Distributed Systems & Event Streaming (Director I - Product Architect)

Hiring Organisation
UST Global
Location
London, UK
Employment Type
Full-time
architecture and evolution of our event platform, including Kafka/Redpanda, event contracts, schema management, event sourcing, CQRS, distributed consistency, high-throughput processing, resilience, observability, and operational excellence. This is not an advisory or documentation-focused architecture role. We are looking for a practitioner who stays close to implementation … resilient systems using retries, circuit breakers, bulkheads, dead-letter queues, replay, rate limiting, and graceful degradation patterns. Strong debugging skills across code, messaging, infrastructure, observability, and production operations. Excellent communication and mentoring capability across globally distributed teams. PreferredRedpanda architecture and operations. Event sourcing and CQRS in production environments. Stream processing ...

DevOps Engineer

Hiring Organisation
Sanderson Recruitment
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
Up to £500 per day + Inside IR-35
Infrastructure as Code using Terraform Design, implement, and support CI/CD pipelines Manage containerised applications using Kubernetes and Docker Implement monitoring, logging, and observability solutions Automate operational tasks and infrastructure processes through scripting Manage environments across development, testing, and production Support cloud security initiatives and best practices Coordinate … Infrastructure as Code Proven experience designing and maintaining CI/CD pipelines Strong knowledge of Kubernetes, Docker, and container technologies Experience implementing monitoring and observability solutions Automation and scripting skills (e.g. Python, Bash, PowerShell) Good understanding of cloud security principles and controls Experience with environment management and release processes Strong ...

Senior ML Engineer

Hiring Organisation
Harnham - Data & Analytics Recruitment
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£75,000 - £85,000 per annum
maintain cloud-native infrastructure using Kubernetes and Infrastructure as Code technologies such as Terraform or Bicep. Develop robust CI/CD processes and observability frameworks to ensure reliable and secure ML operations. Collaborate closely with Data Scientists, software engineers, and client-facing teams to deliver scalable AI solutions. Influence technical … containerised workloads. Expertise in Infrastructure as Code using Terraform, Bicep, Pulumi, or comparable technologies. Experience building CI/CD pipelines and implementing monitoring and observability practices. Experience working with cloud platforms, ideally Azure, although other cloud backgrounds will be considered. Exposure to LLMs, NLP, text analytics, or generative AI applications. ...