1,751 to 1,775 of 2,250 Observability Jobs in London

Senior Associate, Full-Stack Engineer Opportunities

Hiring Organisation
The Bank of New York Mellon
Location
London, UK
Employment Type
Full-time
ways: Design, build, and maintain backend services, batches and APIs, contributing to UI components as needed. Own end-to-end delivery: implementation, testing, deployment, observability, and reliability. Write clean, well-tested code; participate in code reviews and continuous improvement. Collaborate with product, design, and operations to translate business needs into … programming concepts and microservicesProficiency in Java with Spring. Experience with CI/CD, automated testing (JUnit/Spock), and containers (Docker).Familiarity with microservices, observability/telemetry (e.g., Splunk, AppDynamics), and cloud deployments. Curiosity to understand the business domain and translate product strategy into technical solutions. Working knowledge of Groovy ...

Embedded Software Engineer

Location
Greater London, England, United Kingdom
best practices across the device lifecycle Edge and cloud integration: integrate devices with cloud IoT platforms and backend services; improve telemetry, health monitoring and observability; support reliable field operation and debug fleet issues Compliance and testing: support testing and validation for safety, EMC, radio and related requirements; prepare test plans …/CD workflows for embedded software Practical debugging experience using lab and software tools such as logic analysers, protocol analysers, network sniffers or observability platforms Strong problem-solving skills and the ability to work across hardware and software boundaries Bonus: low-power wireless (Thread, Zigbee, BLE); cloud IoT services ...

Infrastructure Engineer

Location
Greater London, England, United Kingdom
tenant environments for enterprise customers. Developer Velocity: Architect CI/CD pipelines (GitHub Actions, Docker) that allow our team to ship safely and quickly. Observability & Reliability: Instrument the stack with OpenTelemetry and Datadog to ensure we detect issues before our users do. AI Performance Tuning: Tune autoscaling and network routing … Haves: Experience with ECS, container orchestration, and distributed task queues (Celery/SQS). Strong Python skills. Familiarity with the Datadog/Grafana observability stack. A background in offensive security, CTFs, or AI/ML infrastructure. What We Offer Competitive Salary + Significant Equity: We want you to have true ...

Software Engineer / AI Engineer

Location
Greater London, England, United Kingdom
within a defined problem, building and testing tool use, retrieval pipelines and agent workflows, integrating AI capabilities into enterprise systems, and contributing to evaluation, observability and guardrails. You will hold a high bar on code quality, flag risks and blockers early, and work alongside host‐function stakeholders to make sure … agentic AI solutions to production standard within a defined technical approach. Implement and test tool use, retrieval pipelines, and agent workflows. Contribute to evaluation, observability and guardrails for agentic systems. Integrate AI capabilities into existing enterprise workflows and systems. Maintain high code quality and documentation so patterns can be reused. ...

Senior Software Engineer - London

Location
Greater London, England, United Kingdom
shorter path each time. Setting the standard for how a benchmark enters the framework, and building the checks that enforce it. Scheduling and observability, so cluster capacity isn't left idle while evaluation jobs queue. Reproducible results across trials, so a release decision rests on numbers that hold. Whatever stack … benchmarks, and started fixing what slows the framework down. By 6 months one part of it is yours, for example scaling the runs, observability, or a group of related benchmarks, and a release will have gone out on your numbers. By 12 months you'll know the design and trade ...

Junior DevOps Engineer

Location
Greater London, England, United Kingdom
/CD processes, automate manual processes to create a self-service environment for our developers, and maintain platform uptime SLAs by improving our observability stack. We use infrastructure-as-code to maintain our platform on AWS, so familiarity with common AWS Services (RDS, S3, ECS, EC2 etc.) and Terraform …/CD pipelines through Jenkins/Github Actions and other technology Support the Development and AI Engineering Teams - help troubleshoot their issues Bring observability through dashboards, alerting and log aggregation Run incident analysis and post mortems Environment management: ephe...dev/staging/prod What you’ll bring A strong foundation ...

Staff Python Engineer (ML)

Location
City Of London, England, United Kingdom
apps in a service architecture. Furthering Developer Experience (DevEx) by mentoring others in writing code that is intuitive, clear, and easy to test Developing observability for new and existing ML applications and GenAI/LLM integrations , making use of the Grafana Stack (Prometheus, Loki, Tempo) Develop integrations and services that … Backend-Engineering Experience owning projects from start to finish, including speccing, architecture, development, testing, deployment, release and monitoring Strong skills in building maintainable tests, observability and tracing systems. Knowledge of best practices for performance optimisation, memory management. Familiarity with Kubernetes , Docker and other cloud infrastructure, ops and containerised tools. Strong ...

Senior AI Engineer - Agentic AI

Location
Greater London, England, United Kingdom
that allow complex AI workflows to operate securely and efficiently at scale. You will be responsible for developing advanced orchestration capabilities, implementing evaluation and observability tooling and embedding enterprise controls for compliance and safety. If you are passionate about innovating with AI in real-world applications and scaling intelligent systems … ensuring graceful degradation and retries Apply enterprise security and governance practices including RBAC, prompt safety checks, traceability and secrets management Implement evaluation pipelines and observability frameworks using tools such as Langfuse, Arize or OpenTelemetry Contribute to architectural design decisions, code reviews and engineering standards for platform development Requirements Bachelor ...

Senior Network Engineer, Studios

Location
Greater London, England, United Kingdom
modelled in a NetBox source of truth, configuration is deployed from code through pipelines rather than hand‐edited, and streaming telemetry feeds the observability stack. Changes are expected to land in days, not months, in a facility where the network carries live content around the clock. The estate spans dedicated … defect. Config as code: build and run the automation that deploys configuration from intent through CI pipelines (Python, Ansible, Git); eliminate hand‐edits. Observability: develop monitoring built on streaming telemetry (gNMI/gRPC), flow analysis and modern tooling (Grafana, Zabbix class), serving the 24/7 operations teams as your ...

Member of Technical Staff (Platform Leaning)

Location
Greater London, England, United Kingdom
light up rather than back away: complex networking scenarios such as site-to-site VPN and private connectivity, multi-cloud and multi-region deployments, observability, and the machinery that makes Tessl straightforward to deploy, operate and support wherever it needs to run. We expect you to be relentlessly AI native … first year, you're shipping product features end to end like any Tessl engineer. When infrastructure work arises, from deployment architecture to networking to observability, you're the one who takes it on and lands it well, with support where you need it. Our deployment and operational practices are stronger ...

Jobshare - Sr Lead Software Engineer - Site Reliability Engineer, Python & Infrastructure management - Part time/Jobshare

Location
Greater London, England, United Kingdom
impact). Champions AI adoption and deliver AI-enabled capabilities to reduce operational toil and improve speed/quality of response. Sets direction for observability across logs/metrics/traces, including instrumentation standards, golden signals, and end-user journey monitoring. Improves alert quality and routing: reduce false positives, improve … design, build, and deliver automation and reliability solutions. Fluency & expertise in Python Deep practical knowledge of: SLOs/SLIs, error budgets, incident management, postmortems, observability design across metrics/logs/traces and distributed systems troubleshooting, resilience engineering, performance/capacity management, and change risk reduction. Proficiency and experience with ...

AI Solutions Architect

Location
Greater London, England, United Kingdom
rather than conceptual designs alone. Support early industrialisation of AI solutions, contributing to the adoption of MLOps, LLMOps and AIOps practices such as evaluation, observability, versioning and responsible deployment. Ensure AI solutions meet relevant standards and regulations for security, privacy and responsible AI, particularly in public sector and regulated environments. … preferred, and an understanding of PaaS, SaaS, real-time and batch processing patterns. Familiarity with Docker, CI/CD, infrastructure approaches, MLOps, LLMOps, AIOps, observability, versioning and production operations would be advantageous. Strong understanding of responsible AI, security, privacy, governance, data lifecycle and data quality considerations, including risks such ...

Senior AI Engineer

Hiring Organisation
MarkIT Placements
Location
London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
From £700 to £900 per day
Deploy AI systems across cloud, on-premises, sovereign and fully offline environments. Engineer for reliability, latency, security and predictable failure behaviour. Build evaluation and observability capabilities to understand model performance, agent behaviour and failure modes. Take end-to-end ownership of technical workstreams, from architecture through deployment. Establish engineering patterns … would be advantageous: Multimodal AI and reasoning. Edge or offline AI deployments. Kubernetes, particularly EKS or OpenShift. MLOps, including model evaluation, monitoring and reproducibility. Observability for agentic AI systems, including model performance, agent behaviour and drift. Agent orchestration and inter-agent communication protocols such as A2A. Model Context Protocol ...

Principal Product Engineer

Location
Greater London, England, United Kingdom
home in NestJS (or a similar Node.js framework) and React. Cloud and DevOps minded. Comfortable across GCP or a similar cloud, CI/CD, observability, and modern infrastructure tooling. Use AI daily. AI assistants are part of how you build. You bring back patterns that help others get more … code. You measure your work by what changed for the customer, not the lines of code shipped. Production minded. Solid grasp of API design, observability, scaling, reliability, and security. Care about craft. Clean APIs, attention to detail, and how the system feels to work in. Comfortable in uncertainty. You move ...

Senior Solution Architect, Data & AI

Location
City Of London, England, United Kingdom
Data Ingestion and Integration Storage and management Processing and Transformation Design of the Semantic Layer Data Sharing Delivery of AI-ready data Governance, security, observability and cost are designed in at every stage. They work closely with Solution Sales Specialists, Partners and their fellow architects in associated domain specialistations. Physical … governance that AI and machine learning need. Work with the Principal AI Architect on AI use cases that depend on the data foundation. Security, Observability & FinOps by Design - Build security, privacy, monitoring, reliability and cost management into every layer of the architecture. Give customers a clear view of consumption-based ...

Principal Software Engineer - Squad Lead Engineer

Location
City Of London, England, United Kingdom
complete complex bug fixes and performance improvements Define and uphold Definition of Ready/Done including code quality, automated test coverage, security checks, and observability Establish/maintain CI/CD pipelines, quality gates, and sensible branching/release strategies Drive a pragmatic quality strategy: test pyramid balance, contract tests … Windows Experience with relational and non‐relational data stores, performance tuning, and data modelling Knowledge of CI/CD platforms, containers, cloud technologies, observability, and monitoring practices Understanding of secure coding, performance optimisation, reliability engineering, and incident response Work in a Way That Works for You We promote a healthy ...

Principal Product Engineer

Hiring Organisation
Zapp
Location
London, UK
Employment Type
Full-time
home in NestJS (or a similar Node.js framework) and React. Cloud and DevOps minded. Comfortable across GCP or a similar cloud, CI/CD, observability, and modern infrastructure tooling. Use AI daily. AI assistants are part of how you build. You bring back patterns that help others get more … code. You measure your work by what changed for the customer, not the lines of code shipped. Production minded. Solid grasp of API design, observability, scaling, reliability, and security. Care about craft. Clean APIs, attention to detail, and how the system feels to work in. Comfortable in uncertainty. You move ...

Staff Data Platform Engineer London, United Kingdom

Location
Greater London, England, United Kingdom
/day scale Contribute to the technical design of the Data Transfer Hub, making pragmatic trade-offs across UX, careliability, throughput, cost, partner constraints, observability and operational support Build parallelised, distributed data transfer pipelines using Flyte for workflow orchestration, Kafka/event-driven patterns for lifecycle tracing, and Azure … Background in autonomous vehicles, robotics, geospatial data, video processing or other data-intensive domains Experience with Kafka or other event-driven systems for transfer observability and lifecycle tracking Experience building dashboards to communicate transfer job status Multi-cloud experience across Azure, AWS and/or GCP Experience improving cost efficiency ...

Data Platform Engineer- BPL-CIO

Location
Greater London, England, United Kingdom
automation, GitOps operating models AWS Platform Engineering Hands‐on AWS engineering with a focus on: IAM and security patterns, Networking integration, Storage and encryption, Observability, Resilience and operational readiness, Automation and supportability Databricks on AWS Experience with: Workspace deployment, Databricks Terraform provider, Identity integration, Storage integration, Network connectivity patterns Platform … Engineering/Internal Developer Platforms Experience building: Self‐service capabilities, Golden paths, Paved roads, Reusable engineering patterns, Internal platform products Observability, Reliability & Operational Engineering Experience with: Monitoring and alerting, Operational telemetry, Incident management, Resilience testing, Recovery automation, Production support You may be assessed on the key critical skills relevant ...

Senior Software Engineer, Data

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
/day scaleContribute to the technical design of the Data Transfer Hub, making pragmatic trade-offs across UX, careliability, throughput, cost, partner constraints, observability and operational supportBuild parallelised, distributed data transfer pipelines using Flyte for workflow orchestration, Kafka/event-driven patterns for lifecycle tracing, and Azure for storage … technical integrationsBackground in autonomous vehicles, robotics, geospatial data, video processing or other data-intensive domainsExperience with Kafka or other event-driven systems for transfer observability and lifecycle trackingExperience building dashboards to communicate transfer job statusMulti-cloud experience across Azure, AWS and/or GCPExperience improving cost efficiency, throughput or reliability ...

Staff Network Engineer

Location
Greater London, England, United Kingdom
root-cause analysis for performance and stability issues, and systematically reducing reactive toil through runbooks, automation, and measurable SLOs. Set the direction for network observability, telemetry, monitoring, and alerting to provide clear visibility into fabric health, performance, and traffic patterns. Ensure the accuracy and reliability of source-of-truth network … Juniper SRX and/or Palo Alto, including security policy architecture, high-availability design, and multi‐tenant segmentation. Experience designing network telemetry and observability for high-throughput, performance-sensitive environments. Proven ability to lead complex technical decisions and incidents across networking, systems, storage, and HPC/AI workload teams, with ...

Software Engineer - Software Delivery

Hiring Organisation
Neo4J
Location
London, UK
Employment Type
Full-time
internal engineers. This involves things likeWorking closely with internal engineers to identify pain pointsMaking sure the product experience is as good as possibleSetting up observability around how the platform is performing but also how users are interacting with the platformExperience creating abstractions to simplify the local developer workflow via tooling … related build tooling like kustomize or helm is also meriting. Experience integrating software with Google Cloud Platform, AWS and AzureSome experience in common software observability practices such as tracing, logging and metrics exporting.#LI-HybridWhy Join Neo4j Neo4j is, without question, the most popular graph intelligence platform in the world. ...

Senior Data Engineer

Location
Greater London, England, United Kingdom
data retrieval layers (pgVector, Pinecone) and ensure efficient embedding pipelines for AI contexts. Partner with AI teams to monitor data latency, cost efficiency, and observability metrics. Collaboration & Governance Partner with Platform Operations and Security to enforce privacy, compliance, and access control frameworks (GDPR, SOC2). Work cross‐functionally with analysts … platform data (Google Ads, Meta, TikTok, DV360, Amazon Ads). Demonstrated expertise in code versioning (GitHub), CI/CD integration, and data observability practices. Ability to write clean, modular, testable code and review peers’ contributions for maintainability and performance. Additional Information Publicis Groupe has fantastic benefits on offer ...

Senior II Product Engineer

Hiring Organisation
9fin
Location
London, UK
Employment Type
Full-time
Engineering, and our editorial and legal domain experts to scope work and ship the right thing. Improve developer experience by investing in tooling, testing, observability, and the paved road so the whole team moves faster. Ramp on legacy areas of the system, find the highest leverage cleanup, and execute … across a product domain. Experience contributing to the design of distributed systems in production, including the operational realities such as failure modes, observability, data consistency, and graceful degradation. A track record of solving scaling problems, whether database scaling, throughput, latency, or cost. You can talk through a real example ...

Agentic Commerce Architect

Hiring Organisation
Accenture
Location
London, UK
Employment Type
Full-time
intelligence — including embedding pipelines, vector retrieval, semantic search, and structured product data — ensuring agent outputs are accurate, grounded, and commercially reliable Embedding AI governance, observability, and responsible AI principles into architecture design — including audit logging, human-in-the-loop escalation points, and performance monitoring — as first-class concerns rather than … deployment, including familiarity with cloud-native services and CI/CD-based delivery practices Understanding of AI governance and responsible AI principles — including observability, tracing, auditability, and how to build guardrails into production AI systems Preferred Experience Experience contributing to technical proposals, reference architectures, or delivery accelerators in a consulting ...