1,901 to 1,925 of 2,452 Observability Jobs in London

Senior Software Engineer

Hiring Organisation
The Portfolio Group
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£90000/annum
Establish automated testing across backend and frontend applications, including unit, contract and end-to-end testing. Work with the platform engineering team on deployment, observability, logging, tracing and operational readiness. Act as the technical owner for the application and integration layer, making and documenting key architectural decisions. Provide technical guidance … such as Lambda, ECS, API Gateway, S3, CloudFront, Cognito and IAM. Experience designing and operating distributed or event-driven systems. A strong understanding of observability, testing and CI/CD practices. Experience working with data platforms or stores such as MongoDB, OpenSearch or Databricks. Experience integrating internal systems and third ...

Tivoli Netcool OMNIbus SME - 11850CF

Hiring Organisation
Proactive Appointments
Location
London, UK
Employment Type
Full-time
effective monitoring coverage. Define platform standards, monitoring policies and best practices. Maintain technical documentation, architecture diagrams and support procedures. Contribute to the monitoring and observability roadmap and identify opportunities for platform modernisation. Provide technical mentoring and knowledge sharing across engineering and operational teams. Significant hands-on experience administering and supporting … APIs, JSON, XML and integration technologies. Understanding of monitoring across AWS, Azure and Google Cloud. Experience with container and Kubernetes monitoring. Knowledge of enterprise observability frameworks and modern SRE practices. Tivoli Netcool OMNIbus SMEDue to the volume of applications received for positions, it will not be possible to respond ...

AI Platform Support Engineer (EMEA)

Location
Greater London, England, United Kingdom
combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in. We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle … post incident reviews and operational improvements Build internal tooling, automation, documentation, and runbooks Partner closely with infrastructure, networking, and platform engineering teams Help improve observability, operational visibility, and troubleshooting workflows Improve the customer experience through better processes and technical guidance What This Role Is Not This is not a traditional ...

Senior Full Stack Engineer - Lyst Shop (11 Month FTC - Maternity Cover)

Location
Greater London, England, United Kingdom
rely heavily on experimentation to validate ideas and guide decisions. Technical Excellence: You will help maintain a high bar for code quality, testing, observability, and system reliability. You’ll contribute to architectural discussions, improve developer experience, and proactively address technical debt where needed. Team Contribution: You will mentor and support … working relationships across Product, Design, QA, Analytics, and Engineering teams while actively participating in team ceremonies and technical discussions. Technical Impact: Improve the stability, observability, and maintainability of our systems through better monitoring, resilient code, and thoughtful testing practices. Growth & Ownership: Gain confidence working across our platform and infrastructure while ...

Embedded Software Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, UK
Employment Type
Full-time
security best practices across the device lifecycleEdge and cloud integration: integrate devices with cloud IoT platforms and backend services; improve telemetry, health monitoring and observability; support reliable field operation and debug fleet issuesCompliance and testing: support testing and validation for safety, EMC, radio and related requirements; prepare test plans …/CD workflows for embedded softwarePractical debugging experience using lab and software tools such as logic analysers, protocol analysers, network sniffers or observability platformsStrong problem-solving skills and the ability to work across hardware and software boundariesBonus: low-power wireless (Thread, Zigbee, BLE); cloud IoT services on AWS, Azure ...

Senior Associate, Full-Stack Engineer Opportunities

Hiring Organisation
The Bank of New York Mellon
Location
London, UK
Employment Type
Full-time
ways: Design, build, and maintain backend services, batches and APIs, contributing to UI components as needed. Own end-to-end delivery: implementation, testing, deployment, observability, and reliability. Write clean, well-tested code; participate in code reviews and continuous improvement. Collaborate with product, design, and operations to translate business needs into … programming concepts and microservicesProficiency in Java with Spring. Experience with CI/CD, automated testing (JUnit/Spock), and containers (Docker).Familiarity with microservices, observability/telemetry (e.g., Splunk, AppDynamics), and cloud deployments. Curiosity to understand the business domain and translate product strategy into technical solutions. Working knowledge of Groovy ...

Embedded Software Engineer

Location
Greater London, England, United Kingdom
best practices across the device lifecycle Edge and cloud integration: integrate devices with cloud IoT platforms and backend services; improve telemetry, health monitoring and observability; support reliable field operation and debug fleet issues Compliance and testing: support testing and validation for safety, EMC, radio and related requirements; prepare test plans …/CD workflows for embedded software Practical debugging experience using lab and software tools such as logic analysers, protocol analysers, network sniffers or observability platforms Strong problem-solving skills and the ability to work across hardware and software boundaries Bonus: low-power wireless (Thread, Zigbee, BLE); cloud IoT services ...

Infrastructure Engineer

Location
Greater London, England, United Kingdom
tenant environments for enterprise customers. Developer Velocity: Architect CI/CD pipelines (GitHub Actions, Docker) that allow our team to ship safely and quickly. Observability & Reliability: Instrument the stack with OpenTelemetry and Datadog to ensure we detect issues before our users do. AI Performance Tuning: Tune autoscaling and network routing … Haves: Experience with ECS, container orchestration, and distributed task queues (Celery/SQS). Strong Python skills. Familiarity with the Datadog/Grafana observability stack. A background in offensive security, CTFs, or AI/ML infrastructure. What We Offer Competitive Salary + Significant Equity: We want you to have true ...

Software Engineer / AI Engineer

Location
Greater London, England, United Kingdom
within a defined problem, building and testing tool use, retrieval pipelines and agent workflows, integrating AI capabilities into enterprise systems, and contributing to evaluation, observability and guardrails. You will hold a high bar on code quality, flag risks and blockers early, and work alongside host‐function stakeholders to make sure … agentic AI solutions to production standard within a defined technical approach. Implement and test tool use, retrieval pipelines, and agent workflows. Contribute to evaluation, observability and guardrails for agentic systems. Integrate AI capabilities into existing enterprise workflows and systems. Maintain high code quality and documentation so patterns can be reused. ...

Senior Software Engineer - London

Location
Greater London, England, United Kingdom
shorter path each time. Setting the standard for how a benchmark enters the framework, and building the checks that enforce it. Scheduling and observability, so cluster capacity isn't left idle while evaluation jobs queue. Reproducible results across trials, so a release decision rests on numbers that hold. Whatever stack … benchmarks, and started fixing what slows the framework down. By 6 months one part of it is yours, for example scaling the runs, observability, or a group of related benchmarks, and a release will have gone out on your numbers. By 12 months you'll know the design and trade ...

Junior DevOps Engineer

Location
Greater London, England, United Kingdom
/CD processes, automate manual processes to create a self-service environment for our developers, and maintain platform uptime SLAs by improving our observability stack. We use infrastructure-as-code to maintain our platform on AWS, so familiarity with common AWS Services (RDS, S3, ECS, EC2 etc.) and Terraform …/CD pipelines through Jenkins/Github Actions and other technology Support the Development and AI Engineering Teams - help troubleshoot their issues Bring observability through dashboards, alerting and log aggregation Run incident analysis and post mortems Environment management: ephe...dev/staging/prod What you’ll bring A strong foundation ...

Staff Python Engineer (ML)

Location
City Of London, England, United Kingdom
apps in a service architecture. Furthering Developer Experience (DevEx) by mentoring others in writing code that is intuitive, clear, and easy to test Developing observability for new and existing ML applications and GenAI/LLM integrations , making use of the Grafana Stack (Prometheus, Loki, Tempo) Develop integrations and services that … Backend-Engineering Experience owning projects from start to finish, including speccing, architecture, development, testing, deployment, release and monitoring Strong skills in building maintainable tests, observability and tracing systems. Knowledge of best practices for performance optimisation, memory management. Familiarity with Kubernetes , Docker and other cloud infrastructure, ops and containerised tools. Strong ...

Senior AI Engineer - Agentic AI

Location
Greater London, England, United Kingdom
that allow complex AI workflows to operate securely and efficiently at scale. You will be responsible for developing advanced orchestration capabilities, implementing evaluation and observability tooling and embedding enterprise controls for compliance and safety. If you are passionate about innovating with AI in real-world applications and scaling intelligent systems … ensuring graceful degradation and retries Apply enterprise security and governance practices including RBAC, prompt safety checks, traceability and secrets management Implement evaluation pipelines and observability frameworks using tools such as Langfuse, Arize or OpenTelemetry Contribute to architectural design decisions, code reviews and engineering standards for platform development Requirements Bachelor ...

Senior Network Engineer, Studios

Location
Greater London, England, United Kingdom
modelled in a NetBox source of truth, configuration is deployed from code through pipelines rather than hand‐edited, and streaming telemetry feeds the observability stack. Changes are expected to land in days, not months, in a facility where the network carries live content around the clock. The estate spans dedicated … defect. Config as code: build and run the automation that deploys configuration from intent through CI pipelines (Python, Ansible, Git); eliminate hand‐edits. Observability: develop monitoring built on streaming telemetry (gNMI/gRPC), flow analysis and modern tooling (Grafana, Zabbix class), serving the 24/7 operations teams as your ...

Member of Technical Staff (Platform Leaning)

Location
Greater London, England, United Kingdom
light up rather than back away: complex networking scenarios such as site-to-site VPN and private connectivity, multi-cloud and multi-region deployments, observability, and the machinery that makes Tessl straightforward to deploy, operate and support wherever it needs to run. We expect you to be relentlessly AI native … first year, you're shipping product features end to end like any Tessl engineer. When infrastructure work arises, from deployment architecture to networking to observability, you're the one who takes it on and lands it well, with support where you need it. Our deployment and operational practices are stronger ...

Jobshare - Sr Lead Software Engineer - Site Reliability Engineer, Python & Infrastructure management - Part time/Jobshare

Location
Greater London, England, United Kingdom
impact). Champions AI adoption and deliver AI-enabled capabilities to reduce operational toil and improve speed/quality of response. Sets direction for observability across logs/metrics/traces, including instrumentation standards, golden signals, and end-user journey monitoring. Improves alert quality and routing: reduce false positives, improve … design, build, and deliver automation and reliability solutions. Fluency & expertise in Python Deep practical knowledge of: SLOs/SLIs, error budgets, incident management, postmortems, observability design across metrics/logs/traces and distributed systems troubleshooting, resilience engineering, performance/capacity management, and change risk reduction. Proficiency and experience with ...

AI Solutions Architect

Location
Greater London, England, United Kingdom
rather than conceptual designs alone. Support early industrialisation of AI solutions, contributing to the adoption of MLOps, LLMOps and AIOps practices such as evaluation, observability, versioning and responsible deployment. Ensure AI solutions meet relevant standards and regulations for security, privacy and responsible AI, particularly in public sector and regulated environments. … preferred, and an understanding of PaaS, SaaS, real-time and batch processing patterns. Familiarity with Docker, CI/CD, infrastructure approaches, MLOps, LLMOps, AIOps, observability, versioning and production operations would be advantageous. Strong understanding of responsible AI, security, privacy, governance, data lifecycle and data quality considerations, including risks such ...

Principal Platform Engineer (12 Month FTC)

Location
Greater London, England, United Kingdom
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Senior AI Engineer

Hiring Organisation
MarkIT Placements
Location
London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
From £700 to £900 per day
Deploy AI systems across cloud, on-premises, sovereign and fully offline environments. Engineer for reliability, latency, security and predictable failure behaviour. Build evaluation and observability capabilities to understand model performance, agent behaviour and failure modes. Take end-to-end ownership of technical workstreams, from architecture through deployment. Establish engineering patterns … would be advantageous: Multimodal AI and reasoning. Edge or offline AI deployments. Kubernetes, particularly EKS or OpenShift. MLOps, including model evaluation, monitoring and reproducibility. Observability for agentic AI systems, including model performance, agent behaviour and drift. Agent orchestration and inter-agent communication protocols such as A2A. Model Context Protocol ...

Principal Product Engineer

Location
Greater London, England, United Kingdom
home in NestJS (or a similar Node.js framework) and React. Cloud and DevOps minded. Comfortable across GCP or a similar cloud, CI/CD, observability, and modern infrastructure tooling. Use AI daily. AI assistants are part of how you build. You bring back patterns that help others get more … code. You measure your work by what changed for the customer, not the lines of code shipped. Production minded. Solid grasp of API design, observability, scaling, reliability, and security. Care about craft. Clean APIs, attention to detail, and how the system feels to work in. Comfortable in uncertainty. You move ...

Senior Data Engineer

Hiring Organisation
Nuffield Hospitals
Location
London, United Kingdom
Salary
£ 60 K
Engineering team operates a distributed leadership model. Each Senior Data Engineer owns a defined functional area, such as ingestion and integration, incident management and observability, governance, security and cost management, or CI/CD and DevOps, and is accountable for the standards and resilience of that area so that nothing … develop it with structured support.Excellent communication skills, including documenting technical design proposals and translating complex technical concepts for technical and non-technical audiences.Desirable:Observability and incident management tooling, such as New Relic and ServiceNow.CI/CD, DevOps and infrastructure as code, such as Terraform and Azure DevOps.Data governance and security ...

Principal Software Engineer - Squad Lead Engineer

Location
City Of London, England, United Kingdom
complete complex bug fixes and performance improvements Define and uphold Definition of Ready/Done including code quality, automated test coverage, security checks, and observability Establish/maintain CI/CD pipelines, quality gates, and sensible branching/release strategies Drive a pragmatic quality strategy: test pyramid balance, contract tests … Windows Experience with relational and non‐relational data stores, performance tuning, and data modelling Knowledge of CI/CD platforms, containers, cloud technologies, observability, and monitoring practices Understanding of secure coding, performance optimisation, reliability engineering, and incident response Work in a Way That Works for You We promote a healthy ...

Principal Product Engineer

Hiring Organisation
Zapp
Location
London, UK
Employment Type
Full-time
home in NestJS (or a similar Node.js framework) and React. Cloud and DevOps minded. Comfortable across GCP or a similar cloud, CI/CD, observability, and modern infrastructure tooling. Use AI daily. AI assistants are part of how you build. You bring back patterns that help others get more … code. You measure your work by what changed for the customer, not the lines of code shipped. Production minded. Solid grasp of API design, observability, scaling, reliability, and security. Care about craft. Clean APIs, attention to detail, and how the system feels to work in. Comfortable in uncertainty. You move ...

Senior Software Engineer, Data

Location
Greater London, England, United Kingdom
/day scale Contribute to the technical design of the Data Transfer Hub, making pragmatic trade-offs across UX, careliability, throughput, cost, partner constraints, observability and operational support Build parallelised, distributed data transfer pipelines using Flyte for workflow orchestration, Kafka/event-driven patterns for lifecycle tracing, and Azure … Background in autonomous vehicles, robotics, geospatial data, video processing or other data-intensive domains Experience with Kafka or other event-driven systems for transfer observability and lifecycle tracking Experience building dashboards to communicate transfer job status Multi-cloud experience across Azure, AWS and/or GCP Experience improving cost efficiency ...

Data Platform Engineer- BPL-CIO

Location
Greater London, England, United Kingdom
automation, GitOps operating models AWS Platform Engineering Hands‐on AWS engineering with a focus on: IAM and security patterns, Networking integration, Storage and encryption, Observability, Resilience and operational readiness, Automation and supportability Databricks on AWS Experience with: Workspace deployment, Databricks Terraform provider, Identity integration, Storage integration, Network connectivity patterns Platform … Engineering/Internal Developer Platforms Experience building: Self‐service capabilities, Golden paths, Paved roads, Reusable engineering patterns, Internal platform products Observability, Reliability & Operational Engineering Experience with: Monitoring and alerting, Operational telemetry, Incident management, Resilience testing, Recovery automation, Production support You may be assessed on the key critical skills relevant ...