3,026 to 3,050 of 3,965 Observability Jobs

Senior/Staff Software Engineer (Nova Core)

Location
Greater London, England, United Kingdom
operations across Nova Cloud deployments. This role focuses on the Nova Core “inner loop”: service architecture, APIs, data models, persistence, authn/authz, observability, and developer experience that other Nova modules and product teams depend on. What you’ll do Own and ship critical Nova Core backend services (e.g., common … engineering and product teams. What success looks like Core Nova services are delivered, adopted, and operated reliably with clear SLIs/SLOs and runbooks. Observability is strong enough that incidents are detected quickly and resolved faster over time (improving MTTD/MTTR). API versioning and compatibility practices reduce integration ...

Sr Director, Platform Engineering – Data Platform & Agentic Platform

Location
Greater London, England, United Kingdom
operate agent workflow platform capabilities aligned to product‐defined standards and interfaces, including traceability, state handling, and convergence patterns Implement production‐grade evaluation, observability, auditability, and guardrail mechanisms required for safe AI workflows Implement security controls, access governance, encryption, and audit requirements in partnership with InfoSec while ensuring enterprise SDLC … large‐scale SaaS systems with production operations accountability Demonstrated success building and operating platforms adopted by multiple product teams, including reliability discipline (SLOs), observability, and incident management Strong hands‐on technical leadership background in distributed systems and platform engineering Deep experience with data platform engineering at scale, including ingestion ...

Operations and SRE Manager

Location
Greater London, England, United Kingdom
internal and external customers. You will be responsible for driving reliability improvements, advancing automation and AI-Ops capabilities, and leading a team focused on observability, incident response, operational excellence, and continuous improvement. Responsibilities Lead the transformation of the Operations function towards an AI-Ops operating model, driving the adoption … improvement actions are owned, tracked and completed. Strengthen operational process adherence, ensuring responsibilities are clear and delegation is effective. Drive SRE practices across observability, automation, disaster recovery, design for reliability, on-call readiness and production support. Protect service levels by ensuring engineering effort is balanced across InfoSec commitments, operational tickets ...

Front Office Equities Trading Technology Support

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
firm's systems to ensure operational stability and availabilityAssist in the monitoring of production environments for anomalies and address issues utilizing standard observability toolsIdentity issues for escalation and communication, and provide solutions to the business and technology stakeholdersAnalyze complex situations and trends to anticipate and solve incident, problem, and change … meaningful relationships to achieve common goalsDemonstrates knowledge of applications or infrastructure in a large-scale technology environment both on premises and public cloudExperience in observability and monitoring tools and techniquesExposure to processes in scope of the Information Technology Infrastructure Library (ITIL) framework Preferred qualifications, capabilities, and skillsExperience with ...

Front Office Equities Trading Technology Support

Location
Greater London, England, United Kingdom
firm’s systems to ensure operational stability and availability Assist in the monitoring of production environments for anomalies and address issues utilizing standard observability tools Identity issues for escalation and communication, and provide solutions to the business and technology stakeholders Analyze complex situations and trends to anticipate and solve incident … achieve common goals Demonstrates knowledge of applications or infrastructure in a large-scale technology environment both on premises and public cloud Experience in observability and monitoring tools and techniques Exposure to processes in scope of the Information Technology Infrastructure Library (ITIL) framework Preferred qualifications, capabilities, and skills Experience with ...

Software Engineer (Simulation, Evaluation, Validation)

Location
Greater London, England, United Kingdom
differences between on-road and simulated execution, identifying issues across data, inference and simulated components Improve simulation reproducibility, reliability and debuggability through automated testing, observability and better developer tooling Profile and improve simulator performance, helping us run increasingly large evaluation workloads efficiently Work with internal users and adjacent engineering teams … sensor data such as camera, radar, lidar or GNSS, including modelling uncertainty or noiseExperience integrating machine-learning inference into production systemsExperience with performance profiling, observability or debugging distributed systemsFamiliarity with large-scale batch processing, cloud infrastructure or GPU-based workloads #J-18808-Ljbffr ...

Professional Services Consultant

Location
Greater London, England, United Kingdom
ITRS, we make society's critical technology work. Our mission is to deliver automated and holistic IT observability solutions that safeguard critical applications and enable innovation. We are the only monitoring and observability platform designed for the most demanding and regulated industries — trusted by 90% of Tier 1 capital markets ...

Customer Solutions Engineer

Location
Greater London, England, United Kingdom
About ITRS At ITRS, we make society's critical technology work. Our mission is to deliver automated and holistic IT observability solutions that safeguard critical applications and enable innovation. We are the only monitoring and observability platform designed for the most demanding and regulated industries — trusted by 90% of Tier ...

AI Product Analyst

Hiring Organisation
The Portfolio Group
Location
London, United Kingdom
Employment Type
Permanent
Salary
£80000 - £85000/annum
fact. Day-to-Day Responsibilities Product Performance & Analytics: Measure how our AI products perform across retrieval quality, correctness of output, and user engagement. Observability Management: Own production quality observability, tracking metrics like thumbs-down rates, regeneration rates, task abandonment, and usage drift. Dashboard Engineering: Build and maintain the daily analysis ...

python developer market data

Location
Greater London, England, United Kingdom
data. Задачи Design and build data platform components using Python; Develop and optimise ETL pipelines, Data Lakes/Lakehouses and distributed systems; Improve infrastructure, observability and CI/CD practices; Maintain and evolve existing systems, including microservices, ETL and Excel add-ins; Support platform operations to ensure reliability and performance. ...

Site Reliability Engineer (DV Security Clearance)

Location
Manchester, England, United Kingdom
Engineer (SRE) to join a high-performing team supporting multiple data product and platform groups. This role is focused on improving the reliability, scalability, observability, deployment, and operational support of critical data-driven platforms and services operating within complex production environments. The successful candidate will work closely with engineering, platform ...

Senior Cloud Data Engineer - KSP

Location
Greater London, England, United Kingdom
optimise streaming data pipelines (Kafka/Flink or equivalent) to enable near real‐time data availability. Ensure data quality and reliability through validation frameworks, observability, and robust handling of late‐arriving or inconsistent data. Design data contracts and schemas that enable reliable integration between upstream event producers and downstream consumers. … medallion architecture or similar data layering approaches. Experience working with streaming technologies (Kafka, Flink, or similar). Strong understanding of data quality, testing, and observability practices. Experience designing schemas and handling data consistency challenges in distributed systems. Ability to work closely with stakeholders to translate business needs into scalable data ...

Data Architect

Hiring Organisation
Tiro Partners
Location
London, United Kingdom
Employment Type
Permanent
connect Design data models, pipelines, storage and governance frameworks Build reusable data capabilities, services and technical components Set standards across APIs, integration, deployment and observability Work closely with engineering, infrastructure, security and data teams Help ensure AI solutions are scalable, secure and maintainable Requirements 6+ years experience Data Architecture/ ...

Service Engineer

Hiring Organisation
Hackajob Ltd
Location
Knutsford, Cheshire, North West, United Kingdom
Employment Type
Permanent
line with ITIL processes, and working with engineering teams to improve application reliability, performance, and operational resilience. You will also leverage monitoring and observability tools to proactively identify issues and minimise service disruption. To be successful as an Application Support Engineer, you should have: Strong AWS knowledge from an application … such as GitLab pipelines, with an understanding of release and deployment practices Experience using monitoring and automation tools such as Kibana, AppDynamics and other observability platforms to support service performance and operational excellence You may be assessed on the key critical skills relevant for success in the role, such ...

Lead DevSecOps

Location
Greater London, England, United Kingdom
technical direction Genuine autonomy and influence from day one Build and mentor the DevSecOps function as the company scales Build security, compliance and observability into the development lifecycle What they're looking for Experience building internal developer platforms or sophisticated CI/CD ecosystems Deep expertise across Kubernetes and cloud ...

Senior Software Engineer - Runtime Platform, Robot Software

Location
Greater London, England, United Kingdom
support new product features, which is critical to the success of Wayve’s mission. The Runtime Platform team equips all Wayve teams with the observability, profiling tools, and infrastructure needed to understand and optimise software performance across our development fleet. We work closely with teams to investigate issues, reduce bottlenecks … caches, context switches, ...), and thread synchronisation Desirable Familiarity with Nvidia performance tools such as NV NSight, NV Lumos and tegrastat Familiarity with observability tools such as Grafana (logs, metrics, traces), Databricks, Datadog Familiarity with QNX and Momentics is a plus This is a full-time role based in London. ...

Software Engineer, ChatGPT Infrastructure

Location
Greater London, England, United Kingdom
diagnosing performance, scalability, or reliability issues in production environments. Understanding of distributed systems, data storage, concurrency, asynchronous processing, or networking. Familiarity with modern deployment, observability, and cloud infrastructure practices. Ability to lead complex technical work and collaborate effectively across product and infrastructure teams. Experience with cloud infrastructure, containerized environments … observability tools is useful but not required. We welcome candidates from backend engineering, distributed systems, platform engineering, and reliability backgrounds. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general‐purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities ...

Backend Developer

Location
Greater London, England, United Kingdom
join our team and help buildthe foundations of the global EV transition. Responsibilities Enable us to build quickly but robustly, using best practices in observability, data analysis and defensive programming, allowing us to safely scale our products and CPO integrations globally. Work with cross-functional stakeholders to break down complex ...

Lead DevSecOps Engineer

Location
Greater London, England, United Kingdom
Build automation that improves the speed and quality of software delivery Embed security best practices throughout the development lifecycle Improve infrastructure reliability, scalability and observability Establish engineering standards across infrastructure and security Provide technical leadership and support the development of other engineers About You: Strong background in DevOps, Platform, Infrastructure ...

Linux Engineer

Location
Milton, Scotland, United Kingdom
changes while improving resilience and reducing risk. Working within a highly skilled engineering team, you’ll help strengthen pre and post-change validation, enhance observability, and support the continuous improvement of automation capabilities. Technology Stack Linux/UNIX Python Ansible Apache Airflow Prometheus Grafana Loki VMware F5 What ...

Senior Full Stack Developer

Location
Greater London, England, United Kingdom
engineers, mentoring junior team members. Partner with product, design and client stakeholders to translate business goals into technical solutions. Drive engineering excellence — testing, observability, performance and security best practice. Contribute to internal frameworks, design systems and reusable libraries across our consulting practice. What we’re looking for 5+ years ...

Staff Software Engineer, AI Reliability Engineering

Location
Greater London, England, United Kingdom
Responsibilities Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity. Design and implement monitoring and observability systems across the token path. Assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud providers Lead incident response … more ML hardware accelerators (GPUs, TPUs, Trainium). Understand ML-specific networking optimizations like RDMA and InfiniBand. Have expertise in AI-specific observability tools and frameworks. Have experience with chaos engineering and systematic resilience testing. Have contributed to open-source infrastructure or ML tooling. The annual compensation range for this ...

Engineering Manager - People & Procurement New London

Location
Greater London, England, United Kingdom
owns the integrations and data flows connecting People & Procurement platforms with the wider FT technology estate. Own operational performance and establish effective approaches to observability, incident and problem management, change management, resilience, disaster recovery and business continuity. Ensure relevant security, privacy, data governance, compliance and technology standards are met, with … system dependencies. Experience owning technology across its lifecycle, from change and delivery through production operation and continuous improvement. Strong understanding of reliability, security, privacy, observability, operational risk and technology lifecycle management. Experience translating business priorities into engineering plans and balancing feature delivery with platform health and technical investment. Experience leading ...

Senior Software Engineer (£70k + benefits)

Location
Manchester, England, United Kingdom
engineering best practices such as TDD, SOLID principles, and pair programming. Mentor other team members and collaborate with product managers. Qualifications TypeScript (Node, React) Observability tools (Datadog, Dynatrace, Honeycomb, CloudWatch, etc.) Experience working in a DevOps-enabled, cloud-native environment is desirable Benefits Salary up to £70k plus benefits. Hybrid ...

SUSE Rancher Platform Deployment Expert

Hiring Organisation
Randstad Technologies Recruitment
Location
Nationwide, United Kingdom
Employment Type
Contract
Contract Rate
£400 - £500/day
controls. Build integration paths between an on-premises VMware vSphere footprint and an enterprise Azure landing zone. Establish native cluster monitoring, logging, and operational observability stacks. Author technical handoff documentation, cluster runbooks, and disaster recovery procedures. Technical Requirements Deep experience in SUSE Rancher Prime implementation and cluster lifecycle management. Direct ...