3,326 to 3,350 of 4,288 Permanent Observability Jobs

Senior Software Engineer - Runtime Platform, Robot Software Sunnyvale, California USA

Location
Greater London, England, United Kingdom
support new product features, which is critical to the success of Wayve’s mission. The Runtime Platform team equips all Wayve teams with the observability, profiling tools, and infrastructure needed to understand and optimise software performance across our development fleet. We work closely with teams to investigate issues, reduce bottlenecks … caches, context switches, ...), and thread synchronisation Desirable Familiarity with Nvidia performance tools such as NV NSight, NV Lumos and tegrastat Familiarity with observability tools such as Grafana (logs, metrics, traces), Databricks, Datadog Familiarity with QNX and Momentics is a plus This is a full-time role based in London. ...

Software Engineer, ChatGPT Infrastructure

Location
Greater London, England, United Kingdom
diagnosing performance, scalability, or reliability issues in production environments. Understanding of distributed systems, data storage, concurrency, asynchronous processing, or networking. Familiarity with modern deployment, observability, and cloud infrastructure practices. Ability to lead complex technical work and collaborate effectively across product and infrastructure teams. Experience with cloud infrastructure, containerized environments … observability tools is useful but not required. We welcome candidates from backend engineering, distributed systems, platform engineering, and reliability backgrounds. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general‐purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities ...

Backend Developer

Location
Greater London, England, United Kingdom
join our team and help buildthe foundations of the global EV transition. Responsibilities Enable us to build quickly but robustly, using best practices in observability, data analysis and defensive programming, allowing us to safely scale our products and CPO integrations globally. Work with cross-functional stakeholders to break down complex ...

Lead DevSecOps Engineer

Location
Greater London, England, United Kingdom
Build automation that improves the speed and quality of software delivery Embed security best practices throughout the development lifecycle Improve infrastructure reliability, scalability and observability Establish engineering standards across infrastructure and security Provide technical leadership and support the development of other engineers About You: Strong background in DevOps, Platform, Infrastructure ...

Staff Site Reliability Engineer

Location
Greater London, England, United Kingdom
multi-trillion-row, petabyte scale — sharding and replication, materialised views, merge and query optimisation, tenant isolation and cost/performance trade-offs. Owning SLOs, observability, capacity planning and incident response for data-intensive systems and pipelines, alongside the orchestration and job execution that power them. Shaping Data-as-a-Service … level rather than as a black box — and ideally have contributed code upstream. Reliability engineering for data platforms. You bring true SRE discipline — SLOs, observability, capacity planning and incident response — to analytical data systems and pipelines. Data-as-a-Service productisation. You think in terms of data as a product ...

Senior Full Stack Developer

Location
Greater London, England, United Kingdom
engineers, mentoring junior team members. Partner with product, design and client stakeholders to translate business goals into technical solutions. Drive engineering excellence — testing, observability, performance and security best practice. Contribute to internal frameworks, design systems and reusable libraries across our consulting practice. What we’re looking for 5+ years ...

Engineering Manager - People & Procurement New London

Location
Greater London, England, United Kingdom
owns the integrations and data flows connecting People & Procurement platforms with the wider FT technology estate. Own operational performance and establish effective approaches to observability, incident and problem management, change management, resilience, disaster recovery and business continuity. Ensure relevant security, privacy, data governance, compliance and technology standards are met, with … system dependencies. Experience owning technology across its lifecycle, from change and delivery through production operation and continuous improvement. Strong understanding of reliability, security, privacy, observability, operational risk and technology lifecycle management. Experience translating business priorities into engineering plans and balancing feature delivery with platform health and technical investment. Experience leading ...

Staff Software Engineer, AI Reliability Engineering

Location
Greater London, England, United Kingdom
Responsibilities Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity. Design and implement monitoring and observability systems across the token path. Assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud providers Lead incident response … more ML hardware accelerators (GPUs, TPUs, Trainium). Understand ML-specific networking optimizations like RDMA and InfiniBand. Have expertise in AI-specific observability tools and frameworks. Have experience with chaos engineering and systematic resilience testing. Have contributed to open-source infrastructure or ML tooling. The annual compensation range for this ...

Senior Software Engineer (£70k + benefits)

Location
Manchester, England, United Kingdom
engineering best practices such as TDD, SOLID principles, and pair programming. Mentor other team members and collaborate with product managers. Qualifications TypeScript (Node, React) Observability tools (Datadog, Dynatrace, Honeycomb, CloudWatch, etc.) Experience working in a DevOps-enabled, cloud-native environment is desirable Benefits Salary up to £70k plus benefits. Hybrid ...

IT Infrastructure Solutions Architect

Location
Cambridge, England, United Kingdom
roadmap.**Key responsibilities*** Define and maintain reference architectures and target-state designs for VMware VCF 9.0 platform architecture and lifecycle patterns.* Define Aria Operations observability strategy (telemetry standards, alert philosophy, capacity/performance governance, service reporting) and ensure operational adoption.* Define VCF Automation platform approach (catalog/service design, templates … iSCSI), VSAN, NAS, and software-defined storage concepts.* Experience or exposure to infrastructure-as-code* Proven capability to architect and operationalize enterprise monitoring/observability standards (Logic Monitor and Aria Operations).* Proven capability to architect, govern, and troubleshoot provisioning automation (VCF Automation).* Proven backup/recovery architecture ...

Frontend Engineer

Location
United Kingdom
engineers, mentoring junior team members. Partner with product, design and client stakeholders to translate business goals into technical solutions. Drive engineering excellence — testing, observability, performance and security best practice. Contribute to internal frameworks, design systems and reusable libraries across our consulting practice. What we're looking for 5+ years ...

IT Infrastructure Solutions Architect

Location
Greater London, England, United Kingdom
roadmap.**Key responsibilities*** Define and maintain reference architectures and target-state designs for VMware VCF 9.0 platform architecture and lifecycle patterns.* Define Aria Operations observability strategy (telemetry standards, alert philosophy, capacity/performance governance, service reporting) and ensure operational adoption.* Define VCF Automation platform approach (catalog/service design, templates … iSCSI), VSAN, NAS, and software-defined storage concepts.* Experience or exposure to infrastructure-as-code* Proven capability to architect and operationalize enterprise monitoring/observability standards (Logic Monitor and Aria Operations).* Proven capability to architect, govern, and troubleshoot provisioning automation (VCF Automation).* Proven backup/recovery architecture ...

AI Productivity Engineer

Hiring Organisation
Aircall
Location
London, UK
Employment Type
Full-time
productionize AI-driven solutions with strong autonomy on how problems are solvedAutomate and streamline workflows across GitLab, Jira, CI/CD, Slack, and observability toolsDesign and operate internal AI services and orchestration layers (e.g. MCP servers)Own solutions end-to-end: discovery design build measure iterateWork hands-on with engineering … injectionAI-powered tooling or internal platformsSolid backend engineering skills (APIs, services, integrations)Experience working with developer tools (CI/CD, GitHub/GitLab, Jira, observability)Strong product mindset and comfort operating in ambiguous problem spacesNice to HaveParticularly interesting profiles are engineers who have built developer tools and are now evolving ...

Associate Director, Connectivity

Location
Oxford, England, United Kingdom
strategy – a clear point of view on where supplier and Restech integrations need to go, aligned to our broader technical direction (faster releases, better observability, simplification, AI driven efficiencies). Build and maintain C‐suite level relationships with Restech and Supplier partners, acting as the senior technical face of Tripadvisor … delivering technical roadmaps, API strategies, and scaling connectivity platforms in high‐traffic environments. Integration Expertise: Deep knowledge of integration architecture, API design, and observability, with the credibility to engage hands‐on with engineering teams. Team Development: Proven ability to build, mentor, and lead high‐performing global engineering and technical teams ...

Associate Director, Connectivity

Location
Greater London, England, United Kingdom
strategy – a clear point of view on where supplier and Restech integrations need to go, aligned to our broader technical direction (faster releases, better observability, simplification, AI driven efficiencies). Build and maintain C‐suite level relationships with Restech and Supplier partners, acting as the senior technical face of Tripadvisor … delivering technical roadmaps, API strategies, and scaling connectivity platforms in high‐traffic environments. Integration Expertise: Deep knowledge of integration architecture, API design, and observability, with the credibility to engage hands‐on with engineering teams. Team Development: Proven ability to build, mentor, and lead high‐performing global engineering and technical teams ...

Engineering Manager Sotheby's · London / Remote

Location
Greater London, England, United Kingdom
priorities and deliver meaningful outcomes Track record of hiring well, onboarding effectively, and building teams that retain good people A focus on engineering quality: observability, analytics and reporting, testing practices, and ensuring the team ships work they can stand behind Useful: Familiarity with our core stack: Go, Scala, React/… feature delivery, technical investment and operational health Set and maintain the quality bar for your team’s area of the platform, including engineering practices, observability, reliability and how work gets reviewed Run the delivery cadence: planning, prioritisation, retrospectives, and the day-to-day work of keeping things moving and unblocked ...

AI / ML Engineer

Location
Greater London, England, United Kingdom
engineers, mentoring junior team members. Partner with product, design and client stakeholders to translate business goals into technical solutions. Drive engineering excellence — testing, observability, performance and security best practice. Contribute to internal frameworks, design systems and reusable libraries across our consulting practice. What we're looking for 5+ years ...

Senior Systems Engineer, Workers AI

Location
Greater London, England, United Kingdom
reduce latency for end-users. System Reliability: Drive significant, measurable improvements in the platform's reliability and resilience by identifying and mitigating systemic risks. Observability: Expand and refine the observability stack (metrics, logging, tracing) and fine-tune alerts to proactively identify and resolve production issues. Technical Leadership: Lead complex, cross ...

MLOps Engineer

Location
Greater London, England, United Kingdom
doing: Designing and maintaining end-to-end ML pipelines across the full lifecycle Deploying and monitoring models in production across multiple compute environments Building observability and alerting layers to catch performance regressions before they reach users Implementing CI/CD pipelines with automated quality checks Enabling reproducible experimentation ...

Engineering Lead – ERP Services

Location
Greater London, England, United Kingdom
impacts. Operational Excellence & Risk Owning the operational performance of ERP Services and ensuring critical services are appropriately monitored and supported. Establishing effective approaches to observability, incident management, problem management, change management and operational readiness. Ensuring recurring incidents and operational problems are investigated to their underlying causes and addressed sustainably. Maintaining … developing engineering capability and strengthening technical ownership within teams. Experience working with Finance or other complex enterprise business domains. Strong understanding of reliability, security, observability, operational risk and technology lifecycle management. Experience translating business priorities into engineering plans and balancing feature delivery with platform health and technical investment. Experience leading ...

Software Engineer

Hiring Organisation
Randstad Technologies
Location
Manchester, Lancashire, United Kingdom
Employment Type
Full-Time
Salary
£55.00 - £62.00 per hour
Looking For Proven experience as a Software Engineer building and scaling high-performance backend systems. Experience working with modern AWS architectures Strong understanding of observability, reliability engineering (SLIs/SLOs), and production monitoring. Experience dealing with highly concurrent systems and A/B testing or experimentation platforms. Strong experience with ...

Forward Deployment Engineer - AI, Software, Full-Stack, Database, SC Cleared

Hiring Organisation
Bangura Solutions
Location
London, UK
Employment Type
Full-time
teams. Up-to-date knowledge of the latest AI advancements. Nice-to-Haves: Experience with our tech stack: NextJS, FastAPI, Postgres. Familiarity with LLM observability and evaluation tools like Langfuse. Cloud infrastructure experience (Terraform, Azure, etc.).Background in startups or entrepreneurial environments. Minorities, women, LGBTQ+ candidates, and individuals with disabilities ...

ml engineer in financial services

Location
Greater London, England, United Kingdom
production systems. Задачи Establish standard, reusable patterns for model serving, pipelines, and deployment; Define standards for model packaging, versioning, CI/CD, monitoring, and observability; Set production-readiness criteria for models entering engineering workflows; Provide technical direction for ML solutions delivered within the team; Evolve platform capabilities to reduce bespoke ...

Global Synthetics Technology Lead

Location
Greater London, England, United Kingdom
global stakeholders across Trading, Risk, Operations, Compliance, and Technology Drive adoption of modern engineering practices, including DevOps, CI/CD, automated testing, and observability Qualifications Proven experience in financial services, with strong exposure to synthetics, swaps, and prime services with appropriate oversight of high-volume swap calculation engines Proven track ...

Governance Architect - £600p/d Inside IR35 - Hybrid

Location
Greater London, England, United Kingdom
cases. Create reusable architecture and cloud patterns for delivery teams. Define cloud patterns for AI workloads across identity, networking, data access, integration, observability, resilience and cost control. Produce technology and cloud roadmaps covering current state, target state, gaps and dependencies. Define operational standards covering naming, tagging, tooling, monitoring, logging, resilience ...