3,726 to 3,750 of 4,060 Permanent Observability Jobs

Senior React Front-End Lead, Equities Tech

Location
Greater London, England, United Kingdom
cross-functional partners in a fast-paced, high-stakes environment. You will mentor mid-level developers, ensure robust CI/CD pipelines, and uphold observability and testing standards while aligning with regulatory and risk controls. #J-18808-Ljbffr ...

Software Engineer, Platform Systems

Location
City Of London, England, United Kingdom
systems that provide visibility into large‐scale training workloads and help operate them reliably at scale. You’ll work on failure detection, tracing, and observability systems that identify slow or faulty nodes, surface performance bottlenecks, and help engineers understand and optimize massive distributed training jobs. This infrastructure is critical … systems for large‐scale AI training jobs Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable visibility into system behavior Improve observability, reliability, and performance across OpenAI’s training platform Debug and resolve issues in complex, high‐throughput distributed systems Collaborate with systems, infrastructure, and research teams ...

Lead GCP Data Platform Architect (Real-Time)

Location
Greater London, England, United Kingdom
service patterns for mature divisions to build trusted models using Core Lake data. You will lead a team of data engineers, ensuring data quality, observability, and SLA/SLO compliance #J-18808-Ljbffr ...

Technology Support Director, Risk Engagement

Location
Greater London, England, United Kingdom
cross-domain incident, problem and change management challenges, identifying process weaknesses, recurring failure patterns and opportunities to improve operational resilience across Infrastructure Platforms. Uses observability, monitoring and diagnostic information to analyse significant, recurring or cross‐domain issues, test root‐cause hypotheses and help accountable teams define sustainable remediation. Leads focused … analysing complex application or infrastructure issues in large‐scale on‐premises and public‐cloud environments, including dependencies across multiple technology domains. Proficient in using observability, monitoring and diagnostic information to identify patterns, test hypotheses and support root‐cause analysis across distributed technology environments. Experience working in a client‐facing technology ...

Infrastructure Engineer - Azure/M365

Hiring Organisation
The Curve Group
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
Salary negotiable
Using PowerShell to automate administration, remediation and operational tasks Investigating complex incidents, carrying out root-cause analysis and driving permanent fixes Supporting monitoring and observability across infrastructure and applications, identifying issues before they become major incidents Working with backup, disaster recovery and resilience, including understanding RPO/RTO requirements Supporting … deployment and renewal Building and managing complex Group Policy environments Using PowerShell to automate repetitive administration or remediate issues at scale Implementing monitoring/observability and using logs and telemetry to investigate incidents Supporting environments with significant numbers of users, devices, servers or applications Working within structured ITIL/change ...

Senior Data Engineer Remote EU/USA and UK Python, Spark, Airflow, DBT £70-130K base • Permanent[...]

Location
United Kingdom
maintain ELT with DBT and Airflow Scale batch and streaming with Spark Define data contracts with product and engineering Improve lineage, testing and observability Requirements 5+ years in data engineering with Python DBT and Airflow in production Experience with Spark or similar frameworks Great communication and documentation habits Remote within ...

AI Platform Engineer

Location
Rochdale, England, United Kingdom
doing something twice, you turn it into a platform. What you'll build Internal SDKs and gateways that abstract over LLM providers. Shared evaluation, observability and cost-attribution tooling. Multi-tenant data infrastructure with strong isolation. Requirements 5+ years engineering with platform/infra emphasis. Strong systems fundamentals; comfortable with ...

Embedded Software Engineer (Space) - Berkshire

Location
England, United Kingdom
mode management, sequencing, and autonomous FDIR for safe and productive operations. Deliver reliable real-time performance – manage concurrency, timing, CPU/memory budgets, and observability under tight constraints. Build verification infrastructure – prototypes, SIL/HIL test harnesses, simulations, and telemetry analysis tooling to validate designs early. Ship code from review … UART, CAN(-FD), GPIO, etc.) and practical lab debugging. Solid software engineering fundamentals: architecture, code review, static analysis, CI/CD, configuration management, and observability/logging. Ability to own systems end-to-end: from requirements and design through implementation, verification, operations support, and iterative improvement. Nice-to-haves Experience ...

DFT (Design for Test) Engineer

Hiring Organisation
ASD GLOBAL
Location
Santa Clara, California, United States
Employment Type
Permanent
Salary
USD Annual
Clara, California • 5+ years of hands-on experience in DFT and ATPG for SoC or ASIC designs • Strong understanding of DFT fundamentals including controllability, observability, and scan-based testing • Proven expertise in ATPG pattern generation, analysis, and debug • Experience with MBIST, including memory test architectures and diagnostics • Knowledge ...

Senior Software Engineer, Trading Team

Hiring Organisation
DRW
Location
London, UK
Employment Type
Full-time
strategies with execution systems and support live trading workflowsBuild tooling for real-time monitoring, diagnostics, performance attribution, and post-trade analysisImprove system design, scalability, observability, operational resilience, and maintainabilityPreferred Qualifications5+ years of hands-on software engineering experience designing, building and operating systems with high complexity. Prior experience in equity stat … data, research, back testing, execution, monitoring, and post-trade analysis. Strong system design skills and sound engineering judgment. High standards for correctness, reliability, testing, observability, and maintainabilityFor more information about DRW's processing activities and our use of job applicants' data, please view our Privacy Notice at . California residents ...

Senior Software Engineer - Systematic Trading Infrastructure

Location
Greater London, England, United Kingdom
execution systems and support live trading workflows Build tooling for real-time monitoring, diagnostics, performance attribution, and post-trade analysis Improve system design, scalability, observability, operational resilience, and maintainability Preferred Qualifications 5+ years of hands-on software engineering experience designing, building and operating systems with high complexity. Prior experience … data, research, back testing, execution, monitoring, and post-trade analysis. Strong system design skills and sound engineering judgment. High standards for correctness, reliability, testing, observability, and maintainability For more information about DRW's processing activities and our use of job applicants' data, please view our Privacy Notice at https:/ ...

Senior Software Engineer, Trading Team DRW · London, United Kingdom 22 hours ago

Location
Greater London, England, United Kingdom
execution systems and support live trading workflows Build tooling for real‐time monitoring, diagnostics, performance attribution, and post‐trade analysis Improve system design, scalability, observability, operational resilience, and maintainability Preferred Qualifications 5+ years of hands‐on software engineering experience designing, building and operating systems with high complexity. Prior experience … data, research, back testing, execution, monitoring, and post‐trade analysis. Strong system design skills and sound engineering judgment. High standards for correctness, reliability, testing, observability, and maintainability For more information about DRW's processing activities and our use of job applicants' data, please view our Privacy Notice at https:/ ...

Site Reliability Engineer II

Hiring Organisation
BC Forward
Location
Charlotte, North Carolina, United States
Employment Type
Permanent
Salary
USD Hourly
Description We are seeking a Site Reliability Engineer II to join our dynamic team. The ideal candidate will have strong experience in SRE strategy, observability, automation, and reliability by design and a proven ability to reduce incidents, improve recovery, and scale engineering-led operations across complex banking and payments environments. … Responsibilities: Design and lead SRE practices, standards, and governance across banking and payments domains. Establish and automate observability, including metrics, logging, tracing, and alerting. Implement SLO/SLI and error budget frameworks to drive engineering priorities and accountability. Reduce incidents, MTTR, and operational toil through automation and resilient engineering patterns. ...

AI Large Language Model (LLM) Technology Architecture Senior Manager/Associate Director

Hiring Organisation
Hackajob Ltd
Location
London, UK
work of multiple domain architects and subject matter specialists — across areas such as agentic application design, AI security and trust, AI operations and observability, data and knowledge engineering, and model platforms and inference — providing the architectural vision, technical governance, and cross-domain coherence that binds their contributions into a unified … deployment within cohesive, production-ready platforms. You will be accountable for ensuring the complete architecture meets the most rigorous non-functional requirements across security, observability, governance, performance, and scalability — and that these concerns are addressed holistically and consistently across all domains. You will produce and govern the authoritative architecture artifacts ...

Sales Sector Lead - Public Sector

Location
Greater London, England, United Kingdom
advantage of all structured and unstructured data - securing and protecting private information more effectively - Elastic's complete, cloud-based solutions for search, security, and observability help organizations deliver on the promise of AI. Elastic is looking for an exceptional invidual to own and grow our Justice & Public Safety business across … Build trusted relationships with senior leaders across government and Blue Light organisations. Lead complex strategic pursuits and expand Elastic's platform across Search, Security, Observability, and AI. Collaborate across Sales, Solutions Architecture, Services, Customer Success, Marketing, and Partners to deliver customer outcomes. Develop customer and partner ecosystems while mentoring colleagues ...

Solutions Engineer

Location
Manchester, England, United Kingdom
solutions into actionable, strategic plans that align with your customers’ priorities. You will work across a comprehensive portfolio of products and services, including Networking, Observability, Security, Collaboration, Compute and AI – crafting multi-architecture solutions to meet customer’s specific needs. You will work on long term, complex and strategic initiatives … these areas: Enterprise Networking technologies (Switching, Wireless, Routing); Network Automation; Programmability; Data Centre and Cloud Technologies; On-Prem and Cloud Security; Applications Observability; Cloud Networking; IoT. Proven ability to work autonomously, taking ownership of tasks and workstreams. Understanding of Enterprise technology, industry trends and market drivers. Excellent communication skills, both ...

AI Large Language Model (LLM) Technology Architecture Senior Manager/Associate Director

Location
City Of London, England, United Kingdom
work of multiple domain architects and subject matter specialists — across areas such as agentic application design, AI security and trust, AI operations and observability, data and knowledge engineering, and model platforms and inference — providing the architectural vision, technical governance, and cross-domain coherence that binds their contributions into a unified … deployment within cohesive, production-ready platforms. You will be accountable for ensuring the complete architecture meets the most rigorous non-functional requirements across security, observability, governance, performance, and scalability — and that these concerns are addressed holistically and consistently across all domains. You will produce and govern the authoritative architecture artifacts ...

Lead Salesforce Implementation Specialist

Location
Sheffield, England, United Kingdom
system of record to an agentic platform.Reporting to the Head of CRM, you will lead the Agentforce function at UniHomes - owning design, build, governance, observability and adoption across Sales and Service, and growing a small team around you as the work scales. You will work shoulder-to-shoulder with … agent governance: trust layer configuration, audit, evaluation and the policies that keep our agents safe to put in front of customers. Stand up agent observability so we know how agents are performing in production and can iterate quickly. Drive adoption across Sales and Service teams - training, change management and feedback ...

Product Marketing Manager, AI Governance (EMEA)

Location
Greater London, England, United Kingdom
part of our business. Working alongside a team of deeply technical product marketers who own positioning for AI traffic management, agent identity, and AI observability, you'll plan the launches, campaigns, and content that drive pipeline. This is a high-leverage execution role. You'll own launches, content, and campaign … market motions is strongly preferred. Familiarity with the AI connectivity landscape — including awareness of competing and complementary vendors across API gateways, AI observability, vector databases, iPaaS, and agentic orchestration frameworks — is a strong plus. About Kong: Kong Inc., the AI Connectivity Company, is building the connectivity layer of AI. Trusted ...

Global Account Manager

Location
Maidenhead, England, United Kingdom
business and the customer in a trusted advisor/consultative approach; and establish credibility quickly with senior-level executives across the organizations You understand Observability, Security and/or technology-as-a-service space (Infrastructure-as-a-Service, Software-as-a-Service, Platform-as-a-Service) You can bring your … have knowledge of the IT ecosystem in large, global, enterprises Why you will love being a Dynatracer Dynatrace is a leader in unified observability and security. We provide a culture of excellence with competitive compensation packages designed to recognize and reward performance. Our employees work with the largest cloud providers ...

AI Large Language Model (LLM) Technology Lead/Principal Architect

Location
Greater London, England, United Kingdom
work of multiple domain architects and subject matter specialists - across areas such as agentic application design, AI security and trust, AI operations and observability, data and knowledge engineering, and model platforms and inference - providing the architectural vision, technical governance, and cross-domain coherence that binds their contributions into a unified … design direction for context assembly and memory that manages prompts, context windows, and conversational state across the platform Be accountable for security, governance, observability, performance, and scalability addressed holistically and consistently across every domain Establish the identity and authorization model - per-agent identity, IAM/IAP binding, and defense ...

Lead Software Engineer (£80k + benefits)

Location
Wigan, England, United Kingdom
production. As a Lead Software Engineer you’d play a key role in system design, helping to modernise their existing microservices and improve observability and testability by using modern approaches like hexagonal architecture and agentic behaviour driven design. Skills: The money is good too – up to £80k plus benefits including ...

Manager - Banking Platform Architecture, TC FS

Location
Greater London, England, United Kingdom
reusable platform services to accelerate delivery with control. Your key responsibilities Define platform services: API gateway, event mesh, identity/access, reference data, observability and developer tooling. Establish API/event patterns and domain contracts across core, channels, ledger, risk and data platforms. Contribute to technology strategy (build/… microservices; understanding of cloud platform services. Experience defining patterns/guardrails and working with multiple delivery teams in parallel. Knowledge of identity, secrets, observability and operational controls. Strong collaboration with architects, product owners and engineering leadership. Ideally, you’ll also have Exposure to FinOps, cost attribution and performance engineering. ...

Senior UX Design Manager, Cloud

Location
Greater London, England, United Kingdom
else will follow.” Responsibilities Drive the development and adoption of a modular, and API-first Continuous Delivery ecosystem that unifies policy enforcement, observability, and infrastructure orchestration across all Google environments to accelerate developer velocity, supporting Google's CD tools that handle 154+ million deployments annually across Alphabet. Transform how production … agents which will forecast and mitigate potential outages before they ever manifest. Deliver a simple, intelligent, and comprehensive platform that ensures unified and reliable observability across Alphabet by leveraging AI/ML for proactive insights and automated solutions, while providing personalized, adaptive user experiences. Build Cloud’s unified, enterprise-grade ...

Technology Support Director, Risk Engagement

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
cross-domain incident, problem and change management challenges, identifying process weaknesses, recurring failure patterns and opportunities to improve operational resilience across Infrastructure Platforms. Uses observability, monitoring and diagnostic information to analyse significant, recurring or cross-domain issues, test root-cause hypotheses and help accountable teams define sustainable remediation. Leads focused … analysing complex application or infrastructure issues in large-scale on-premises and public-cloud environments, including dependencies across multiple technology domains. Proficient in using observability, monitoring and diagnostic information to identify patterns, test hypotheses and support root-cause analysis across distributed technology environments. Experience working in a client-facing technology ...