3,076 to 3,100 of 4,448 Observability Jobs

Senior Software Engineer

Location
Greater London, England, United Kingdom
TypeScript experience. Solid AWS and Serverless experience. Experience building distributed, resilient APIs. Experience with MongoDB or another document database. A strong understanding of testing, observability and performance. Experience working in financial services, fintech, startups or another regulated environment. Experience mentoring engineers and taking ownership of production services. ...

DevOps Engineer - ELK

Hiring Organisation
ECS Resource Group Ltd
Location
City of London, London, United Kingdom
Employment Type
Contract
Contract Rate
£250 - £400/day
DevOps Engineer, you will be responsible for: Designing, deploying, and maintaining Elasticsearch clusters. Monitoring platform performance and resolving technical issues. Supporting logging, monitoring, and observability solutions. Automating operational processes through scripting and tooling. Working closely with engineering teams to improve system reliability and scalability. Supporting cloud-native and containerised environments. ...

Network Automation Engineer

Hiring Organisation
Project Recruit
Location
Manchester, Lancashire, United Kingdom
Employment Type
Contract
Contract Rate
GBP Annual
diagnostic workflows for incidents and monitored events, with clear guardrails, approvals and exception handling. Integrate NetBrain with ServiceNow, DX NetOps Spectrum and other monitoring, observability and orchestration platforms using supported APIs and integration patterns. Automate collection of device state, topology, paths, configurations, logs and relevant telemetry, and make outputs available … across multi-vendor network estates such as Cisco, Juniper, Arista, Fortinet, Palo Alto Networks or F5. Experience integrating network platforms with ITSM, monitoring or observability tools, preferably ServiceNow and DX NetOps Spectrum. Ability to interpret alarms, telemetry, configurations, routing tables, logs and packet or path information to develop reliable diagnostic ...

Senior Software Engineer

Location
Bath, England, United Kingdom
design scalable, resilient systems and take ownership beyond simply delivering tickets. You will take ownership across system design, architecture, CI/CD, observability, incident response and production systems, working closely with Engineering, Product and Operations teams to build and continuously improve reliable platforms. Required Skills: Strong backend engineering experience with ...

Senior Product Manager - Storage & Networking

Location
Greater London, England, United Kingdom
underpinning Radiant’s GPU platform. You’ll work across bare-metal GPU clusters, Kubernetes, high-performance storage, data‐centre networking, infrastructure inventory, automation, and observability, partnering closely with engineering, SRE, infrastructure, networking, and operations. This is a highly technical product role focused on ensuring GPU workloads have reliable, high‐throughput … infrastructure through to customer and workload connectivity. Partner with engineering to turn requirements into scalable, reliable platform services. Drive improvements in automation, self‐service, observability, and operational efficiency. Define success metrics and use them to guide performance, reliability, capacity, and investment decisions. Qualifications: Experience owning technical infrastructure, platform, storage, networking ...

Model Release Engineer London, United Kingdom

Location
Greater London, England, United Kingdom
with teams across Wayve to understand their requirements, agree interfaces and resolve technical or delivery conflicts across shared workflows. Improve the reliability, scalability and observability of the platform through effective monitoring, alerting and operational tooling. Provide clear visibility of model candidates, their progress, evaluation results, approvals and release status. … work through conflicting priorities. Experience operating cloud-based services using Kubernetes, with a good understanding of reliability, scalability and performance. Practical knowledge of observability, monitoring and alerting, including defining meaningful service-health metrics. Confidence using AI coding tools and agents to improve engineering productivity and automate repeatable work. A pragmatic ...

Principal AI Engineer

Location
Greater London, England, United Kingdom
Anaplan’s platform and third‐party integrations Optimise model inference pipelines for performance, cost, and scalability in production environments Implement monitoring, logging, and observability for GenAI systems to track usage, errors, and model behaviour Collaborate with data scientists to productionise ML models and forecasting algorithms Your Skills Extensive hands … Experience with A/B testing and experimentation frameworks for AI features Contributions to open‐source ML projects or research publications Experience with model observability tools (LangSmith, W&B;, MLflow) Our Commitment to Diversity, Equity, Inclusionand Belonging (DEIB) We believe attracting and retaining the best talent and fostering an inclusive ...

Senior Manager - HR AI Solution Architect

Location
Hursley, England, United Kingdom
experiences, use cases, and large‐scale enterprise environments. You will bring an understanding of cloud-based software design, including considerations for sensitive data handling, observability, and audit readiness. You will contribute to the development of best practices, technical assets, and reference architectures (e.g., white papers, code samples, blog posts), while … process automation Data architectures and knowledge management LLM integration patterns: RAG architectures, vector databases, embedding pipelines API/integration patterns across core HCM platforms Observability, monitoring, and governance frameworks AI/agent-based architecture principles and their enterprise implications. AI related and core HR Technology Transformation expertise having worked ...

Senior AI Engineer (Amazon Bedrock)

Location
City of Edinburgh, Scotland, United Kingdom
other sources Apply the A2A (Agent-to-Agent) protocol to enable interoperability between agents across systems and workflows Instrument AI systems with observability and tracing tooling — CloudWatch, spans, and traces — to support debugging, performance monitoring, and compliance requirements Integrate LLMs into client applications through prompt engineering, context management, and function … strong, applied proficiency in an AWS and AI context AWS infrastructure — working knowledge of Lambda, DynamoDB, and S3 as components of AI system backends Observability — experience instrumenting AI systems with tracing, logging, and monitoring tooling (CloudWatch preferred) LLM integration — prompt engineering, tool/function calling, context window management, and output ...

Partner Sales Manager - EMEA

Location
Greater London, England, United Kingdom
March Capital, Lightspeed, Sorenson Ventures, Industry Ventures, and Emergent Ventures, we are a Series-C funded company headquartered in Silicon Valley. Our Enterprise Data Observability Platform - the first of its kind - helps enterprises build and operate world-class data products by ensuring data is reliable, trusted, and ready to power … strategy with hyperscaler priorities and customer modernization initiatives. Product and Industry Expertise and Demonstration Maintain a strong understanding of Acceldata's platform, including data observability use cases across modern and legacy data architectures. Confidently articulate Acceldata's value to partner sales, technical, and executive audiences. Support and, when needed, deliver ...

Solutions Architect

Location
Greater London, England, United Kingdom
Maintain comprehensive documentation of design decisions, patterns, standards, and trade-offs. Define non-functional platform architecture standards covering resilience, backup/restore, disaster recovery, observability, service levels, auditability and operational readiness. Define platform security and privacy architecture in partnership with Cyber Security and Data Governance, including PII handling, access recertification … authority in environments transitioning from outsourced to in-house ownership is beneficial. Experience defining non-functional requirements and architecture patterns for enterprise-grade resilience, observability, disaster recovery, data lifecycle management and operational readiness. Experience using architecture decision records, design authorities, exception management and measurable standards adoption to embed architectural governance ...

Associate Director, Platform Engineering

Location
Greater London, England, United Kingdom
pipelines, internal services and developer tooling, treating the platform as a product and our internal engineers and computational scientists as its customers. Establish comprehensive observability and robust disaster recovery and failover strategies across on-premises and cloud environments. Own Relation’s platform security controls, including IAM, secrets management, network policy … understanding of the reliability, scalability and reproducibility requirements of scientific data pipelines. A track record of maturing CI/CD pipelines, developer tooling and observability across complex, multi-environment platforms. Strong security fundamentals, with hands-on experience designing and implementing appropriate security controls. Excellent written and verbal communication skills, with ...

Lead Infrastructure Engineer - Proxy/SSE Network Security

Location
Greater London, England, United Kingdom
resilience outcomes. Drive operational excellence at scale for perimeter, proxy, and SSE services in the US, including incident, change, and problem management rigor, observability and resiliency validation practices, automation to improve repeatability and evidence quality, reduction of client and partner impact, and execution of Technology Lifecycle Management (TLM) and modernization … design, exception frameworks, audit-ready traceability, and measurable risk reduction reporting. Experience with large-scale operations for externally facing or security enforcement services, including observability strategy, resilience testing, incident response alignment, and reduction of repeat incidents and client-impacting events. Experience designing and operating hybrid edge architectures and cloud interconnect ...

Senior Director, Data and Information Marketplace

Location
Cambridge, England, United Kingdom
ensuring intuitive experiences for both people and agents across discovery, access, sharing, understanding and use. Advise the development of capabilities for data access, lineage, observability and quality so that data assets are transparent, trusted and usable at scale. Shape enterprise approaches to data, information and knowledge lifecycle management, embedding governance … seamless experiences that are widely adopted by users and machines across multiple enterprise business units. Deep expertise in relevant capability areas, including data quality, observability, lineage, access management, lifecycle management and information governance. Measurable evidence of optimising the data P&L across covering cost/FinOps, value realisation and sustainability. ...

Senior Software Engineer (£80k + 15% bonus)

Location
Manchester, England, United Kingdom
product they work on, from ideation through to supporting systems in production, so you should be confident in practices like system design, testing, deployment, observability and monitoring. This position would suit someone looking for a role with a high level of autonomy and the opportunity to work on complex technical ...

Expansion Account Executive

Location
Maidenhead, England, United Kingdom
sell internally across all supporting resources to maximize your effectiveness and advance the sales process (familiar with MEDDPIC). You are familiar with the observability and modern application market.## **Why you will love being a Dynatracer*** Dynatrace is a leader in unified observability and security.* We provide a culture ...

Frontend Engineer

Hiring Organisation
Zopa
Location
London, UK
Employment Type
Full-time
performance, reliability and accessibility, using data, experimentation and customer feedback to continuously improve our products. Participate in code reviews, testing, CI/CD and observability, helping maintain high engineering standards while learning from experienced teammates. Take ownership of your work, proactively identifying opportunities to improve our applications, developer experience … crawlability, structured data, indexation and semantic HTML), with experience optimising applications for both traditional search engines and AI-powered search. Experience with GraphQL and observability tools such as Honeycomb or Grafana. Working on customer-facing products within fintech, regulated industries or fast-growing technology businesses. Contributing to shared libraries, design ...

Senior Backend Engineer

Location
Greater London, England, United Kingdom
execution capabilities to agents Correctness, auditability and data integrity across financial workflows, so material actions remain traceable long after they happen Production ownership : observability, alerting, incident response, performance and scalability for the services you build Shared backend libraries, patterns and platform capabilities , plus technical direction, design and code review standards … data-critical systems across SQL and NoSQL, with a clear sense of APIs, service boundaries and data contracts Designing for correctness, resilience and observability , and building reliable integrations with external systems you do not control Delivery at scale of months. A record of scoping and delivering work measured in months ...

Senior Architect AI

Location
West of England, England, United Kingdom
Dataflow Cloud Functions GKE Pub/Sub IAM and Security Controls MLOps/LLMOps CI/CD for ML Version Control Automated Evaluation Monitoring & Observability Rollback Frameworks Required Experience 14+ years of overall IT experience. 5+ years in AI/ML solution architecture. 3+ years in Generative ...

Senior AI Engineer

Hiring Organisation
Lynx Recruitment Limited
Location
South West London, London, United Kingdom
Employment Type
Permanent
Salary
£80,000
deliver production GenAI and multi-agent systems Build agentic workflows, RAG solutions, tool use and multi-step orchestration Define evaluation strategies and implement AI observability Design secure AWS deployment architectures Build CI/CD pipelines and production-ready AI applications Lead technical discussions with clients and senior stakeholders Mentor engineers ...

Senior AI Engineer

Hiring Organisation
Lynx Recruitment Ltd
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£60,000 - £80,000 per annum
deliver production GenAI and multi-agent systems Build agentic workflows, RAG solutions, tool use and multi-step orchestration Define evaluation strategies and implement AI observability Design secure AWS deployment architectures Build CI/CD pipelines and production-ready AI applications Lead technical discussions with clients and senior stakeholders Mentor engineers ...

Principal Software Engineer (Platforms)

Location
West of England, England, United Kingdom
business: run evaluations, evidence the benefits, and train teams on effective use. Build the infrastructure that makes agents trustworthy: sandboxing, permissioning, audit logging, and observability so that what an agent did is efficient, secure, and fully transparent to the engineer who has to stand behind it. Create repeatable development environments … experience implementing agentic AI in real engineering workflows (beyond individual experimentation). Experience designing the guardrails around agentic AI: sandboxing, permissioning, audit logging, or observability for tools operating with elevated access. Deep, hands‐on experience of CI/CD, infrastructure as code, containers, and Kubernetes. Experience of building and deploying ...

Principal AI Engineer

Location
Greater London, England, United Kingdom
company's platform and third‐party integrations Optimise model inference pipelines for performance, cost, and scalability in production environments Implement monitoring, logging, and observability for GenAI systems to track usage, errors, and model behaviour Collaborate with data scientists to productionise ML models and forecasting algorithms Your Skills Extensive hands … Experience with A/B testing and experimentation frameworks for AI features Contributions to open‐source ML projects or research publications Experience with model observability tools (LangSmith, W& B;, MLflow) Our Commitment to Diversity, Equity, Inclusionand Belonging (DEIB) We believe attracting and retaining the best talent and fostering an inclusive ...

Software Engineer, Simulation

Location
Greater London, England, United Kingdom
differences between on-road and simulated execution, identifying issues across data, inference and simulated components. Improve simulation reproducibility, reliability and debuggability through automated testing, observability and better developer tooling. Profile and improve simulator performance, helping us run increasingly large evaluation workloads efficiently. Work with internal users and adjacent engineering teams … such as camera, radar, lidar or GNSS, including modelling uncertainty or noise. Experience integrating machine-learning inference into production systems. Experience with performance profiling, observability or debugging distributed systems. Familiarity with large-scale batch processing, cloud infrastructure or GPU-based workloads. This is a full-time role based ...

Engineering Manager, ML Infrastructure, London

Location
Greater London, England, United Kingdom
problems that decide whether the platform is fast, affordable, and operable: context and cache management, model asset management and lifecycle, throughput and latency, observability, and the developer and test infrastructure that everything else is built on.You would own set of components in this stack. The generative-AI landscape … build on, including API and compatibility stewardship across versions and hardware generations. Familiarity with production operations for latency-sensitive services: SLOs and error budgets, observability, capacity planning, canary and rollback discipline. Background in developer experience and build or test infrastructure, and a view on how to reduce cycle time without ...