3,051 to 3,075 of 4,448 Observability Jobs

Software Developer - Data Reliability

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
platforms, and automation that our technology and investment teams depend on. The team operates at the intersection of engineering and operations. We build the observability tooling that surfaces problems before they become incidents, the job orchestration platform that runs production workloads at scale, the self-serve systems that let teams … with a global team and external vendorsMindset: Proactive, detail-oriented, and self-driven with a strong sense of ownership and accountabilityNice to haveExperience with observability and monitoring tools such as Grafana, Kibana, or PrometheusExperience developing automation tooling and implementing configuration managementExperience with cloud platforms such as Google Cloud or AWSExperience ...

Agentic AI Engineer - Rapid Prototyping

Hiring Organisation
PCR Recruitment Limited
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
Up to £500 per day + Inside IR35
Kernel, AutoGen, CrewAI or similar RAG, tool/function calling and grounding LLMs against enterprise data Prompt and agent instruction design LLM evaluation, guardrails, observability and security Strong Python and SQL REST APIs, JSON, webhooks and enterprise integrations Querying and manipulating data for AI use cases Rapid prototyping and taking … role is exploring emerging AI technology outside the existing stack - from new agent frameworks and foundation models to AI coding, automation, evaluation and observability tools. We're looking for someone who enjoys experimenting with new technology but can apply it to real enterprise problems. Build quickly, test with users, measure ...

Senior Engineering Manager (GenAI & Agentic Platforms)

Location
Greater London, England, United Kingdom
Agentic Platform: approximately 8-12 engineers across two teams, while partnering closely with ML Platform leadership. GenAI Platform owns shared model access, routing, evaluation, observability, prompt and configuration lifecycle, retrieval, grounding, and guardrails. Agentic Platform builds on those foundations with durable execution, tools, state, context, memory, permissions, human controls, agent … coaching. Own hiring quality, team composition, evolving team boundaries, and operating models as the platforms and company change. Guide architecture across model access, evaluation, observability, retrieval, agent orchestration, tools, state, memory, permissions and control planes. Translate strategy into realistic plans, balancing foundational investment with near-term product needs while surfacing ...

Senior Database Administrator

Hiring Organisation
N P Associates
Location
London, United Kingdom
Employment Type
Full-Time
Salary
£80,000 - £95,000 per annum
Time-Series: QuestDB (Used for massive market data ingestion & operational metrics) Version-Controlled SQL: Dolt (Git-for-data branching technology) Caching & Search: Redis & Elasticsearch Observability: Grafana & Linux-based monitoring tools Key Responsibilities Estate Ownership: Administer, monitor, and scale the database ecosystem to ensure continuous high availability and reliability. Developer Collaboration … Partner closely with engineering teams on schema design, query optimization, and database access patterns. Build Observability: Write complex SQL queries to surface vital business and operational metrics onto Grafana dashboards. Infrastructure Resilience: Plan and execute seamless upgrades, patch management, and routinely test point-in-time recovery (RPO/RTO) strategies. ...

Data Consultant

Hiring Organisation
Lynx Recruitment Ltd
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£40,000 - £50,000 per annum
/or analytics Client-facing or consulting experience Strong SQL and Python skills Experience with AWS, GCP or Azure Understanding of data quality, governance, observability and dimensional modelling ...

Head of Data Science

Location
United Kingdom
model registry; inference services with real latency, availability and rollback requirements; historically accurate features where decisions need backdated reconstruction; eval harnesses and trace‐observability for agentic workflows; and monitoring for data quality, drift, model performance and economic outcomes. Start ad hoc where that's sufficient and platformise once the first … uplift models, feature pipelines and drift on one side; tool and context design, prompt and retrieval iteration, evals against golden answer sets and trace‐observability on the other. Practised at running live systems rather than just launching them — monitoring, retraining, incident response, rollback, and the on‐call reality ...

Test Engineer

Hiring Organisation
SRG
Location
Manchester, Lancashire, United Kingdom
Employment Type
Full-Time
Salary
£40,000 - £45,000 per annum
level, identifying risks early, and carrying out exploratory testing to uncover issues that scripted and automated tests may miss. You'll also contribute to observability-led quality, using monitoring, metrics and logging to understand system reliability, performance and user experience across live environments. The role will suit someone who enjoys … testing pyramid. Carry out exploratory testing to uncover defects and validate user journeys. Perform backend testing across APIs, databases and integrated systems. Contribute to observability-led quality through monitoring, metrics, logging and alerting. Collaborate closely with Developers, Product Owners, UX teams and Technical Leads. Identify, communicate and manage quality risks ...

Gen AI Architect

Location
Greater London, England, United Kingdom
production-grade AI systems using Amazon Bedrock, retrieval-augmented generation (RAG), agentic workflows, and cloud-native AWS services. Drive architecture standards, model orchestration, governance, observability, and operational excellence across the GenAI lifecycle while collaborating with engineering, security, compliance, and business stakeholders**Hybrid working:**The places that you work from … customization, prompt orchestration, retrieval pipelines, and agentic workflows* Design agentic AI systems incorporating tool use, workflow orchestration, memory management, and autonomous decision flows* Implement observability for prompts, model responses, vector retrieval quality, and agent execution workflows* Integrate GenAI capabilities into enterprise applications, APIs, workflow platforms, and data ecosystems* Work with ...

Senior Software Development Engineer

Location
City Of London, England, United Kingdom
expectations and can be reused across multiple brands and platforms. Drive engineering excellence for the services you own by championing code quality, automated testing, observability, performance optimization, and simplification, taking technical responsibility for service health, scalability, resilience, and the ongoing reduction of technical debt and operational overhead. Provide technical mentorship … integrations, and communicating trade-offs to both technical and non-technical stakeholders. Track record of improving operational excellence at the team level through enhanced observability, automation, performance tuning, and data-driven analysis of incidents and customer impact. Hands-on experience integrating or consuming AI/ML-enabled services or platforms ...

Data Platform SRE

Location
Greater London, England, United Kingdom
wider business depends on for analytics, reporting, and product features. This is a hands‐on engineering role: you'll write infrastructure‐as‐code, build observability into data systems from the ground up, lead incident response for data platform outages, and work directly with data engineers to raise the reliability … maintain the infrastructure that underpins the data platform (compute, storage, orchestration, networking) using infrastructure-as-code primarily focussed in Microsoft Azure. Build and improve observability for data systems — metrics, logging, tracing, and data‐quality/freshness monitoring — so issues are caught before they reach downstream consumers. Lead incident response ...

Remote Sr. Software Engineer, Fullstack (UK)

Hiring Organisation
First Up
Location
Remote, UK
industry standards and best practices to help others solve complex problems. Identify, analyse, and resolve system availability, reliability, and performance issues, contributing to observability and resiliency improvements. Bachelor's degree in Computer Science or related field, or equivalent experience. ~5+ years building reliable, performant applications and microservices. ~ Strong proficiency … Experience building and maintaining integrations with HRIS platforms such as Workday, SAP SuccessFactors, Dayforce, or similar enterprise HR systems. Experience with Datadog or similar observability and monitoring platforms. Demonstrated use of Generative AI tools or coding agents in development workflows. Familiarity with ML/AI integration in production systems. Open ...

Remote Sr. Software Engineer, Fullstack (UK)

Hiring Organisation
First Up
Location
Wrexham, Wales, UK
industry standards and best practices to help others solve complex problems. Identify, analyse, and resolve system availability, reliability, and performance issues, contributing to observability and resiliency improvements. Bachelor's degree in Computer Science or related field, or equivalent experience. ~5+ years building reliable, performant applications and microservices. ~ Strong proficiency … Experience building and maintaining integrations with HRIS platforms such as Workday, SAP SuccessFactors, Dayforce, or similar enterprise HR systems. Experience with Datadog or similar observability and monitoring platforms. Demonstrated use of Generative AI tools or coding agents in development workflows. Familiarity with ML/AI integration in production systems. Open ...

Remote Sr. Software Engineer, Fullstack (UK)

Hiring Organisation
First Up
Location
Ipswich, Suffolk, UK
industry standards and best practices to help others solve complex problems. Identify, analyse, and resolve system availability, reliability, and performance issues, contributing to observability and resiliency improvements. Bachelor's degree in Computer Science or related field, or equivalent experience. ~5+ years building reliable, performant applications and microservices. ~ Strong proficiency … Experience building and maintaining integrations with HRIS platforms such as Workday, SAP SuccessFactors, Dayforce, or similar enterprise HR systems. Experience with Datadog or similar observability and monitoring platforms. Demonstrated use of Generative AI tools or coding agents in development workflows. Familiarity with ML/AI integration in production systems. Open ...

Remote Sr. Software Engineer, Fullstack (UK)

Hiring Organisation
First Up
Location
Braintree, Essex, UK
industry standards and best practices to help others solve complex problems. Identify, analyse, and resolve system availability, reliability, and performance issues, contributing to observability and resiliency improvements. Bachelor's degree in Computer Science or related field, or equivalent experience. ~5+ years building reliable, performant applications and microservices. ~ Strong proficiency … Experience building and maintaining integrations with HRIS platforms such as Workday, SAP SuccessFactors, Dayforce, or similar enterprise HR systems. Experience with Datadog or similar observability and monitoring platforms. Demonstrated use of Generative AI tools or coding agents in development workflows. Familiarity with ML/AI integration in production systems. Open ...

Remote Sr. Software Engineer, Fullstack (UK)

Hiring Organisation
First Up
Location
Ilminster, Somerset, UK
industry standards and best practices to help others solve complex problems. Identify, analyse, and resolve system availability, reliability, and performance issues, contributing to observability and resiliency improvements. Bachelor's degree in Computer Science or related field, or equivalent experience. ~5+ years building reliable, performant applications and microservices. ~ Strong proficiency … Experience building and maintaining integrations with HRIS platforms such as Workday, SAP SuccessFactors, Dayforce, or similar enterprise HR systems. Experience with Datadog or similar observability and monitoring platforms. Demonstrated use of Generative AI tools or coding agents in development workflows. Familiarity with ML/AI integration in production systems. Open ...

Remote Sr. Software Engineer, Fullstack (UK)

Location
Mansfield, Nottinghamshire, United Kingdom
industry standards and best practices to help others solve complex problems. Identify, analyse, and resolve system availability, reliability, and performance issues, contributing to observability and resiliency improvements. Bachelor's degree in Computer Science or related field, or equivalent experience. 5+ years building reliable, performant applications and microservices. Strong proficiency … Experience building and maintaining integrations with HRIS platforms such as Workday, SAP SuccessFactors, Dayforce, or similar enterprise HR systems. Experience with Datadog or similar observability and monitoring platforms. Demonstrated use of Generative AI tools or coding agents in development workflows. Familiarity with ML/AI integration in production systems. Open ...

Remote Sr. Software Engineer, Fullstack (UK)

Hiring Organisation
First Up
Location
Sunderland, Tyne and Wear, UK
industry standards and best practices to help others solve complex problems. Identify, analyse, and resolve system availability, reliability, and performance issues, contributing to observability and resiliency improvements. Bachelor's degree in Computer Science or related field, or equivalent experience. ~5+ years building reliable, performant applications and microservices. ~ Strong proficiency … Experience building and maintaining integrations with HRIS platforms such as Workday, SAP SuccessFactors, Dayforce, or similar enterprise HR systems. Experience with Datadog or similar observability and monitoring platforms. Demonstrated use of Generative AI tools or coding agents in development workflows. Familiarity with ML/AI integration in production systems. Open ...

Integration Engineering Manager

Location
Greater London, England, United Kingdom
Agent-ready API products, reusable integration assets, and governed API catalogs AI-enabled integration patterns across Salesforce, Data, ERP, customer, operational, and platform domains Observability, policy enforcement, auditability, and compliance for AI-driven integration flows The manager will oversee a broad portfolio of APIs, connectors, MCP-enabled tools, and integration … routing, policy enforcement, and agent-to-system orchestration. Ensure AI-enabled integrations follow strong standards for identity, consent, access control, data minimisation, rate limiting, observability, and auditability. Provide architectural guidance on complex use cases involving multi-system orchestration, agent workflows, human-in-the-loop controls, and event-driven automation. Ensure ...

QA Engineer

Location
Manchester, England, United Kingdom
left testing by considering testability early in designs and features and helping the team validate functionality before it reaches production. You will contribute to observability-led quality, using monitoring, metrics, and logging to understand system reliability, performance, and user experience in production. By leveraging these insights, you will help … testability early in designs and features. Expertise in performing exploratory testing to uncover defects and validate real-world usage scenarios. Expertise in contributing to observability-led quality, leveraging monitoring, metrics, logging, and alerts to understand system reliability and performance. Experience of identifying and reporting risks, issues, and defects to support ...

Principal Product Engineer

Location
Greater London, England, United Kingdom
typed APIs, clear service and data boundaries, robust processing workflows and platform capabilities that can handle millions of records and events without compromising correctness, observability or operability. A key part of the role is separating the data layer from the application layer. You will help ensure an action taken … propagation separately, ensuring that customer-facing actions can be reversed safely without creating hidden inconsistency underneath. Reliability Engineering : Improve idempotency, retry behaviour, failure isolation, observability, alerting and recovery across critical workflows and integrations. Technical Leadership : Lead design reviews, mentor through code review and pairing, make technical standards explicit, and help ...

Senior ServiceNow Developer

Location
Leeds, England, United Kingdom
service operations tooling, ensuring integration and automation opportunities are maximised. Stay current with industry best practices and emerging technologies in service operations tooling, including observability and alerting. Drive innovation by integrating and optimising monitoring solutions within service management workflows. Develop proof-of-concept (POC) solutions for new features and capabilities … across observability and ITSM platforms. Lead cross-team initiatives to drive innovation, with the potential for reuse across multiple customer environments. Service Operations Reporting & Insights: Design and implement reporting solutions to provide real-time insights into service performance, incident trends, and operational health. Utilise data analytics and explore machine learning ...

Staff Backend Software Engineer - Python & GO

Location
Greater London, England, United Kingdom
create, train, and deploy physics-informed models at PhysicsX. You'll take ownership of your work from implementation to production, ensuring quality, scalability, and observability at every step. By engaging with our Guilds and leveraging domain knowledge from end users, you’ll not only refine your craft but also help … Python and Go. A willingness to take ownership from implementation to production, including testing, containerisation, continuous integration and delivery, authentication/authorisation, telemetry/observability/monitoring A working understanding of messaging in event-driven systems, which implies some experience using tools such as NATS, RabbitMQ, or Kafka for example. ...

Lead Software Engineer - Agent Safety

Location
Greater London, England, United Kingdom
enforce pre- and post-generation guardrails, managing the overarching governance of AI models operating within internal tools and platforms. Drive AI security and observability: Build out dedicated auth/permissions for internal AI agents, establish deep monitoring/observability pipelines, and define incident response protocols for AI-specific anomalies. Verification … Inspect AI, Ragas, OpenAI Evals, NeMo Guardrails). Django experience and strong backend engineering patterns (security, performance, maintainability). Experience with Datadog for complex observability, tracing, and monitoring in AI environments. Familiarity with foundational AI engineering tooling like Pydantic AI, LiteLLM, or LangChain. Are you ready for a career with ...

Senior Software Engineer, Banking Connectivity London, UK

Location
Greater London, England, United Kingdom
that high-integrity financial data is correctly distributed across internal systems. Your work will focus on scaling integrations while improving the system’s resilience, observability, and overall structure. You will play a key role in evolving the platform to support new banking partners, products, and regulatory requirements while addressing technical … real‐world banking constraints Collaborate with product, operations, and external partners to unblock integrations and accelerate delivery Improve system quality through pragmatic enhancements in observability, testing, and resilience. This is a high‐impact role. What You'll Bring Experience building and supporting reliable backend systems with external integrations (APIs, webhooks ...

Product Engineer (Backend/AI @Briink)

Location
Greater London, England, United Kingdom
that support multiple user journeys and reporting workflows rather than solving each problem in isolation Improve the robustness of our systems, strengthening reliability, testing, observability, performance, and maintainability as we scale Work directly with users, product, and data, using qualitative feedback and product data to understand the real problem … combined with the ability and willingness to become productive in Python quickly) A solid software-quality mindset, including testing, code review, CI/CD, observability, and pragmatic approaches to reliability and security Experience in a small, high-ownership product team, ideally within a startup or scale-up engineering organisation ...