76 to 100 of 101 Observability Jobs in the Thames Valley

Senior Site Reliability Engineer

Location
Milton Keynes, England, United Kingdom
experience with both Azure, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways … Engineer Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles. Regularly use Datadog and other observability tools for application performance monitoring. Implement new ways of working, helping to shape how the organisation responds and recovers to incidents. Take ownership of incident ...

Senior Site Reliability Engineer

Hiring Organisation
VIQU IT Recruitment
Location
Milton Keynes, Buckinghamshire, United Kingdom
Employment Type
Full-Time
Salary
£65,000 - £75,000 per annum
experience with both Azure, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways … Engineer Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles. Regularly use Datadog and other observability tools for application performance monitoring. Implement new ways of working, helping to shape how the organisation responds and recovers to incidents. Take ownership of incident ...

AI Platform Engineer

Location
Reading, England, United Kingdom
across the business adopt AI solutions safely and effectively. This is a hands-on engineering role that combines feature delivery with responsibility for resilience, observability, governance and measurable outcomes. You’ll work closely with Product Owners, Architects, Engineers, Security teams and business stakeholders to deliver secure, scalable and reliable platform … applications, APIs, services or platforms. Experience using logs, metrics, traces and telemetry to improve platform performance and reliability. Experience designing solutions with resilience, scalability, observability and operability in mind. Evidence of improving the dependability of production systems through practical engineering changes. A track record of delivering technology capabilities into production ...

SRE: Cloud Reliability, DevSecOps & Observability — Hybrid London

Location
Slough, England, United Kingdom
will apply software engineering principles to automate, scale, and secure cloud-native environments. Responsibilities include building and maintaining production and demo environments, implementing observability with Prometheus, Grafana, and Loki, and guiding project teams in DevSecOps practices. Hybrid London model, SC level clearance may be required. #J-18808-Ljbffr ...

Fullstack Engineer

Location
Bracknell, England, United Kingdom
Design and maintain scalable Go-based microservices Build and support REST and gRPC APIs Develop integrations and event-driven solutions across distributed systems Improve observability, reliability, scalability, and security Deploy and support applications in AWS Work with Docker, Kubernetes, and CI/CD pipelines Participate in production support … Nice to have: AWS experience Nice to have: Kubernetes and container orchestration Nice to have: Event-driven architectures and messaging platforms Nice to have: Observability, monitoring, and distributed tracing Nice to have: Experience working in a SaaS product organisation Core Competencies Demonstrates expertise in building modern web applications using React ...

Observability Engineer: Metrics, Logs & UX

Location
Oxford, England, United Kingdom
Oxford Nanopore Technologies is seeking a Software Engineer with an observability focus to join the Operational Software Engineering team in Oxford. You will build and improve observability systems across on-prem VMs, HPC, Kubernetes and AWS Elastic Container Service, supporting R&D, Tech Transfer and Manufacturing operations. The role emphasizes ...

Site Reliability Engineer

Location
Milton Keynes, England, United Kingdom
hands‐on role in ensuring it is reliable, scalable, and observable. You will help establish and mature SRE practices, focusing on: Monitoring and observability Incident response Post‐incident review Reliability testing and capacity planning Toil reduction Enabling development velocity We offer a hybrid working arrangement with one day per week … Build dashboards, alerts, and runbooks to improve visibility Automate repetitive tasks to reduce operational toil Collaborate with cross‐functional teams to enhance reliability and observability Support performance testing and capacity planning Proactively identify and prioritise reliability improvements Experience & Skills Required Hands‐on experience with Azure Monitoring (Application Insights, Alerts, Action ...

Data Platform DevOps Analyst

Location
Reading, England, United Kingdom
journey supporting the migration from Teradata to Databricks, validating production readiness, improving monitoring and automation capabilities and helping embed modern DataOps practices across deployment, observability and support processes. What You’ll Bring Here at Primark, we want everyone to feel valued – so please bring your authentic self to work … platforms and DataOps practices with exposure to Azure DevOps, Databricks, dbt Cloud, Azure data services, SQL, Python, ETL/ELT processes, job scheduling, monitoring, observability and deployment automation. Experience supporting production environments and service operations including monitoring, alerting, incident management, ticketing systems, release management, environment management, change governance and enterprise ...

Senior Data Engineer

Location
Maidenhead, England, United Kingdom
data solutions are reliable, scalable, performant, secure, and production‐ready Monitor, troubleshoot, and continuously improve pipeline performance, data quality, and platform stability Drive automation, observability, and supportability across data, analytics, and AI/ML solutions Our Ideal Candidate Strong data engineering experience with hands‐on delivery of scalable data pipelines … productionization, LLM‐based applications, or agentic AI patterns will be an added advantage Experience with DevOps and DataOps practices, including CI/CD, monitoring, observability, and incident support Maersk is committed to a diverse and inclusive workplace, and we embrace different styles of thinking. Maersk is an equal opportunities employer ...

Senior AI Engineer

Hiring Organisation
MarkIT Placements
Location
Didcot, Oxfordshire, South East, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
From £700 to £900 per day
predictable failure behaviour. Deploy AI systems across cloud and on-premises environments, with an understanding of the constraints associated with each. Build evaluation and observability capabilities to measure model performance, agent behaviour and system reliability. Take end-to-end ownership of technical workstreams, from architecture and implementation through to deployment. … would be advantageous: Multimodal AI and reasoning. Edge or offline AI deployments. Kubernetes, particularly EKS or OpenShift. MLOps, including model evaluation, monitoring and reproducibility. Observability for agentic AI systems, including model performance, agent behaviour and drift. Agent orchestration and inter-agent communication protocols such as A2A. Model Context Protocol ...

Platform Engineering Manager (SRE)

Location
Bracknell, England, United Kingdom
accelerate software delivery and operational performance. Build and maintain core platform capabilities: CI/CD pipelines, build and release automation, infrastructure automation, environment management, observability, and developer tooling. Drive standardisation and reduce engineering friction through automation and self-service. Site Reliability Engineering Introduce and embed SRE practices across Engineering. Improve … scale. A background built in cloud product or SaaS companies, where reliability is something customers feel directly. Deep, hands-on SRE expertise: monitoring, observability, alerting, incident management, and operational readiness. You can define what good looks like and stand up foundational SRE practices, from SLOs and error budgets to public ...

Live Quantum Observability Engineer

Location
Reading, England, United Kingdom
leading quantum computing company, is seeking a Monitoring and Observability Engineer to keep our live cryogenic systems performing at peak reliability. You’ll build dashboards, refine alerts, and automate responses to maximise uptime, collaborating across engineering, operations and reliability teams. Ideal candidates have experience in monitoring, dashboards, Python development … call incident management, with a passion for observability and continuous improvement in complex #J-18808-Ljbffr ...

Senior AI Engineer

Location
Maidenhead, England, United Kingdom
other teams build against. Build evaluation frameworks and developer tooling robust enough for production yet simple enough for non‐specialist developers to adopt. Establish observability standards for AI systems – quality, performance, cost, and regression signals – and build dashboards and reporting that turn those signals into actionable decisions. Drive engineering rigor … machine translation, or content‐generation systems, including metrics such as COMET, chrF++, BLEU, MetricX, and MQM‐style human evaluation. Experience with experimentation and observability tooling, data/test‐set versioning, and rigorous benchmarking workflows. Established practice in AI governance and documentation - model cards, system cards, reproducibility, and responsible‐AI considerations ...

Data Platform Engineer

Hiring Organisation
Hays
Location
Milton Keynes, Buckinghamshire, UK
Employment Type
Full-time
Reference: 4819090Job ID: 5408590Posted: 2026-08-05Closing date: 2026-11-02Location: Milton Keynes, (Hybrid - 1 day a week on site)Salary: 55 - 65k (DOE) (per annum)Job type: PermanentWorking pattern: Full timeIndustry: Property ...

Strategic Enterprise Account Executive - Public Sector

Location
Maidenhead, England, United Kingdom
executive leader for named strategic accounts, building trusted relationships with senior business and technology stakeholders and driving adoption of Dynatrace's observability, security and AI-powered platform capabilities. Success in this role comes from understanding how large government organisations operate, navigating complex stakeholder landscapes, influencing major transformation programmes and orchestrating … within assigned accounts. Driving Adoption and Expansion Increase adoption of Dynatrace across departments, programmes and technology domains. Identify opportunities to expand platform usage across observability, application security, infrastructure monitoring, logs, cloud and AI-powered capabilities. Help customers realise greater business value from existing investments while identifying new strategic growth opportunities. ...

Senior Connectivity Engineer

Hiring Organisation
Hackajob Ltd
Location
Wallingford, Oxfordshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
WireGuard/Tailscale or equivalent): access-as-code, policy patterns, posture/health automation, and resilience/disaster recovery planning. Deliver fleet-wide connectivity observability: monitoring, alerting, reporting, and actionable signals that help teams diagnose end-to-end issues quickly. Improve cellular/SIM lifecycle management: provisioning automation, usage/…/PMTUD, conntrack, nftables/iptables) and diagnosing kernel-level networking behaviour. Proficient in Go and/or Python and experienced with modern observability tooling; bonus points for containers/IoT OS, ACL-as-code patterns, and carrier/router API integrations. Benefits Starting from the interview process and continuing ...

Senior Software Engineer, Non-Realtime Controllers

Location
Oxford, England, United Kingdom
common to many controllers rather than specific to one. Work closely with the systems teams and scientists who depend on these services, and strengthen observability, health checking and alerting. Raise the standard of the code around you through review, and bring less experienced engineers on by working alongside them. Requirements … Rust codebase, including async programming. Experience designing and operating networked services such as gRPC or REST, with a real grasp of error handling, observability and failure modes. Experience of hardware and instrument protocols and interfaces such as SCPI, Modbus, serial and Ethernet. Comfortable owning the production behaviour of what ...

Lead ClickHouse Solutions Architect Greenfield Observability

Location
Slough, England, United Kingdom
Colehouse Group is seeking a ClickHouse Solutions Architect to lead a greenfield, enterprise-scale observability platform for a global banking client in Slough. You will shape architecture from first principles to meet high throughput, multi-tenant isolation and storage cost constraints. You'll design data models, ingestion and retention strategies ...

Fleet Connectivity Engineer — Automation & Observability

Location
Wallingford, England, United Kingdom
/or Python, and possess experience in robotics or connected devices. Responsibilities include evolving the connectivity stack, building self-serve workflows, and ensuring observability across the fleet. #J-18808-Ljbffr ...

Senior Connectivity Engineer / Network Engineer

Location
Wallingford, England, United Kingdom
WireGuard/Tailscale or equivalent): access‐as‐code, policy patterns, posture/health automation, and resilience/disaster recovery planning. Deliver fleet‐wide connectivity observability: monitoring, alerting, reporting, and actionable signals that help teams diagnose end‐to‐end issues quickly. Improve cellular/SIM lifecycle management: provisioning automation, usage/…/PMTUD, conntrack, nftables/iptables) and diagnosing kernel‐level networking behaviour. Proficient in Go and/or Python and experienced with modern observability tooling; bonus points for containers/IoT OS, ACL‐as‐code patterns, and carrier/router API integrations. #J-18808-Ljbffr ...

Staff Site Reliability Engineer

Hiring Organisation
Genomics
Location
Oxford, Oxfordshire, UK
Employment Type
Full-time
multi-trillion-row, petabyte scale — sharding and replication, materialised views, merge and query optimisation, tenant isolation and cost/performance trade-offs. Owning SLOs, observability, capacity planning and incident response for data-intensive systems and pipelines, alongside the orchestration and job execution that power them. Shaping Data-as-a-Service … level rather than as a black box — and ideally have contributed code upstream. Reliability engineering for data platforms. You bring true SRE discipline — SLOs, observability, capacity planning and incident response — to analytical data systems and pipelines. Data-as-a-Service productisation. You think in terms of data as a product ...

Strategic Public Sector Account Exec (Observability & AI)

Location
Maidenhead, England, United Kingdom
Account Executive to lead a portfolio of 2-3 UK Public Sector customers and a select set of prospects. You will drive adoption across observability, security and AI-powered platform capabilities while managing executive-level relationships. You will orchestrate cross-functional teams, navigate complex procurement environments and partner with large ...

Clickhouse Solutions Architect

Location
Slough, England, United Kingdom
Role We are looking for a ClickHouse Solutions Architect to join our team supporting the design and implementation of a greenfield, enterprise-scale ClickHouse observability platform for a global banking client. This is a genuine greenfield build at significant scale — there is no incumbent platform to inherit or work around. … where benchmarks disprove the design Establish infrastructure-as-code, CI/CD and environment promotion for schema and configuration changes Productionisation Define and implement observability — system table monitoring, metrics, alerting thresholds, capacity headroom tracking Establish backup, restore and disaster recovery, and validate them by test Implement security and governance — RBAC ...

Global Account Manager

Location
Maidenhead, England, United Kingdom
business and the customer in a trusted advisor/consultative approach; and establish credibility quickly with senior-level executives across the organizations You understand Observability, Security and/or technology-as-a-service space (Infrastructure-as-a-Service, Software-as-a-Service, Platform-as-a-Service) You can bring your … have knowledge of the IT ecosystem in large, global, enterprises Why you will love being a Dynatracer Dynatrace is a leader in unified observability and security. We provide a culture of excellence with competitive compensation packages designed to recognize and reward performance. Our employees work with the largest cloud providers ...

Software Engineer - Life AI Platform AI & Robotics Oxford, England, United Kingdom

Location
Oxford, England, United Kingdom
hand a ticket to at the beginning. Build and operate the technical foundations for agentic systems, including orchestration, tool interfaces, state management, evaluation, observability and failure recovery. Design evaluative frameworks to understand if the deployed model is working Work directly with scientists, observe how they operate and translate what … have dealt with agentsand their typical failure modes. You’ve designed andbuiltproduction-readysystems.You can make sound decisions about APIs, data models, testing, deployment, observability and failure recovery. You’ve built policy-aware systems with durable audit trails.Constraints are enforced by design, and every decision leaves a trustworthy, queryable record. ...