1 to 25 of 26 Observability Jobs in Oxfordshire

Senior Infrastructure Engineer

Hiring Organisation
Hackajob Ltd
Location
Wallingford, Oxfordshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Build infrastructure-as-code (IaC) with Terraform on Terraform Cloud, establishing standards for repeatable environments, reducing manual processes and increasing automation. Support monitoring, logging, observability and alerting systems, via Grafana, to support production health, incident detection, root-cause analysis and rapid recovery. Drive a culture of DevOps best-practice across ...

DevOps Engineer

Location
Wallingford, England, United Kingdom
Build infrastructure-as-code (IaC) with Terraform on Terraform Cloud, establishing standards for repeatable environments, reducing manual processes and increasing automation. Support monitoring, logging, observability and alerting systems, via Grafana, to support production health, incident detection, root-cause analysis and rapid recovery. Drive a culture of DevOps best-practice across ...

Senior Platform Engineer

Location
Abingdon, England, United Kingdom
data visualisation or monitoring dashboards. Knowledge of UI/UX principles and development of web‐based or GUI tools, particularly for internal platforms, observability systems, or scientific workflow interfaces. Exposure to modern or emerging programming languages and ecosystems such as Rust, Julia, or SYCL, and interest in evaluating new tools ...

Embedded DevOps Engineer

Location
Kidlington, England, United Kingdom
environments and automated test rigs. Familiarity with embedded Linux, cross-compilation toolchains, RTOS environments or FPGA development and deployment workflows. Experience with monitoring and observability platforms such asOpenTelemetry, Prometheus, Grafana or Loki. Knowledge of secure software supply-chain practices, including dependency scanning, artifact signing, software bills of materials, secrets management ...

Engineering Manager - Data

Location
Oxford, England, United Kingdom
Lead the transition from traditional DBA operations towards code-defined, automated and rebuildable database environments, including automated schema migrations and immutable infrastructure approaches. Own observability, reliability and cost management across the data estate, establishing meaningful monitoring, alerting and clear service levels for the services teams depend on. Make data quality ...

Senior DevOps Engineer

Hiring Organisation
MarkIT Placements
Location
Didcot, Oxfordshire, South East, United Kingdom
Employment Type
Permanent
image scanning. Implement and monitor infrastructure and application security controls. Support the organisation's ongoing compliance and certification requirements. Reliability & SRE Establish and maintain observability across distributed systems. Develop proactive monitoring, alerting and performance-tuning strategies. Help maintain service-level objectives and platform availability. Investigate and resolve infrastructure and application … advantageous: MLOps or LLMOps experience. Experience with platforms such as SageMaker, Kubeflow or ZenML . Extensive on-premises Kubernetes deployment experience. Prometheus or comparable observability platforms. AWS Karpenter. AWS Compute Optimizer. Experience operating highly distributed systems. Familiarity with ISO 27001, NIST SSDF, OWASP SAMM or similar security frameworks. Understanding ...

Senior Software Engineer (Python)

Location
Oxford, England, United Kingdom
maintainability Working across software, edge infrastructure and integrations Introducing and strengthening modern engineering practices including CI/CD, automated testing, peer review and observability Diagnosing complex issues across software, Linux, networking and infrastructure boundaries Mentoring other engineers and helping shape technical direction as the team grows What we're looking ...

Software Engineer - Orchestration AI & Robotics Oxford, England, United Kingdom

Location
Oxford, England, United Kingdom
turns out to be the wrong answer. The systems you write control real hardware doing real biology, so correctness, recoverability, and observability matter more here than they do in most software work. Key Responsibilities Design and build the software that orchestrates autonomous laboratory systems, including scheduling, workflow execution, state management ...

Software Engineer, Observability

Location
Oxford, England, United Kingdom
Description We are looking for a Software Engineer with an observability focus to join our Operational Software Engineering team. As part of a cross-functional team supporting ONT’s R&D, Tech Transfer and Manufacturing operations, you will build and improve observability systems across multiple applications and deployment environments including … prem VMs, HPC compute and storage systems, Kubernetes and AWS Elastic Container Service. Job Description We are looking for a Software Engineer with an observability focus to join our Operational Software Engineering team. As part of a cross-functional team supporting ONT’s R&D, Tech Transfer and Manufacturing operations ...

Senior Python Engineer - Edge & Autonomous Systems

Location
Oxford, England, United Kingdom
infrastructure, networking, and system integrations, delivering production-grade reliability and scalability. You will mentor engineers, advocate modern engineering practices (CI/CD, automated testing, observability), and help shape the architectural direction as #J-18808-Ljbffr ...

Senior ML Infrastructure Engineer Enterprise Operations Oxford, England, United Kingdom

Location
Oxford, England, United Kingdom
Proactively benchmark, profile, and resolve performance bottlenecks across the compute, network, and orchestration layers to maximise efficiency for distributed training and inference. Establish comprehensive observability, resilience, and automated security controls to ensure compliance and robust operation of sensitive research environments. Partner with Research, Data, and Applied teams to forecast capacity ...

Senior Electronics Engineer

Location
Oxford, England, United Kingdom
development and production hardware . You’ll have the opportunity to drive system upgrades, standardize control and drive electronics across systems, and build the observability that catches problems before they cause downtime , contributing directly to the success of the Reliability Engineering team . Key responsibilities include: Triage & root cause: Actively ...

Artificial Intelligence Engineer

Location
Oxford, England, United Kingdom
solutions in enterprise or regulated environments — aviation, land transport, public safety, telecommunications, or government Familiarity with cloud platforms, APIs, CI/CD, analytics, and observability Experience shaping early-stage products with users and iterating based on feedback Additional information Why Join NCS? Grow with Us Work on cutting-edge ...

Staff Software Engineer, Non-Realtime Controllers

Location
Oxford, England, United Kingdom
codebase, including async programming. Experience building and operating networked services in production such as gRPC or REST, with a real grasp of error handling, observability and failure modes. Experience of hardware and instrument protocols and interfaces such as SCPI, Modbus, serial and Ethernet. A reputation as the person people bring ...

Principal Java Engineer

Location
Wallingford, England, United Kingdom
continuous improvement. Production systems are reliable, observable and operationally excellent Lead root cause analysis and resolution of complex production issues. Drive improvements in system observability, monitoring and operational performance. Ensure applications are designed and operated to meet reliability, availability and performance targets. Partner with Operations, DevOps and QA teams … Claude, Codex, Gitlab Duo, etc) REST APIs, OpenAPI, Microservices, Event-driven architecture (RabbitMQ) Containers, Docker, AWS, Linux CI/CD with GitLab Pipelines & Jenkins Observability: logging, metrics and monitoring MySQL, Apache Solr Front-end UI (e.g. Angular) Person Specification Strategic and systems-thinking mindset Excellent communication and stakeholder management skills ...

Automation & Platform Engineer

Location
Abingdon, England, United Kingdom
services, agent services, APIs, and microservices. Implement infrastructure-as-code for platform environments and network automation resources. Ensure automation platforms meet security, compliance, availability, observability, and operational resilience requirements. Your Profile Experience in network automation, platform engineering, DevOps, cloud engineering, or telecom automation. Strong hands-on experience with Ansible, Terraform … vendor APIs. API management and orchestration. Intent-to-configuration workflows. Data pipelines. Vector databases and graph APIs. MCP integration. Security frameworks and compliance controls. Observability and logging. Preferred Certifications Kubernetes CKA/CKAD. Terraform Associate. Red Hat Ansible certification. Google Cloud, AWS, or Azure certification. Cisco, Juniper, Nokia, or Ericsson ...

Advanced Engineer, Investment Technology

Location
Henley-on-Thames, England, United Kingdom
subject matter experts. Their primary focus is on execution within defined parameters, applying modern engineering practices including CI/CD, automated testing, code quality, observability and data quality controls. They may also be accountable for regular reporting or process administration within the squad. Working in partnership with more experienced staff … business requirements into practical, well-engineered technical solutions. Build production-ready solutions using modern engineering practices, including CI/CD, automated testing, code quality, observability and data quality controls. Support data quality, monitoring, reporting and process administration activities to help ensure accurate and consistent outcomes. Identify and resolve technical problems ...

Observability Engineer: Metrics, Logs & UX

Location
Oxford, England, United Kingdom
Oxford Nanopore Technologies is seeking a Software Engineer with an observability focus to join the Operational Software Engineering team in Oxford. You will build and improve observability systems across on-prem VMs, HPC, Kubernetes and AWS Elastic Container Service, supporting R&D, Tech Transfer and Manufacturing operations. The role emphasizes ...

Senior AI Engineer

Hiring Organisation
MarkIT Placements
Location
Didcot, Oxfordshire, South East, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
From £700 to £900 per day
predictable failure behaviour. Deploy AI systems across cloud and on-premises environments, with an understanding of the constraints associated with each. Build evaluation and observability capabilities to measure model performance, agent behaviour and system reliability. Take end-to-end ownership of technical workstreams, from architecture and implementation through to deployment. … would be advantageous: Multimodal AI and reasoning. Edge or offline AI deployments. Kubernetes, particularly EKS or OpenShift. MLOps, including model evaluation, monitoring and reproducibility. Observability for agentic AI systems, including model performance, agent behaviour and drift. Agent orchestration and inter-agent communication protocols such as A2A. Model Context Protocol ...

Senior Connectivity Engineer

Hiring Organisation
Hackajob Ltd
Location
Wallingford, Oxfordshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
WireGuard/Tailscale or equivalent): access-as-code, policy patterns, posture/health automation, and resilience/disaster recovery planning. Deliver fleet-wide connectivity observability: monitoring, alerting, reporting, and actionable signals that help teams diagnose end-to-end issues quickly. Improve cellular/SIM lifecycle management: provisioning automation, usage/…/PMTUD, conntrack, nftables/iptables) and diagnosing kernel-level networking behaviour. Proficient in Go and/or Python and experienced with modern observability tooling; bonus points for containers/IoT OS, ACL-as-code patterns, and carrier/router API integrations. Benefits Starting from the interview process and continuing ...

Senior Software Engineer, Non-Realtime Controllers

Location
Oxford, England, United Kingdom
common to many controllers rather than specific to one. Work closely with the systems teams and scientists who depend on these services, and strengthen observability, health checking and alerting. Raise the standard of the code around you through review, and bring less experienced engineers on by working alongside them. Requirements … Rust codebase, including async programming. Experience designing and operating networked services such as gRPC or REST, with a real grasp of error handling, observability and failure modes. Experience of hardware and instrument protocols and interfaces such as SCPI, Modbus, serial and Ethernet. Comfortable owning the production behaviour of what ...

Fleet Connectivity Engineer — Automation & Observability

Location
Wallingford, England, United Kingdom
/or Python, and possess experience in robotics or connected devices. Responsibilities include evolving the connectivity stack, building self-serve workflows, and ensuring observability across the fleet. #J-18808-Ljbffr ...

Senior Connectivity Engineer / Network Engineer

Location
Wallingford, England, United Kingdom
WireGuard/Tailscale or equivalent): access‐as‐code, policy patterns, posture/health automation, and resilience/disaster recovery planning. Deliver fleet‐wide connectivity observability: monitoring, alerting, reporting, and actionable signals that help teams diagnose end‐to‐end issues quickly. Improve cellular/SIM lifecycle management: provisioning automation, usage/…/PMTUD, conntrack, nftables/iptables) and diagnosing kernel‐level networking behaviour. Proficient in Go and/or Python and experienced with modern observability tooling; bonus points for containers/IoT OS, ACL‐as‐code patterns, and carrier/router API integrations. #J-18808-Ljbffr ...

Software Engineer - Life AI Platform AI & Robotics Oxford, England, United Kingdom

Location
Oxford, England, United Kingdom
hand a ticket to at the beginning. Build and operate the technical foundations for agentic systems, including orchestration, tool interfaces, state management, evaluation, observability and failure recovery. Design evaluative frameworks to understand if the deployed model is working Work directly with scientists, observe how they operate and translate what … have dealt with agentsand their typical failure modes. You’ve designed andbuiltproduction-readysystems.You can make sound decisions about APIs, data models, testing, deployment, observability and failure recovery. You’ve built policy-aware systems with durable audit trails.Constraints are enforced by design, and every decision leaves a trustworthy, queryable record. ...