51 to 75 of 99 Observability Jobs in the Thames Valley

Principal Java Engineer

Location
Wallingford, England, United Kingdom
continuous improvement. Production systems are reliable, observable and operationally excellent Lead root cause analysis and resolution of complex production issues. Drive improvements in system observability, monitoring and operational performance. Ensure applications are designed and operated to meet reliability, availability and performance targets. Partner with Operations, DevOps and QA teams … Claude, Codex, Gitlab Duo, etc) REST APIs, OpenAPI, Microservices, Event-driven architecture (RabbitMQ) Containers, Docker, AWS, Linux CI/CD with GitLab Pipelines & Jenkins Observability: logging, metrics and monitoring MySQL, Apache Solr Front-end UI (e.g. Angular) Person Specification Strategic and systems-thinking mindset Excellent communication and stakeholder management skills ...

Automation & Platform Engineer

Location
Abingdon, England, United Kingdom
services, agent services, APIs, and microservices. Implement infrastructure-as-code for platform environments and network automation resources. Ensure automation platforms meet security, compliance, availability, observability, and operational resilience requirements. Your Profile Experience in network automation, platform engineering, DevOps, cloud engineering, or telecom automation. Strong hands-on experience with Ansible, Terraform … vendor APIs. API management and orchestration. Intent-to-configuration workflows. Data pipelines. Vector databases and graph APIs. MCP integration. Security frameworks and compliance controls. Observability and logging. Preferred Certifications Kubernetes CKA/CKAD. Terraform Associate. Red Hat Ansible certification. Google Cloud, AWS, or Azure certification. Cisco, Juniper, Nokia, or Ericsson ...

Production Python & AI Services Engineer (SC Cleared)

Location
Reading, England, United Kingdom
APIs within a large codebase, integrating LLM capabilities and improving security, reliability, and maintainability of live platforms. The role involves owning production issues, testing, observability, and documentation, with a focus on secure software development and scalable deployments. #J-18808-Ljbffr ...

Senior Cloud Infrastructure & Automation Engineer

Location
Bracknell, England, United Kingdom
platform reliability, automation, and scalable systems that power mission-critical CX and CCaaS services. You’ll design and implement IaC, strengthen security and observability, provide 3rd line support, coordinate zero-downtime deployments, and collaborate across disciplines to push forward a culture of continuous improvement. #J-18808-Ljbffr ...

Principal Java Engineer - Payments, Low-Latency Real-Time

Location
Milton Keynes, England, United Kingdom
this hybrid role, you will deploy and monitor on OpenShift and AWS, integrate with MongoDB, and mentor peers while driving architectural decisions and observability enhancements across services. #J-18808-Ljbffr ...

Operations Team Lead (Production & Reliability)

Location
High Wycombe, England, United Kingdom
Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under pressure, clear in communication Systems thinker, fixes root causes, not symptoms How We Think Production is sacred. Clear ownership beats ambiguity. ...

Operations Team Lead (Production & Reliability)

Location
Slough, England, United Kingdom
Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under pressure, clear in communication Systems thinker, fixes root causes, not symptoms How We Think Production is sacred. Clear ownership beats ambiguity. ...

Senior Forward Deployment Engineer

Location
Slough, England, United Kingdom
analysis, upgrading Java and NPM runtimes, modernizing Spring and legacy middleware applications, improving CI/CD pipelines, containerizing applications, automating deployments, and introducing standard observability and resilience patterns. The Expert FDE is expected to lead complex engagements, work directly with development and client stakeholders, define the technical remediation approach, implement … testing, release, resilience, and legacy technology challenges with development teams. Assess application code, dependencies, runtime environment, test coverage, deployment architecture, CI/CD pipelines, observability, and operational risks. Write, debug, review, and enhance production-quality code and configuration throughout engagements. Define and implement practical modernization and remediation plans with clear ...

Senior DevOps Engineer

Hiring Organisation
Halian Technology Limited
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
reliability, and availability Implement self-service tooling to empower development teams Drive DevOps best practices across the digital product lifecycle Develop and enhance monitoring, observability, and incident response processes Support global engineering teams delivering high-traffic platforms Key Requirements Proven experience supporting digital product delivery in a DevOps or platform … with Infrastructure as Code (Terraform, Ansible, Puppet or similar) Hands-on experience with Kubernetes, Docker, and cloud platforms (AWS preferred) Experience with monitoring/observability tools (Prometheus, Grafana, ELK, APM tools) Solid understanding of system performance, scalability, and resilience Strong collaboration and communication skills within cross-functional product teams Desirable ...

Senior Site Reliability Engineer

Location
Reading, England, United Kingdom
inference infrastructure (GPU-backed endpoints, autoscaling, latency and cost tradeoffs), with SLOs, on-call, and incident response that cover models, not just services Observability(includes ML models) - drift and performance monitoring for ML, plus LLM-specific tracing, evals, and guardrails, wired into the same metrics and logging stacks … model registries, feature stores, and lineage (Kubeflow, MLflow, Feast, Weights & Biases, or equivalents) LLMOps in production - inference serving, prompt/version management, and LLM observability (tracing, evals, drift, guardrails, cost per request) Governing ML/LLM workloads as platform capabilities: data-residency and PII controls, and audit trails Any other ...

Monitoring & Observability Engineer

Location
Reading, England, United Kingdom
million Series C funding round – the largest fundraise ever completed by a quantum computing company in Europe. The Purpose As a Monitoring and Observability Engineer, you'll help keep OQC's live quantum computing systems running at their best. By improving system visibility, developing intelligent monitoring solutions and driving operational … Live Services team, you'll monitor the health and performance of our live cryogenic systems, respond to operational incidents and continuously improve our observability capabilities. You'll collaborate across engineering, operations, software and reliability teams to develop dashboards, refine alerting strategies and automate operational responses that improve reliability and reduce ...

Logs Specialist

Location
Maidenhead, England, United Kingdom
demonstrate the unique value of the Dynatrace GrailTM data lakehouse, helping customers transition from high-cost, fragmented logging silos to a unified, AI-powered observability platform.**Core Responsibilities****Domain Expertise:*** Act as the domain "subject matter expert" (SME) for Logs, staying ahead of industry trends like OpenTelemetry (OTel), log pipelines … management impacts MTTR and operational overhead.* Education: Bachelor's degree in - Computer science, Engineering, or equivalent practical experience.* Hands‐on exposure to modern observability pipelines, including Cribl solutions or OpenTelemetry logging specifications.* Familiarity with the Cribl ecosystem or OpenTelemetry logging specs.* Experience with scripting (Python, Go, or Bash ...

AI Platform Engineer — Remote-Eligible Platform Automation

Location
Reading, England, United Kingdom
engineers and security across the business to deliver scalable, secure platform capabilities. This hands-on engineering role covers platform delivery, AI-driven operations, resilience, observability and governance, with a focus on measurable outcomes and value across #J-18808-Ljbffr ...

Advanced Engineer, Investment Technology

Location
Henley-on-Thames, England, United Kingdom
subject matter experts. Their primary focus is on execution within defined parameters, applying modern engineering practices including CI/CD, automated testing, code quality, observability and data quality controls. They may also be accountable for regular reporting or process administration within the squad. Working in partnership with more experienced staff … business requirements into practical, well-engineered technical solutions. Build production-ready solutions using modern engineering practices, including CI/CD, automated testing, code quality, observability and data quality controls. Support data quality, monitoring, reporting and process administration activities to help ensure accurate and consistent outcomes. Identify and resolve technical problems ...

Fullstack Engineer

Location
Bracknell, England, United Kingdom
Design and maintain scalable Go-based microservices Build and support REST and gRPC APIs Develop integrations and event-driven solutions across distributed systems Improve observability, reliability, scalability, and security Cloud & DevOps Deploy and support applications in AWS Work with Docker, Kubernetes, and CI/CD pipelines Participate in production support … communication skills Nice to Have gRPC and Protocol Buffers MongoDB experience AWS experience Kubernetes and container orchestration Event-driven architectures and messaging platforms Observability, monitoring, and distributed tracing Experience working in a SaaS product organisation #UKJobs #SoftwareEngineering #Golang #TechjobsUK #LI-AD1 Flexera is proud to be an equal opportunity employer. ...

Production Reliability Lead

Location
Slough, England, United Kingdom
change coordination. In this hands-on leadership role, you’ll build a scalable operating model, foster strong runbooks, and drive continuous improvement across observability, MTTR, and incident prevention #J-18808-Ljbffr ...

Senior DevOps Engineer

Hiring Organisation
Halian Technology Limited
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£95,000
Drive platform improvements and DevOps best practices. Design and implement self-service infrastructure and tooling. Deliver scalable, secure, and highly available systems. Enhance monitoring, observability, and operational performance. Support engineering teams with technical expertise and guidance. Skills & Experience Experience designing and implementing CI/CD pipelines and software delivery processes. … Infrastructure as Code experience using tools such as Terraform or Ansible. Experience with monitoring and observability tools. Strong knowledge of Docker, Kubernetes, AWS, and cloud technologies. Excellent communication skills and ability to collaborate across teams. A passion for automation, platform engineering, and continuous improvement. This is a full-time, permanent ...

Security & Network Engineer - 12 months Fixed term

Hiring Organisation
Techtronic Industries - Europe HQ
Location
Maidenhead, Berkshire, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
cloud landing zones) Manage incident response for critical infrastructure events; lead post-mortems and remediation Collaborate with infrastructure teams to build monitoring, alerting, and observability stacks that surface security signals Required Experience & Skills Over 5 years of experience in network engineering, infrastructure architecture, or systems engineering roles Proven hands … technical and non-technical stakeholders Experience with infrastructure-as-code tools (Terraform, CloudFormation, ARM) and configuration management (Ansible, etc.) Proficiency with network monitoring and observability tools (e.g., Splunk, Datadog, New Relic, Elasticsearch, Prometheus, RSA NetWitness, Tenable) Strong incident response and troubleshooting background; comfort operating in high-pressure environments Experience mentoring ...

Security & Network Engineer - 12 months Fixed term

Location
Maidenhead, England, United Kingdom
cloud landing zones) Manage incident response for critical infrastructure events; lead post-mortems and remediation Collaborate with infrastructure teams to build monitoring, alerting, and observability stacks that surface security signals Required Experience & Skills Over 5 years of experience in network engineering, infrastructure architecture, or systems engineering roles Proven hands … technical and non-technical stakeholders Experience with infrastructure-as-code tools (Terraform, CloudFormation, ARM) and configuration management (Ansible, etc.) Proficiency with network monitoring and observability tools (e.g., Splunk, Datadog, New Relic, Elasticsearch, Prometheus, RSA NetWitness, Tenable) Strong incident response and troubleshooting background; comfort operating in high-pressure environments Experience mentoring ...

Operations Team Lead — Reliability & Incident Leader

Location
High Wycombe, England, United Kingdom
will shape on-call rotations, maintain high availability, and ensure safe production changes. The role requires hands-on leadership and a strong focus on observability and operational discipline. #J-18808-Ljbffr ...

Infrastructure & Security Manager

Location
Milton Keynes, England, United Kingdom
WLAN Cloud infrastructure Identity & Access Management End User Computing Security technologies including SIEM, EDR/XDR, ZTNA, CASB and next-generation firewalls Monitoring and observability DR/BCP Managed service providers We're particularly interested in people who have worked across complex, multi-site environments such as retail, hospitality, leisure ...

Platform Engineer

Location
Milton Keynes, England, United Kingdom
Terraform, CloudFormation, or CDK), CI/CD pipelines, API Gateway, Lambda, and Aurora PostgreSQL — and confident applying fundamentals such as high availability, fault tolerance, observability, and cost control in a live environment. Payments or fintech exposure is an advantage but not required; what matters more is a practical, first-principles … rollback processes Support deployment of AI generated applications and tooling Support change control and release management alongside Engineering and the Information Security Officer Reliability, Observability & Data Infrastructure Design and maintain systems for high availability, fault tolerance, and resilience by default Implement logging, monitoring, and alerting (for example CloudWatch) across services ...

Senior Site Reliability Engineer

Location
Milton Keynes, England, United Kingdom
experience with both Azure, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways … Engineer Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles. Regularly use Datadog and other observability tools for application performance monitoring. Implement new ways of working, helping to shape how the organisation responds and recovers to incidents. Take ownership of incident ...

Senior Site Reliability Engineer

Hiring Organisation
VIQU IT Recruitment
Location
Milton Keynes, Buckinghamshire, United Kingdom
Employment Type
Full-Time
Salary
£65,000 - £75,000 per annum
experience with both Azure, and on-premise virtual machines. Experience withInfrastructure as Code/Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor). Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways … Engineer Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles. Regularly use Datadog and other observability tools for application performance monitoring. Implement new ways of working, helping to shape how the organisation responds and recovers to incidents. Take ownership of incident ...