26 to 50 of 57 Observability Jobs in Berkshire

Senior AWS Platform Engineer

Location
Reading, England, United Kingdom
implementation and operational handover. The successful candidate will have experience delivering AWS platform solutions across multiple environments and a strong understanding of automation, security, observability and cloud engineering best practice. Due to the nature of this role successful candidates may need to undergo security clearance. More information on SC Clearance … from discovery through to operational handover. Lead AWS platform implementation activities, coordinating technical dependencies, guiding engineering decisions, promoting best practices across automation, security and observability, and supporting the development of early‐career colleagues through mentoring and knowledge sharing. About You Strong hands‐on AWS platform engineering experience, with a proven ...

Senior Platform Software Engineer

Location
Reading, England, United Kingdom
Job Description We are seeking a Software Developer 4 to build and operate AI and cloud platform capabilities that support a strategic enterprise customer engagement. This role will work in a fast-paced environment where ...

Fullstack Engineer - React/Go Microservices on AWS

Location
Bracknell, England, United Kingdom
TypeScript, while also designing scalable Go-based microservices and deploying in AWS. You will work across REST and gRPC APIs, and contribute to observability, reliability, and performance improvements. You will collaborate with Product Managers, UX Designers, and cross-functional teams in an Agile environment, participating in architecture reviews, code reviews ...

Senior Platform Engineer (Python)

Location
Slough, England, United Kingdom
workload scheduling, configuration, and runtime resource access. Experience building CI/CD pipelines with GitLab CI, GitHub Actions, Jenkins, or similar tools. Experience with observability tools such as Prometheus, Grafana, and OpenTelemetry. Good understanding of platform security: secrets management, IAM, network isolation, and dependency/supply-chain risks. Strong troubleshooting … teams to identify recurring pain points and turn them into scalable platform features. Automate manual infrastructure operations and build reliable self-service workflows. Implement observability through metrics, structured logging, and alerting. Own platform production issues from troubleshooting and root cause analysis to permanent resolution. Manage and scale on-prem compute ...

Production Python & AI Services Engineer (SC Cleared)

Location
Reading, England, United Kingdom
APIs within a large codebase, integrating LLM capabilities and improving security, reliability, and maintainability of live platforms. The role involves owning production issues, testing, observability, and documentation, with a focus on secure software development and scalable deployments. #J-18808-Ljbffr ...

Senior Cloud Infrastructure & Automation Engineer

Location
Bracknell, England, United Kingdom
platform reliability, automation, and scalable systems that power mission-critical CX and CCaaS services. You’ll design and implement IaC, strengthen security and observability, provide 3rd line support, coordinate zero-downtime deployments, and collaborate across disciplines to push forward a culture of continuous improvement. #J-18808-Ljbffr ...

Operations Team Lead (Production & Reliability)

Location
Slough, England, United Kingdom
Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under pressure, clear in communication Systems thinker, fixes root causes, not symptoms How We Think Production is sacred. Clear ownership beats ambiguity. ...

Senior Forward Deployment Engineer

Location
Slough, England, United Kingdom
analysis, upgrading Java and NPM runtimes, modernizing Spring and legacy middleware applications, improving CI/CD pipelines, containerizing applications, automating deployments, and introducing standard observability and resilience patterns. The Expert FDE is expected to lead complex engagements, work directly with development and client stakeholders, define the technical remediation approach, implement … testing, release, resilience, and legacy technology challenges with development teams. Assess application code, dependencies, runtime environment, test coverage, deployment architecture, CI/CD pipelines, observability, and operational risks. Write, debug, review, and enhance production-quality code and configuration throughout engagements. Define and implement practical modernization and remediation plans with clear ...

Senior DevOps Engineer

Hiring Organisation
Halian Technology Limited
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
reliability, and availability Implement self-service tooling to empower development teams Drive DevOps best practices across the digital product lifecycle Develop and enhance monitoring, observability, and incident response processes Support global engineering teams delivering high-traffic platforms Key Requirements Proven experience supporting digital product delivery in a DevOps or platform … with Infrastructure as Code (Terraform, Ansible, Puppet or similar) Hands-on experience with Kubernetes, Docker, and cloud platforms (AWS preferred) Experience with monitoring/observability tools (Prometheus, Grafana, ELK, APM tools) Solid understanding of system performance, scalability, and resilience Strong collaboration and communication skills within cross-functional product teams Desirable ...

Senior Site Reliability Engineer

Location
Reading, England, United Kingdom
inference infrastructure (GPU-backed endpoints, autoscaling, latency and cost tradeoffs), with SLOs, on-call, and incident response that cover models, not just services Observability(includes ML models) - drift and performance monitoring for ML, plus LLM-specific tracing, evals, and guardrails, wired into the same metrics and logging stacks … model registries, feature stores, and lineage (Kubeflow, MLflow, Feast, Weights & Biases, or equivalents) LLMOps in production - inference serving, prompt/version management, and LLM observability (tracing, evals, drift, guardrails, cost per request) Governing ML/LLM workloads as platform capabilities: data-residency and PII controls, and audit trails Any other ...

Monitoring & Observability Engineer

Location
Reading, England, United Kingdom
million Series C funding round – the largest fundraise ever completed by a quantum computing company in Europe. The Purpose As a Monitoring and Observability Engineer, you'll help keep OQC's live quantum computing systems running at their best. By improving system visibility, developing intelligent monitoring solutions and driving operational … Live Services team, you'll monitor the health and performance of our live cryogenic systems, respond to operational incidents and continuously improve our observability capabilities. You'll collaborate across engineering, operations, software and reliability teams to develop dashboards, refine alerting strategies and automate operational responses that improve reliability and reduce ...

Logs Specialist

Location
Maidenhead, England, United Kingdom
demonstrate the unique value of the Dynatrace GrailTM data lakehouse, helping customers transition from high-cost, fragmented logging silos to a unified, AI-powered observability platform.**Core Responsibilities****Domain Expertise:*** Act as the domain "subject matter expert" (SME) for Logs, staying ahead of industry trends like OpenTelemetry (OTel), log pipelines … management impacts MTTR and operational overhead.* Education: Bachelor's degree in - Computer science, Engineering, or equivalent practical experience.* Hands‐on exposure to modern observability pipelines, including Cribl solutions or OpenTelemetry logging specifications.* Familiarity with the Cribl ecosystem or OpenTelemetry logging specs.* Experience with scripting (Python, Go, or Bash ...

AI Platform Engineer — Remote-Eligible Platform Automation

Location
Reading, England, United Kingdom
engineers and security across the business to deliver scalable, secure platform capabilities. This hands-on engineering role covers platform delivery, AI-driven operations, resilience, observability and governance, with a focus on measurable outcomes and value across #J-18808-Ljbffr ...

Fullstack Engineer

Location
Bracknell, England, United Kingdom
Design and maintain scalable Go-based microservices Build and support REST and gRPC APIs Develop integrations and event-driven solutions across distributed systems Improve observability, reliability, scalability, and security Cloud & DevOps Deploy and support applications in AWS Work with Docker, Kubernetes, and CI/CD pipelines Participate in production support … communication skills Nice to Have gRPC and Protocol Buffers MongoDB experience AWS experience Kubernetes and container orchestration Event-driven architectures and messaging platforms Observability, monitoring, and distributed tracing Experience working in a SaaS product organisation #UKJobs #SoftwareEngineering #Golang #TechjobsUK #LI-AD1 Flexera is proud to be an equal opportunity employer. ...

Production Reliability Lead

Location
Slough, England, United Kingdom
change coordination. In this hands-on leadership role, you’ll build a scalable operating model, foster strong runbooks, and drive continuous improvement across observability, MTTR, and incident prevention #J-18808-Ljbffr ...

Senior DevOps Engineer

Hiring Organisation
Halian Technology Limited
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£95,000
Drive platform improvements and DevOps best practices. Design and implement self-service infrastructure and tooling. Deliver scalable, secure, and highly available systems. Enhance monitoring, observability, and operational performance. Support engineering teams with technical expertise and guidance. Skills & Experience Experience designing and implementing CI/CD pipelines and software delivery processes. … Infrastructure as Code experience using tools such as Terraform or Ansible. Experience with monitoring and observability tools. Strong knowledge of Docker, Kubernetes, AWS, and cloud technologies. Excellent communication skills and ability to collaborate across teams. A passion for automation, platform engineering, and continuous improvement. This is a full-time, permanent ...

Security & Network Engineer - 12 months Fixed term

Hiring Organisation
Techtronic Industries - Europe HQ
Location
Maidenhead, Berkshire, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
cloud landing zones) Manage incident response for critical infrastructure events; lead post-mortems and remediation Collaborate with infrastructure teams to build monitoring, alerting, and observability stacks that surface security signals Required Experience & Skills Over 5 years of experience in network engineering, infrastructure architecture, or systems engineering roles Proven hands … technical and non-technical stakeholders Experience with infrastructure-as-code tools (Terraform, CloudFormation, ARM) and configuration management (Ansible, etc.) Proficiency with network monitoring and observability tools (e.g., Splunk, Datadog, New Relic, Elasticsearch, Prometheus, RSA NetWitness, Tenable) Strong incident response and troubleshooting background; comfort operating in high-pressure environments Experience mentoring ...

Security & Network Engineer - 12 months Fixed term

Location
Maidenhead, England, United Kingdom
cloud landing zones) Manage incident response for critical infrastructure events; lead post-mortems and remediation Collaborate with infrastructure teams to build monitoring, alerting, and observability stacks that surface security signals Required Experience & Skills Over 5 years of experience in network engineering, infrastructure architecture, or systems engineering roles Proven hands … technical and non-technical stakeholders Experience with infrastructure-as-code tools (Terraform, CloudFormation, ARM) and configuration management (Ansible, etc.) Proficiency with network monitoring and observability tools (e.g., Splunk, Datadog, New Relic, Elasticsearch, Prometheus, RSA NetWitness, Tenable) Strong incident response and troubleshooting background; comfort operating in high-pressure environments Experience mentoring ...

AI Platform Engineer

Location
Reading, England, United Kingdom
across the business adopt AI solutions safely and effectively. This is a hands-on engineering role that combines feature delivery with responsibility for resilience, observability, governance and measurable outcomes. You’ll work closely with Product Owners, Architects, Engineers, Security teams and business stakeholders to deliver secure, scalable and reliable platform … applications, APIs, services or platforms. Experience using logs, metrics, traces and telemetry to improve platform performance and reliability. Experience designing solutions with resilience, scalability, observability and operability in mind. Evidence of improving the dependability of production systems through practical engineering changes. A track record of delivering technology capabilities into production ...

SRE: Cloud Reliability, DevSecOps & Observability — Hybrid London

Location
Slough, England, United Kingdom
will apply software engineering principles to automate, scale, and secure cloud-native environments. Responsibilities include building and maintaining production and demo environments, implementing observability with Prometheus, Grafana, and Loki, and guiding project teams in DevSecOps practices. Hybrid London model, SC level clearance may be required. #J-18808-Ljbffr ...

Fullstack Engineer

Location
Bracknell, England, United Kingdom
Design and maintain scalable Go-based microservices Build and support REST and gRPC APIs Develop integrations and event-driven solutions across distributed systems Improve observability, reliability, scalability, and security Deploy and support applications in AWS Work with Docker, Kubernetes, and CI/CD pipelines Participate in production support … Nice to have: AWS experience Nice to have: Kubernetes and container orchestration Nice to have: Event-driven architectures and messaging platforms Nice to have: Observability, monitoring, and distributed tracing Nice to have: Experience working in a SaaS product organisation Core Competencies Demonstrates expertise in building modern web applications using React ...

Data Platform DevOps Analyst

Location
Reading, England, United Kingdom
journey supporting the migration from Teradata to Databricks, validating production readiness, improving monitoring and automation capabilities and helping embed modern DataOps practices across deployment, observability and support processes. What You’ll Bring Here at Primark, we want everyone to feel valued – so please bring your authentic self to work … platforms and DataOps practices with exposure to Azure DevOps, Databricks, dbt Cloud, Azure data services, SQL, Python, ETL/ELT processes, job scheduling, monitoring, observability and deployment automation. Experience supporting production environments and service operations including monitoring, alerting, incident management, ticketing systems, release management, environment management, change governance and enterprise ...

Senior Data Engineer

Location
Maidenhead, England, United Kingdom
data solutions are reliable, scalable, performant, secure, and production‐ready Monitor, troubleshoot, and continuously improve pipeline performance, data quality, and platform stability Drive automation, observability, and supportability across data, analytics, and AI/ML solutions Our Ideal Candidate Strong data engineering experience with hands‐on delivery of scalable data pipelines … productionization, LLM‐based applications, or agentic AI patterns will be an added advantage Experience with DevOps and DataOps practices, including CI/CD, monitoring, observability, and incident support Maersk is committed to a diverse and inclusive workplace, and we embrace different styles of thinking. Maersk is an equal opportunities employer ...

Platform Engineering Manager (SRE)

Location
Bracknell, England, United Kingdom
accelerate software delivery and operational performance. Build and maintain core platform capabilities: CI/CD pipelines, build and release automation, infrastructure automation, environment management, observability, and developer tooling. Drive standardisation and reduce engineering friction through automation and self-service. Site Reliability Engineering Introduce and embed SRE practices across Engineering. Improve … scale. A background built in cloud product or SaaS companies, where reliability is something customers feel directly. Deep, hands-on SRE expertise: monitoring, observability, alerting, incident management, and operational readiness. You can define what good looks like and stand up foundational SRE practices, from SLOs and error budgets to public ...

Principal Consultant – Agentic AI, Integration & API Architecture

Hiring Organisation
NeosAlpha Technologies
Location
Maidenhead, England, United Kingdom
secure, governed, and production-ready agentic architectures. The role spans Agentic Architecture, Agent Control Planes, AI Gateways, MCP/A2A, Agentic Security & Governance, Observability and Agentic Quality Engineering, alongside enterprise integration and API architecture. As Agentic AI is an emerging field, we do not expect candidates to have many years …/tool discovery and registries o Agent identity and delegated identity o Human-in-the-loop controls o Context engineering and RAG o Agent observability, auditability and cost management o Agentic Security and Governance o Agentic Quality Engineering o Hybrid and multi-cloud agent architectures Assess how platforms such ...