14 of 14 Permanent Observability Jobs in Slough

Cloud Architect

Location
Slough, England, United Kingdom
Engineering Collaborate with platform engineering to deliver secure, scalable cloud platforms. Drive IaC adoption using Terraform and native tooling. Support automation, CI/CD, observability, and reliability initiatives. Define standards for networking, identity, security, monitoring, and resilience. Governance & Security Ensure solutions comply with banking regulatory requirements. Work with Security Architecture ...

Senior Platform Engineer: AI-Ready Infra & Security

Location
Slough, England, United Kingdom
scalable platform features, automate operations, and build self-service workflows. The role emphasizes strong Python, Linux, Terraform/Ansible, Docker and Kubernetes proficiency, plus observability with Prometheus, Grafana and OpenTelemetry. #J-18808-Ljbffr ...

Platform Engineer

Location
Slough, England, United Kingdom
Core Focus) Build secure, scalable Azure and GCP cloud platforms. Drive IaC adoption using Terraform and cloud-native tooling. Implement automation, CI/CD, observability, and reliability engineering practices. Define standards for networking, identity, security, monitoring, and resilience. Improve developer experience through platform consistency, automation, and self-service. Governance, Security ...

Java Developer

Location
Slough, England, United Kingdom
Exposure to data-driven or analytical systems; experience with AI/ML-related features is beneficial but not required. Familiarity with CI/CD, observability, and operating production systems. Understanding of front-office or revenue-generating technology in an investment bank, together with knowledge of FIC products (FX, Credit, Rates ...

Senior Product Manager

Location
Slough, England, United Kingdom
continuous improvement. Desirable Skills: Experience with data platforms such as Snowflake, Azure Data Lake, or similar cloud-based technologies. Understanding of model monitoring, observability, and performance tracking practices. Exposure to AI/ML-enabled products and data monetization initiatives. Rewards & Benefits TCS is consistently voted a Top Employer ...

Site Reliability Engineer

Location
Slough, England, United Kingdom
Site Reliability Engineer (SRE) DevSecOps | Cloud Engineering | Observability | Production Environments | London SR2 is supporting a major 3-year programme and looking for an experienced Site Reliability Engineer (SRE) to join the Production Engineering team. This function underpins the reliability, security, and performance of all live environments, from production systems … native infrastructure. Beyond supporting live systems, this team also acts as a centre of excellence, guiding project teams in adopting best practices across DevSecOps, observability, and cost optimisation. Key Responsibilities: Build, maintain, and support production and demo environments Automate infrastructure provisioning and deployment workflows (Terraform, GitHub Actions, GitOps) Package ...

Senior Software Engineer

Location
Slough, England, United Kingdom
developments and bring relevant patterns back to the team. Break down stories into tasks, provide estimates, and surface risks early in sprint planning. Maintain observability standards: structured logging, metrics, and distributed tracing. Support CI/CD pipelines and participate in production readiness reviews. Mentor junior engineers through pair programming, code … sprint planning, story decomposition, backlog grooming, retrospectives. Strong unit and component testing discipline; exposure to BDD or contract testing is a plus. Appreciation for observability: structured logging, distributed tracing, alerting hygiene. Desirable Qualifications and Skills Insurance or Insurtech domain knowledge — Policy Administration, Claims, or Underwriting workflows. Familiarity with Kafka ...

Senior Platform Engineer (Python)

Location
Slough, England, United Kingdom
workload scheduling, configuration, and runtime resource access. Experience building CI/CD pipelines with GitLab CI, GitHub Actions, Jenkins, or similar tools. Experience with observability tools such as Prometheus, Grafana, and OpenTelemetry. Good understanding of platform security: secrets management, IAM, network isolation, and dependency/supply-chain risks. Strong troubleshooting … teams to identify recurring pain points and turn them into scalable platform features. Automate manual infrastructure operations and build reliable self-service workflows. Implement observability through metrics, structured logging, and alerting. Own platform production issues from troubleshooting and root cause analysis to permanent resolution. Manage and scale on-prem compute ...

Operations Team Lead (Production & Reliability)

Location
Slough, England, United Kingdom
Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under pressure, clear in communication Systems thinker, fixes root causes, not symptoms How We Think Production is sacred. Clear ownership beats ambiguity. ...

Senior Forward Deployment Engineer

Location
Slough, England, United Kingdom
analysis, upgrading Java and NPM runtimes, modernizing Spring and legacy middleware applications, improving CI/CD pipelines, containerizing applications, automating deployments, and introducing standard observability and resilience patterns. The Expert FDE is expected to lead complex engagements, work directly with development and client stakeholders, define the technical remediation approach, implement … testing, release, resilience, and legacy technology challenges with development teams. Assess application code, dependencies, runtime environment, test coverage, deployment architecture, CI/CD pipelines, observability, and operational risks. Write, debug, review, and enhance production-quality code and configuration throughout engagements. Define and implement practical modernization and remediation plans with clear ...

Production Reliability Lead

Location
Slough, England, United Kingdom
change coordination. In this hands-on leadership role, you’ll build a scalable operating model, foster strong runbooks, and drive continuous improvement across observability, MTTR, and incident prevention #J-18808-Ljbffr ...

SRE: Cloud Reliability, DevSecOps & Observability — Hybrid London

Location
Slough, England, United Kingdom
will apply software engineering principles to automate, scale, and secure cloud-native environments. Responsibilities include building and maintaining production and demo environments, implementing observability with Prometheus, Grafana, and Loki, and guiding project teams in DevSecOps practices. Hybrid London model, SC level clearance may be required. #J-18808-Ljbffr ...

Lead ClickHouse Solutions Architect Greenfield Observability

Location
Slough, England, United Kingdom
Colehouse Group is seeking a ClickHouse Solutions Architect to lead a greenfield, enterprise-scale observability platform for a global banking client in Slough. You will shape architecture from first principles to meet high throughput, multi-tenant isolation and storage cost constraints. You'll design data models, ingestion and retention strategies ...

Clickhouse Solutions Architect

Location
Slough, England, United Kingdom
Role We are looking for a ClickHouse Solutions Architect to join our team supporting the design and implementation of a greenfield, enterprise-scale ClickHouse observability platform for a global banking client. This is a genuine greenfield build at significant scale — there is no incumbent platform to inherit or work around. … where benchmarks disprove the design Establish infrastructure-as-code, CI/CD and environment promotion for schema and configuration changes Productionisation Define and implement observability — system table monitoring, metrics, alerting thresholds, capacity headroom tracking Establish backup, restore and disaster recovery, and validate them by test Implement security and governance — RBAC ...