351 to 375 of 502 Remote/Hybrid Observability Jobs

Principal Platform Engineer

Hiring Organisation
SF Partners Admin
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent, Work From Home
capabilities. Design and operate production-grade Kubernetes platforms, including EKS, AKS or OpenShift. Define engineering standards, golden paths, reusable modules and platform patterns. Build observability strategies using Prometheus, Grafana, OpenTelemetry and modern APM tooling. Improve reliability through SLOs, incident reviews and Site Reliability Engineering (SRE) practises. Embed DevSecOps, supply-chain … Infrastructure as Code (IaC). CI/CD automation. GitOps tools such as ArgoCD or Flux. Internal Developer Platforms or self-service engineering. Observability tools including Prometheus, Grafana, OpenTelemetry, ELK, Datadog, Dynatrace or New Relic. DevSecOps and supply-chain security. SRE practises, SLOs, SLIs and incident management. Platform governance, cloud ...

Principal Infrastructure Engineer

Hiring Organisation
Sidram tech
Location
San Francisco, California, United States
Employment Type
Permanent
Salary
USD Annual
secure, highly available Kubernetes platforms across multi-cloud environments. Collaborate with customers, product teams, and engineering to deliver scalable infrastructure solutions. Drive platform reliability, observability, security, and performance optimization. Conduct infrastructure code reviews, security reviews, and architecture improvements. Support production environments and ensure operational excellence. Required Qualifications Bachelor's Degree … platforms. Knowledge of data and ML platforms including Snowflake and Databricks. Experience with cloud networking (AWS VPC, Azure VNet, GCP VPC). Experience with observability tools (Prometheus, ELK Stack, Grafana, or similar). Experience with service mesh technologies such as Istio or Linkerd. Strong understanding of cloud security, scalability, reliability ...

Performance Test Engineer - SC Cleared

Hiring Organisation
Lorien
Location
London, UK
Employment Type
Full-time
performance testing strategies and non-functional requirements. \n Identify performance bottlenecks and provide actionable recommendations for improvement. \n Monitor application and infrastructure performance using observability and diagnostic tools. \n Collaborate with engineering teams to improve system scalability and resilience. \n Integrate performance testing into CI/CD pipelines and automated … analysing performance, load, stress, and volume tests. \n Strong knowledge of cloud platforms including Azure, AWS, or GCP. \n Experience with monitoring and observability tools such as Grafana, Splunk, New Relic, or Dynatrace. \n Scripting and automation skills using Python, JavaScript, or similar technologies. \n Understanding of Docker, Kubernetes ...

Senior DevOps Engineer

Hiring Organisation
Halian Technology Limited
Location
Basingstoke, Hampshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£95,000
Drive platform improvements and DevOps best practices. Design and implement self-service infrastructure and tooling. Deliver scalable, secure, and highly available systems. Enhance monitoring, observability, and operational performance. Support engineering teams with technical expertise and guidance. Skills & Experience Experience designing and implementing CI/CD pipelines and software delivery processes. … Infrastructure as Code experience using tools such as Terraform or Ansible. Experience with monitoring and observability tools. Strong knowledge of Docker, Kubernetes, AWS, and cloud technologies. Excellent communication skills and ability to collaborate across teams. A passion for automation, platform engineering, and continuous improvement. This is a full-time, permanent ...

Senior Tech Lead - FinTech

Hiring Organisation
Carousel Consultancy Ltd
Location
London, South East, England, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
Shaping the platform architecture Working closely with the tech team to evolve the platform Designing scalable backend systems and services Improving reliability, performance and observability Helping modernise legacy parts of the platform Enhancing the platform and DevOps - collaborating with AWS infrastructure and cloud-native services, refining CI/CD pipelines … Native Infrastructure: AWS (ECS, EKS, RDS, S3, Lambda) Containers and Orchestration: Docker, Kubernetes CI/CD: Jenkins, GitHub Actions Databases: MySQL, PostgreSQL Monitoring and Observability: Sentry, CloudWatch, Grafana Skills and experience required: Solid software engineering experience (c8+ years), working as a Senior, Staff, Principal Engineer or Tech Lead FinTech ...

Senior Engineering Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
health: tech debt, refactoring, security, and governance Take ownership of features from start to finish using agile methodologies, delivering with feature flags, tests, and observability Champion high‐quality technical communications: proposals, specs, testing reports, and release planning Drive AI‐first ways of working within the squad — embedding AI tooling into … relates to the Manage domain Contribute to CI/CD pipeline improvements and progressive delivery practices across squads Drive reliability monitoring and observability within Manage (Prometheus, Grafana, Sentry) Contribute to security posture improvements: vulnerability scanning, pen testing coordination, and enforcement of standards Cross‐Squad & Leadership Collaboration Work closely with ...

AWS API Engineer / Data Architect

Hiring Organisation
Capgemini
Location
Hampshire, United Kingdom
Employment Type
Full Time
will design RESTful APIs and service boundaries, define integration standards, and collaborate with product, delivery and cybersecurity stakeholders to ensure secure access patterns, strong observability, and smooth deployments and cutovers. Where required, you will also contribute to data architecture decisions (e.g., PostgreSQL and document stores) to enable robust … will be able to design secure, scalable and well-governed RESTful APIs and microservices on AWS, applying consistent standards for identity, access control, observability, auditability and operational resilience. You will be comfortable collaborating with stakeholders to translate outcomes into pragmatic architectures and delivery plans, and you will bring strong engineering ...

AI Consulting –Sr. AI Architect & Client Partner

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
services that underpin enterprise AI ecosystems. Guide teams on software architecture, performance, scalability, security, and maintainability. AI governance & LLMOps Architect governance frameworks for auditability, observability, explainability, and compliance. Design guardrails for hallucination, prompt injection, toxicity, and model safety. Establish LLMOps: evaluation pipelines, automated testing, CI/CD, monitoring, and production … with one of Azure, OpenAI, AWS Bedrock, Claude; Kubernetes and cloud-native deployment. LLMOps & evaluation : CI/CD for AI, automated evals, experiment tracking, observability, model lifecycle management. Responsible AI : governance frameworks, guardrails, model safety, compliance, and auditability. Preferred but not required Orchestration frameworks : LangChain/LangGraph, LlamaIndex, CrewAI, AutoGen ...

Senior Software Engineer II, Developer Experience / Operational Excellence

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
confidently. Within DevEx, the Operational Excellence (OPX) team is the group that keeps production healthy at scale. We provide engineering teams the platform capabilities, observability tooling, automated safeguards, incident management tooling, and safe feature release systems they need to deliver highly available systems, ship features with confidence, and investigate … health. Reduce alert noise, surface actionable signals, and empower engineering teams to operate their services confidently with minimal operational burden Develop and evolve our observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection, to give teams real-time, actionable visibility into system health and latency Contribute to AI-driven ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
South West London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
17918
Location
United Kingdom
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Cardiff, South Glamorgan, Wales, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Belfast, County Antrim, Northern Ireland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer (SRE Squad Lead)

Hiring Organisation
Inspire People
Location
Darlington, County Durham, North East, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Platform Chapter Lead - Engineering

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
paved paths, less friction — using metrics (e.g. DORA) and real feedback to keep improving it. Keep it reliable and compliant. Oversee performance, resilience and observability for revenue‐critical services through peak traffic, and maintain security and compliance (e.g. PCI‐DSS, GDPR). Bring the business with you. Align platform strategy … balance cost, speed and risk. Strong cloud‐native and modern DevOps background — cloud (ideally AWS) and Kubernetes, CI/CD, infrastructure as code and observability — with enough engineering depth (e.g. Java, .NET, Python) to be credible with strong engineers. Excellent communication and the ability to influence at executive level, plus ...

Lead Java Developer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
adoption and ensure successful rollout of new capabilities. Lead root cause analysis on production issues, drive long‐term stability improvements, and strengthen monitoring and observability across the platform. Recommended Experience Strong experience in Core Java, J2EE, Spring Framework Exposure to Python scripting and data analysis Experience in fast moving Capital … such as Kafka, JMS, gRPC etc Proficient in latency measurement and performance optimization of Java based platforms with focus on JVM tuning Experience with observability stacks like ELK, Prometheus, Grafana, Kiali, Jaeger etc. Sound knowledge for persistence technologies such as relational databases, NoSQL databases, off heap storages and distributed caches ...

Lead Java Developer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
adoption and ensure successful rollout of new capabilities.* Lead root cause analysis on production issues, drive long‐term stability improvements, and strengthen monitoring and observability across the platform.**Recommended Experience:*** Strong experience in Core Java, J2EE, Spring Framework* Exposure to Python scripting and data analysis* Experience in fast moving Capital … such as Kafka, JMS, gRPC etc* Proficient in latency measurement and performance optimization of Java based platforms with focus on JVM tuning* Experience with observability stacks like ELK, Prometheus, Grafana, Kiali, Jaeger etc.* Sound knowledge for persistence technologies such as relational databases, NoSQL databases, off heap storages and distributed caches ...

Staff or Principal Data Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
create unnecessary complexity, risk, duplicated capability or long‐term support burden. Raise the quality bar for data products through clear ownership, robust testing, reconciliation, observability, lineage, documentation, performance and supportability. Collaborate with cross‐functional teams to address security, GDPR, PII handling, role‐based access, auditability and data governance are designed … services across batch, streaming and event‐driven patterns. Deep understanding of engineering practice: clean design, testing strategy, CI/CD, infrastructure as code, observability, performance, security, incident response and DevSecOps. Experience with cloud data services and modern data stacks. Relevant technologies may include Snowflake, Azure/AWS/GCP data ...

Head of Technology Operations

Hiring Organisation
Jobleads-UK
Location
Halifax, England, United Kingdom
adoption of infrastructure as code (IaC), CI/CD pipelines, and automated testing within platform operations. Champion site reliability engineering (SRE) practices, embedding monitoring, observability, and incident response, and continuously improving performance metrics including application load times, throughput, and error rates. Partner with the Head of Development and the Director … technologies. Deep expertise in Kubernetes, cloud networking, CI/CD pipelines, infrastructure as code, and platform security. Experience leading site reliability engineering (SRE), monitoring, observability, and performance optimisation (load times, application speed, latency management). Experience providing senior‐level escalation support for complex infrastructure, firewall, networking, and systems issues. ITIL ...

Global Head of SRE & Reliability – Hybrid Role

Hiring Organisation
Jobleads-UK
Location
Bristol, England, United Kingdom
office and collaboration across Engineering, Infrastructure Operations and Security to boost reliability and performance of critical platforms. You will define reliability strategy, drive automation, observability, and incident maturity, and scale the organization while embedding reliability into the #J-18808-Ljbffr ...

Data Platform Lead — AI-Ready, Scalable Pipelines (Hybrid)

Hiring Organisation
Jobleads-UK
Location
Cardiff, Wales, United Kingdom
models, ensuring reliability and AI readiness while coaching engineers and setting technical standards. You will balance immediate delivery with long-term platform evolution, improve observability and data quality, and drive governance across datasets to empower self-service analytics. #J-18808-Ljbffr ...

Senior Developer Experience Engineer - Platform Tools

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
create paved roads that help engineers ship quickly and securely. You’ll work across the lifecycle—from code through testing, deployment and production observability—collaborating with the Platform team to reduce friction and measure adoption. Hybrid work is supported in London and beyond. #J-18808-Ljbffr ...