251 to 275 of 4,103 Observability Jobs

Principal Security Architect — DevSecOps (Group Manager II - Information Security)

Hiring Organisation
UST Global
Location
London, UK
Employment Type
Full-time
repository controls, signed artifacts, environment promotion, drift detection, and deployment guardrails. Software supply-chain security: SBOMs, artifact signing, provenance, dependency controls, and build integrity. Observability and security monitoring: Prometheus, Grafana, CloudWatch, OpenTelemetry, SIEM, tracing, logging, and event correlation. API security: OAuth2/OIDC, mTLS, service mesh, gateway security, rate limiting ...

Observability SRE

Location
Greater London, England, United Kingdom
want you to find your spark. Because that’s what drives you to be better, be more and ultimately, be more fulfilled. Job Title: Observability SRE Location: London, UK Employment Type: Fixed term contract (12 months duration) Job type: Onsite Job/Group Overview: SRE within the Group Platform Services …/support position responsible for administering and supporting Production environment as well as engineering reliability into the products/services we support i.e. monitoring & observability platform. The successful candidate will have a vital role in shaping future monitoring strategy and direction. A fantastic opportunity for somebody with 3+ years ...

Principal Site Reliability Engineer, Infrastructure Observability

Location
Greater London, England, United Kingdom
opportunity to grow and make a difference in ways that matter to you. Role Summary In this role as Principal Site Reliability Engineer, Infrastructure Observability you will help formulate, develop, and implement a team of Site Reliability Engineers (SREs) focused on the observability, sustainability, scalability, measurability and recoverability … Proficiency with understanding and explaining incident situations and their recovery plans to prevent recurrence Knowledge/experience driving dashboard standardization across the ecosystem for observability, APM and infrastructure monitoring, and application‐specific logging Knowledge/experience with observability tools such as New Relic, SolarWinds DPA, Elastic Stack, Prometheus, Grafana, Splunk ...

Senior Azure DevOps Engineer

Location
Greater London, England, United Kingdom
position focuses on building and improve containerised environments using Docker and Kubernetes. A highly varied technical environment spanning Azure, Kubernetes, Terraform, CI/CD, observability and automation. What you'll bring Proficiency in at least one programming or scripting language such as Python, Go or Node.js Familiarity with … large-scale systems using tools such as GitLab CI/CD, GitHub Actions, Jenkins or similar You’ll need experience with monitoring, logging and observability technologies such as Prometheus, Grafana, Loki, Open You’ll need experience designing and implementing scalable systems capable of handling high workloads Role details Work model ...

Sr Lead Infrastructure Engineer- Devops/AWS

Location
Auchentibber, Scotland, United Kingdom
machine learning products. As the team takes end-to-end ownership of the platforms it runs, you will build the CI/CD, observability, and incident-management practices that keep those services stable, secure, and performant across international markets. This is a Vice President-level role and an integral part … release automation, and deployment tooling Establishes reliability practices (SLOs, error budgets, runbooks) and leads production incident response and post-incident review Builds and operates observability across the team's AI/ML services (metrics, logging, tracing, alerting) Automates infrastructure provisioning and configuration through infrastructure-as-code Implements operational security, secrets ...

AI Native DevOps Platform Engineer

Hiring Organisation
Sanderson Recruitment
Location
London, United Kingdom
Employment Type
Permanent
within an ambitious, product-led environment. You'll work closely with Product Engineering teams and Technical Leadership to deliver cloud platforms, infrastructure, deployment pipelines, observability and developer tooling, while introducing AI-assisted and agent-driven approaches to improve engineering productivity, reliability and operational efficiency. What You'll Do Design, build … governance and compliance throughout the software delivery lifecycle Optimise Azure environments for performance, scalability, reliability and cost efficiency Support model-serving infrastructure and AI observability capabilities Lead incident response and drive continuous improvement across production environments Work closely with engineering teams delivering microservices and distributed cloud-native applications Required Experience ...

Sr Lead Infrastructure Engineer- Devops/AWS

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
machine learning products. As the team takes end-to-end ownership of the platforms it runs, you will build the CI/CD, observability, and incident-management practices that keep those services stable, secure, and performant across international markets. This is a Vice President-level role and an integral part … pipelines, release automation, and deployment toolingEstablishes reliability practices (SLOs, error budgets, runbooks) and leads production incident response and post-incident reviewBuilds and operates observability across the team's AI/ML services (metrics, logging, tracing, alerting)Automates infrastructure provisioning and configuration through infrastructure-as-codeImplements operational security, secrets management ...

Sr Lead Infrastructure Engineer- Devops/AWS

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
machine learning products. As the team takes end-to-end ownership of the platforms it runs, you will build the CI/CD, observability, and incident-management practices that keep those services stable, secure, and performant across international markets. This is a Vice President-level role and an integral part … release automation, and deployment tooling Establishes reliability practices (SLOs, error budgets, runbooks) and leads production incident response and post-incident review Builds and operates observability across the team's AI/ML services (metrics, logging, tracing, alerting) Automates infrastructure provisioning and configuration through infrastructure-as-code Implements operational security, secrets ...

DevOps Engineer

Location
Greater London, England, United Kingdom
automation, platforms, and tooling that allow development teams to deploy and manage applications reliably and securely. Key areas include IaC, CI/CD, Kubernetes, observability, AWS services, and developer self-service. The engineer will collaborate closely with global cloud, infrastructure, security, and application development teams. As our cloud and platform … pipelines using GitHub Actions and related tooling. Manage and improve Kubernetes platforms, GitOps workflows, Helm charts, and application onboarding. Develop monitoring, logging, alerting, and observability capabilities using Grafana, Prometheus, DataDog, Loki, and CloudWatch. Support cloud networking and connectivity across AWS accounts, regions, on-prem environments, and other cloud platforms ...

Lead Software Engineer

Location
Greater London, England, United Kingdom
stack technical leadership role with responsibility for the platform end-to-end, including user interfaces, APIs, backend services, integrations, data processing, cloud infrastructure, security, observability, and operational support. The role combines software engineering, solution architecture, operational ownership, and team leadership. You will work closely with Product, Operations, Compliance, and Architecture … PRIIP Cloud platform in production. Support the team's responsibility for production incidents, customer issues, operational troubleshooting, and service restoration activities. Drive monitoring, observability, alerting, incident management, capacity planning, and service improvement initiatives. Ensure operational processes, runbooks, support procedures, and resilience capabilities are maintained and continuously improved. Work closely with ...

Senior Consultant Snowflake

Hiring Organisation
Infosys Technologies
Location
London, United Kingdom
Salary
£ 70 K
ecosystem. The ideal candidate will combine deep technical expertise with strong delivery and people leadership capabilities to drive platform reliability, operational excellence, governance, automation, observability, and stakeholder management.The role requires hands-on expertise in at least two platform technologies, with mandatory expertise in either Snowflake or Confluent Kafka. The individual … SLAs, KPIs, operational metrics, and customer commitments.•Proactively manage risks, dependencies, and technical blockers.Drive compliance, security, audit readiness, and cost optimization. •Implement effective monitoring, observability, and service management processes.•Hands on experience•Mentor and guide Snowflake, Kafka, Cloud, DevOps, and Platform engineers.•Provide architectural oversight and technical governance across ...

Platform Engineer (Mid-level)

Location
Wakefield, England, United Kingdom
PowerShell, Python, Bash, or similar technologies. Assist with Infrastructure as Code and platform configuration management. Improve repeatability, documentation, and operational consistency across platform services. Observability, Support & Reliability Support monitoring, alerting, and operational dashboards across platform services. Troubleshoot and resolve infrastructure, cloud, and platform issues. Participate in incident response, root-cause … source control platforms such as Perforce, GitHub, GitLab. Experience supporting TeamCity, Jenkins, GitHub Actions, or similar CI/CD platforms. Experience with monitoring and observability platforms such as Datadog, Grafana, Prometheus, or similar. Experience with Docker, Kubernetes, or container-based platforms. Experience with package management platforms such as Artifactory ...

Systems Engineer I

Location
City Of London, England, United Kingdom
passionate about building reliable cloud infrastructure and automating systems that power products used around the world? Do you enjoy solving operational challenges, improving observability, and collaborating with engineers to deliver resilient, scalable platforms? About our Team Join a collaborative and supportive engineering team powering Scopus, working closely with colleagues across … cause analysis and postmortems Build and maintain automation for deployment, monitoring, and infrastructure management (Infrastructure as Code) Collaborate with development teams to improve system observability (logging, metrics, tracing, alerting) Contribute to capacity planning, performance tuning, and cost optimization efforts Write and maintain runbooks, documentation, and operational procedures Support CI/ ...

Platform Engineer

Location
City Of London, England, United Kingdom
overhead. Support and enhance CI/CD platforms and engineering workflows, including the safe promotion of infrastructure and platform changes. Improve platform reliability, security, observability, performance and operational excellence. Partner with software engineers, quantitative developers, research teams, Technology Operations and security colleagues to deliver secure, scalable and resilient platforms. Education … experience with AWS and cloud‐native infrastructure. Experience with Kubernetes, including Amazon EKS, and Docker or other container runtimes. Experience with monitoring and observability tools such as Grafana, Prometheus, Loki or Datadog. Experience with artifact‐management platforms such as JFrog Artifactory. Knowledge of AWS Batch, AWS Step Functions, AWS Identity ...

Senior TypeScript Engineer

Location
Greater London, England, United Kingdom
reliable solutions. Use AI tools to accelerate coding, debugging, testing, research, and documentation, while validating outputs carefully and applying sound judgment. Strengthen service reliability, observability, and engineering quality by improving monitoring, incident response, testing, and development practices. The candidate 5 to 8 years of professional software engineering experience delivering production … experience supporting or mentoring other engineers within a product engineering team. Familiarity with infrastructure as code, for example CDK or Terraform. Exposure to observability tooling, incident response, or production monitoring practices. Experience working with React when contributing to end‐to‐end product delivery. What’s in it for me? They ...

Senior MLOps Engineer 201043

Hiring Organisation
Harnham - Data & Analytics Recruitment
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£75,000 - £85,000 per annum
deployment, inference, and retraining workflows. Implement infrastructure-as-code solutions using tools such as Terraform, Bicep, or similar technologies. Establish monitoring, logging, alerting, and observability frameworks to ensure reliability and performance. Support scalable deployment across multiple client environments while maintaining security and operational excellence. Collaborate with data scientists, engineers … pipelines using Azure DevOps or similar tools. Knowledge of infrastructure-as-code approaches using Terraform, Bicep, Pulumi, or equivalent technologies. Experience implementing monitoring and observability solutions using tools such as Prometheus, Grafana, or similar. Familiarity with orchestration platforms including Dagster, Airflow, Prefect, or related technologies. Desirable experience includes: Model serving ...

Senior AWS Platform Engineer

Location
United Kingdom
implementation and operational handover. The successful candidate will have experience delivering AWS platform solutions across multiple environments and a strong understanding of automation, security, observability and cloud engineering best practice. Due to the nature of this role successful candidates may need to undergo security clearance. More information on SC Clearance … from discovery through to operational handover. Lead AWS platform implementation activities, coordinating technical dependencies, guiding engineering decisions, promoting best practices across automation, security and observability, and supporting the development of early-career colleagues through mentoring and knowledge sharing. About You Strong hands-on AWS platform engineering experience, with a proven ...

Software Engineering Manager

Hiring Organisation
Halian Technology Limited
Location
Central London, London, United Kingdom
Employment Type
Permanent
making sound architectural and design decisions. Champion software quality, security, performance, and operational excellence. Encourage modern engineering practices, including CI/CD, automated testing, observability, and cloud-native development. Stakeholder Management Build strong relationships with business and technology stakeholders. Communicate progress, risks, and dependencies effectively. Align engineering activities with organisational … would be beneficial: .NET, Java, Python, or Node.js, React Microservices architecture RESTful APIs Kubernetes and Docker AWS, Azure, or GCP CI/CD tooling Observability and monitoring platforms Modern data platforms and event-driven architectures There is a 2 - 3 stage interview process, with interview slots now available with ...

Senior Platform Engineer

Location
Manchester, England, United Kingdom
experience with OpenTelemetry. You don't just look atmetrics; you design the telemetry that allows for deep‐dive distributed tracing and root‐cause analysis. Observability Infrastructure: Hands‐on experience scaling Prometheus and Grafana tohandle high‐cardinality data. Your Skills BS degree in Computer Science, related technical field, or equivalent practical … Platform Engineer. Substantial experience in Java, RUST programming experience Proven experience in building and operating a SaaS product at scale. Proven cloud experience using observability and operating workloads in Kubernetes. Experience with cloud storage S3, NFS, NetApp OnTap, AWS FSX Distributed Systems and programming techniques. Monitoring and metrics infrastructure (OTEL ...

Systems Engineer I

Location
Greater London, England, United Kingdom
passionate about building reliable cloud infrastructure and automating systems that power products used around the world? Do you enjoy solving operational challenges, improving observability, and collaborating with engineers to deliver resilient, scalable platforms? About our Team Join a collaborative and supportive engineering team powering Scopus, working closely with colleagues across … cause analysis and postmortems Build and maintain automation for deployment, monitoring, and infrastructure management (Infrastructure as Code) Collaborate with development teams to improve system observability (logging, metrics, tracing, alerting) Contribute to capacity planning, performance tuning, and cost optimization efforts Write and maintain runbooks, documentation, and operational procedures Support CI/ ...

Lead AI Software Engineer - TRP Labs London

Location
Greater London, England, United Kingdom
Guide the development of reusable AI capabilities and shared platform components across areas such as agent orchestration, tool use, retrieval-based systems, evaluation frameworks, observability, and guardrails.* Champion engineering excellence through strong software design, code quality, automated testing, continuous integration, and continuous delivery practices.* Ensure AI systems are built with … measurable quality, production readiness, operational observability, and appropriate safety controls.* Oversee technical debt and drive continuous improvement across AI platforms, services, and development standards.* Identify and pursue opportunities to apply AI in ways that accelerate workflows, improve decision-making, and create scalable business impact across the firm.* Present and demonstrate ...

Senior Backend Engineer

Hiring Organisation
Inspire People
Location
South West London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
part of a multidisciplinary agile team, you will design, build and run platform services that underpin critical digital products, helping development teams improve observability, monitoring, CI/CD and service resilience. As a Senior Backend Engineer (Site Reliability), you will: * Design, build and operate reliable, secure and scalable cloud platform … services supporting critical digital products. * Develop and maintain platform tooling, automation, observability, monitoring and CI/CD capabilities. * Build software solutions using Python and modern engineering practices. * Write clean, maintainable code and infrastructure-as-code solutions to support service delivery. * Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets ...

Senior Backend Engineer

Location
United Kingdom
part of a multidisciplinary agile team, you will design, build and run platform services that underpin critical digital products, helping development teams improve observability, monitoring, CI/CD and service resilience. As a Senior Backend Engineer (Site Reliability), you will: Design, build and operate reliable, secure and scalable cloud platform … services supporting critical digital products. Develop and maintain platform tooling, automation, observability, monitoring and CI/CD capabilities. Build software solutions using Python and modern engineering practices. Write clean, maintainable code and infrastructure-as-code solutions to support service delivery. Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets ...

Senior Backend Engineer

Hiring Organisation
Inspire People
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
part of a multidisciplinary agile team, you will design, build and run platform services that underpin critical digital products, helping development teams improve observability, monitoring, CI/CD and service resilience. As a Senior Backend Engineer (Site Reliability), you will: * Design, build and operate reliable, secure and scalable cloud platform … services supporting critical digital products. * Develop and maintain platform tooling, automation, observability, monitoring and CI/CD capabilities. * Build software solutions using Python and modern engineering practices. * Write clean, maintainable code and infrastructure-as-code solutions to support service delivery. * Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets ...