276 to 300 of 4,102 Observability Jobs

Senior Backend Engineer

Hiring Organisation
Inspire People
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
part of a multidisciplinary agile team, you will design, build and run platform services that underpin critical digital products, helping development teams improve observability, monitoring, CI/CD and service resilience. As a Senior Backend Engineer (Site Reliability), you will: * Design, build and operate reliable, secure and scalable cloud platform … services supporting critical digital products. * Develop and maintain platform tooling, automation, observability, monitoring and CI/CD capabilities. * Build software solutions using Python and modern engineering practices. * Write clean, maintainable code and infrastructure-as-code solutions to support service delivery. * Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets ...

Senior Backend Engineer

Hiring Organisation
Inspire People
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
part of a multidisciplinary agile team, you will design, build and run platform services that underpin critical digital products, helping development teams improve observability, monitoring, CI/CD and service resilience. As a Senior Backend Engineer (Site Reliability), you will: * Design, build and operate reliable, secure and scalable cloud platform … services supporting critical digital products. * Develop and maintain platform tooling, automation, observability, monitoring and CI/CD capabilities. * Build software solutions using Python and modern engineering practices. * Write clean, maintainable code and infrastructure-as-code solutions to support service delivery. * Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets ...

Senior Backend Engineer

Hiring Organisation
Inspire People
Location
Cardiff, South Glamorgan, Wales, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
part of a multidisciplinary agile team, you will design, build and run platform services that underpin critical digital products, helping development teams improve observability, monitoring, CI/CD and service resilience. As a Senior Backend Engineer (Site Reliability), you will: * Design, build and operate reliable, secure and scalable cloud platform … services supporting critical digital products. * Develop and maintain platform tooling, automation, observability, monitoring and CI/CD capabilities. * Build software solutions using Python and modern engineering practices. * Write clean, maintainable code and infrastructure-as-code solutions to support service delivery. * Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets ...

Senior Backend Engineer

Hiring Organisation
Inspire People
Location
Belfast, County Antrim, Northern Ireland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
part of a multidisciplinary agile team, you will design, build and run platform services that underpin critical digital products, helping development teams improve observability, monitoring, CI/CD and service resilience. As a Senior Backend Engineer (Site Reliability), you will: * Design, build and operate reliable, secure and scalable cloud platform … services supporting critical digital products. * Develop and maintain platform tooling, automation, observability, monitoring and CI/CD capabilities. * Build software solutions using Python and modern engineering practices. * Write clean, maintainable code and infrastructure-as-code solutions to support service delivery. * Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets ...

Senior Backend Engineer

Hiring Organisation
Inspire People
Location
Darlington, County Durham, North East, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
part of a multidisciplinary agile team, you will design, build and run platform services that underpin critical digital products, helping development teams improve observability, monitoring, CI/CD and service resilience. As a Senior Backend Engineer (Site Reliability), you will: * Design, build and operate reliable, secure and scalable cloud platform … services supporting critical digital products. * Develop and maintain platform tooling, automation, observability, monitoring and CI/CD capabilities. * Build software solutions using Python and modern engineering practices. * Write clean, maintainable code and infrastructure-as-code solutions to support service delivery. * Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets ...

Senior Backend Developer

Hiring Organisation
Inspire People
Location
Darlington, County Durham, North East, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
part of a multidisciplinary agile team, you will design, build and run platform services that underpin critical digital products, helping development teams improve observability, monitoring, CI/CD and service resilience. As a Senior Backend Engineer (Site Reliability), you will: * Design, build and operate reliable, secure and scalable cloud platform … services supporting critical digital products. * Develop and maintain platform tooling, automation, observability, monitoring and CI/CD capabilities. * Build software solutions using Python and modern engineering practices. * Write clean, maintainable code and infrastructure-as-code solutions to support service delivery. * Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets ...

Senior Cloud Engineer - Azure DevOps

Location
Manchester, England, United Kingdom
into deployable cloud platforms and services. Support the implementation of Azure networking, identity, security, compute, storage and platform services. Implement monitoring, logging, alerting and observability across cloud environments. Troubleshoot complex platform, deployment and infrastructure issues. Improve reliability, scalability and operational performance through automation and engineering best practice. Support containerised workloads … Experience with GitOps approaches and tooling. Experience with policy‐as‐code and automated governance. Experience with Azure Monitor, Log Analytics, Application Insights or equivalent observability tooling. Experience implementing automated security scanning within CI/CD pipelines. Experience working with Microsoft Entra ID. Experience with configuration management and automation tooling. Experience ...

Full Stack Lead Architect, Vice President

Hiring Organisation
Citigroup
Location
Belfast, UK
Employment Type
Full-time
frameworks, AI Agents, and intelligent automation solutions. Architect event-driven systems using Kafka, Solace, and distributed caching technologies. Establish engineering standards, SDLC best practices, observability, resiliency, and security controls. Collaborate with Architecture, Product, Data Engineering, Infrastructure, Security, and Business teams globally. Lead technical reviews, solution governance, and critical production issue …/Financial Services experience. Experience with Data Lake, Lakehouse, and Analytics platforms. Knowledge of OAuth2, JWT, Zero Trust, and API Security standards. Experience with observability platforms such as Grafana, Dynatrace, Splunk, Prometheus, or OpenTelemetry. Leadership ExpectationsProvide technical leadership across multiple teams and strategic initiatives. Drive architecture decisions and engineering excellence. ...

Staff Software Engineer - Commercial Planning

Location
Greater London, England, United Kingdom
modernisation plans. Your expertise will help us on this journey, creating solutions for the business that are robust, scalable and secure, with strong observability, meaningful metrics and best-in-class engineering practices. Working closely with M&S teams and strategic partners, you’ll play a key role in shaping … Mentor and coach engineers at all levels, sharing knowledge and helping develop technical capability across the wider engineering community. Drive operational excellence through monitoring, observability, alerting and incident management practices, ensuring learnings from production environments are fed back into development. Stay informed of emerging technologies, AI capabilities and engineering trends ...

Senior Site Reliability Engineer

Hiring Organisation
Brevan Howard
Location
London, UK
Employment Type
Full-time
operational efficiency. Infrastructure-as-Code (IaC): Own and evolve our declarative infrastructure using Terraform for cloud resources and Helm for Kubernetes application deployment. Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification. Reliability & Performance: Define, measure, and enforce Service … major programming language, preferably Python, for automation and tool development. Tooling & ConceptsCI/CD: Experience setting up and maintaining modern CI/CD pipelines. Observability: Practical experience implementing and managing monitoring and logging tools. Networking: Solid understanding of TCP/IP, load balancing, DNS, and cloud-native networking within Kubernetes. ...

Senior Platform Engineer Cheltenham

Location
Cheltenham, England, United Kingdom
delivery pipelines. Work across greenfield and brownfield platform engineering projects. Develop infrastructure-as-code solutions using modern tooling and engineering practices. Improve platform reliability, observability and operational maturity through automation and engineering excellence. Work closely with developers, architects and security teams to understand challenges and deliver pragmatic solutions. Champion platform … Experience with container technologies such as Docker and Kubernetes. Experience with cloud platforms such as AWS, Azure or GCP. Experience implementing monitoring, logging and observability solutions. Ability to work effectively across engineering, security and customer teams. Strong problem-solving skills with the ability to operate in complex and ambiguous environments. ...

Software Engineer II (AWS Infrastructure)

Location
Auchentibber, Scotland, United Kingdom
Contribute to the build and maintenance of CI/CD pipelines (e.g., Jenkins) that apply infrastructure changes safely and reliably Instrument the estate for observability using tools such as Datadog, CloudWatch, and Dynatrace, and use telemetry insights to support improvements to infrastructure hygiene Uses enterprise-authorized AI capabilities within … supporting cluster operations Practical experience working with cloud infrastructure on AWS in a production or near-production environment Familiarity with production readiness practices including observability (metrics, tracing, logging) and incident management in distributed systems Working knowledge of using enterprise-authorized AI capabilities within the work environment to support software engineering ...

Solution Architect

Location
Greater London, England, United Kingdom
Kubernetes, or GKE. Experience with data formats such as JSON, CSV, Avro, and Parquet. Working knowledge of Git, automated testing, CI/CD pipelines, observability, and release management. Ability to design solutions for scalability, resilience, security, performance, operability, and cost efficiency. Preferred Skills Experience with Cloud Composer or Apache Airflow … Kubernetes, or GKE. Experience with data formats such as JSON, CSV, Avro, and Parquet. Working knowledge of Git, automated testing, CI/CD pipelines, observability, and release management. Ability to design solutions for scalability, resilience, security, performance, operability, and cost efficiency. Preferred Skills Experience with Cloud Composer or Apache Airflow ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, United Kingdom
Salary
£ 70 K
across the practice.Partner with sales and pre-sales to scope AI solutions, size infrastructure, and validate technical feasibility of proposed architectures.Define automation, orchestration, and observability standards across the AI stack, from GPU cluster provisioning through to model monitoring in production.Architect integration points connecting AI platforms to existing enterprise networks, third … automated model testing, validation gates, and promotion pipelines (continuous training/continuous delivery) that move models safely from experimentation to production.Infrastructure & GPU observability: NVIDIA DCGM, Prometheus/Grafana, and related telemetry stacks for GPU utilization, thermal, and cluster health monitoring.Model & LLM observability: production model performance monitoring, data/concept drift ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, UK
Employment Type
Full-time
practice. Partner with sales and pre-sales to scope AI solutions, size infrastructure, and validate technical feasibility of proposed architectures. Define automation, orchestration, and observability standards across the AI stack, from GPU cluster provisioning through to model monitoring in production. Architect integration points connecting AI platforms to existing enterprise networks … automated model testing, validation gates, and promotion pipelines (continuous training/continuous delivery) that move models safely from experimentation to production. Infrastructure & GPU observability: NVIDIA DCGM, Prometheus/Grafana, and related telemetry stacks for GPU utilization, thermal, and cluster health monitoring. Model & LLM observability: production model performance monitoring, data/ ...

Senior Back-End Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, UK
dependencies, and delivery challenges. Conduct code reviews and design reviews to ensure high engineering standards across teams. Define and improve best practices for testing, observability, performance, reliability, and operational excellence. Investigate complex production issues and drive continuous improvement through root-cause analysis. Ensure solutions meet security, compliance, reliability, and performance … GitOps practices. Experience operating containerized workloads on Kubernetes. Exposure to front-end technologies such as TypeScript, Vue.js, or modern JavaScript frameworks. Experience implementing monitoring, observability, and operational tooling. What success looks like Delivering scalable, secure, and resilient backend services that support our Ignite's products and customers. Driving strategic architectural ...

Engineering Manager

Location
Greater London, England, United Kingdom
features and manage the dates, by communicating and delivering on time. Conduct interviews on the hiring processes. Drive CI/CD, test automation, and observability practices across the team. Collaborate with product managers to align priorities and solutions. Observe and enforce the standards set by the Architects. Provide Level … Engineering, or related field. Hands-on expertise with AWS architecture, serverless services, EKS computing, and event-driven design. Experience with CI/CD systems, observability, and infrastructure-as-code (e.g., Terraform). Fluency in one or more backend languages (Java, Python, Node.js). Fluency in SQL (MySQL, Oracle, PostgreSQL ...

Senior Data Platform Engineer - Data Enablement

Location
Greater London, England, United Kingdom
engineering.depop.com/What You’ll Do Pave a path for data as product: Champion data as a first‐class citizen by introducing robust data observability and governance tooling into the platform Software engineering: Develop microservices, libraries, data pipelines Technical design: implement and evolve platform services that enable teams to work … automation‐first mindset Experience delivering data compliance & privacy solutions, ensuring that we uphold data subject rights for our customers Experience introducing a data governance & observability stack enabling rich data lineage, data contracts, SLA/SLO, tagging and data quality monitoring capabilities both on our own platform but also for data ...

Lead Java Developer (London)

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
automation, security and resilience. Champion effective CI/CD practices and continuous improvement across the development lifecycle. Implement and promote effective monitoring, logging and observability practices. Investigate and resolve production issues, taking ownership through to resolution and ensuring lessons are incorporated into future development. Collaborate with engineers, architects, product teams … environments. AWS cloud services and cloud-native application development. Terraform and Infrastructure as Code. CI/CD tooling and modern DevOps practices. Monitoring and observability tools, such as Grafana. Automated testing and engineering quality practices. Why join our client? We have three values: wonder, share, and delight. These values inform ...

DevOps Engineer

Location
Leicester, England, United Kingdom
Azure App Services, Azure Functions, Azure SQL, Azure Front Door, Azure Storage, Service Bus, Key Vault and related platform services. Implement monitoring, alerting and observability solutions using tools such as Azure Monitor and Application Insights. Support incident investigation, troubleshooting and root cause analysis to minimise disruption and improve platform reliability. … Code concepts and environment management practices. Understanding of cloud architecture, networking, application security and identity management principles. Experience with monitoring, logging, alerting and observability platforms. Strong troubleshooting, problem-solving and incident management skills. Experience collaborating with software engineering, QA and technical stakeholders. Excellent communication and stakeholder management skills. A proactive ...

Lead Java Developer (London)

Location
Greater London, England, United Kingdom
automation, security and resilience. Champion effective CI/CD practices and continuous improvement across the development lifecycle. Implement and promote effective monitoring, logging and observability practices. Investigate and resolve production issues, taking ownership through to resolution and ensuring lessons are incorporated into future development. Collaborate with engineers, architects, product teams … environments. AWS cloud services and cloud-native application development. Terraform and Infrastructure as Code. CI/CD tooling and modern DevOps practices. Monitoring and observability tools, such as Grafana. Automated testing and engineering quality practices. Why join AND Digital? We have three values: wonder, share, and delight. These values inform ...

Principal AI Engineer - Hybrid

Hiring Organisation
Genesis10
Location
Columbus, Ohio, United States
Employment Type
Permanent
Salary
USD Hourly
adoption, support, and continuous improvement Establish and promote responsible AI, privacy, access control, governance, cost management, and model-risk practices Lead the implementation of observability capabilities for logging, metrics, token utilization, tracing, and operational health on GCP-hosted services Partner with product management, security, architecture, and lines of business … Experience designing and integrating secure REST APIs and OAuth 2.0/OpenID Connect authorization flows Strong understanding of software development lifecycle, CI/CD, observability, reliability, security, privacy, and responsible AI practices Desired Skills: Hands-on experience using Claude Code and GitHub Copilot in an engineering workflow Experience building ...

Software Engineer III - AI/ML Platform Reliability

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
enhance the reliability and scalability of AI/ML platforms and applications to accommodate fast-growing demands. Own NFRs and develop tooling for observability, security, resilience, infrastructure management and operations excellence. Build and maintain scalable infrastructure to support the deployment and operation of large-scale AI platforms and apps. Build … architecture. Experience building large scale infrastructure and and cloud-native delivery practice in Google Cloud, AWS, or Azure and Terraform. Extensive experience implementing advanced observability using tools like Open Telemetry, Dynatrace, Grafana, and/or cloud-native services. Systematic problem-solving and troubleshooting skills in a complex system. Hands ...

Lead Platform Operations Engineer

Location
Greater London, England, United Kingdom
maintain platform standards, patterns, and best practices Own Platform Reliability, Security & Performance Lead incident response, root cause analysis, and platform improvements Implement robust monitoring, observability, and alerting strategies Drive security improvements aligned to ISO27001, SOC2, and modern SDLC practices Ensure strong governance across infrastructure, applications, and data Deliver Scalable & Secure … Experience implementing security tooling (SAST, DAST, container scanning, WAF) Strong knowledge of cloud security, encryption, TLS/SSL, certificates, and access control Experience with observability, monitoring, and alerting tools Security & Compliance Practical experience implementing ISO27001 and SOC2 controls Knowledge of OWASP methodologies and secure development lifecycle practices Experience with vulnerability ...

Sr. Technical Lead / Architect

Location
Greater London, England, United Kingdom
integration patterns, messaging, and event-driven architectures; Produce high-level and low-level design documents; Ensure applications meet non-functional requirements including availability, reliability, observability, and security; Drive API-first architecture and reusable component design. Software Development Lead full-stack application development using: Python, SQL, AWS, React JS; Contribute … implement cloud-native architectures; Collaborate with DevOps teams to automate deployments and infrastructure; Define CI/CD pipelines and release strategies; Improve application observability using logging and monitoring tools; Ensure infrastructure scalability and operational excellence. Quality & Security Promote secure coding practices; Ensure applications comply with security and compliance standards; Drive ...