101 to 125 of 3,940 Observability Jobs

Site Reliability Engineer - Service Assurance Systems

Location
Greater London, England, United Kingdom
range of systems across the Viasat estate. This data is shared between internal platforms and distributed through Google Cloud Platform (GCP), where it underpins observability and monitoring capabilities that give operations teams real-time insight into service health and performance. To ensure the data is accurate and fit for purpose … SLAs. Manage and oversee application deployments across development, staging and production environments, ensuring smooth and reliable release processes. Monitor application and infrastructure health using observability tools such as Prometheus and AWS CloudWatch; proactively identify and respond to anomalies and performance degradation. Own and maintain CI/CD pipelines, working ...

Cloud Engineer II

Hiring Organisation
HTC Global Services Inc
Location
Dearborn, Michigan, United States
Employment Type
Permanent
Salary
USD Annual
capabilities. This role will work within a cross-functional Agile team to build cloud services, data integrations, user-facing capabilities, APIs, and platform observability solutions. The position involves designing, building, and deploying cloud-based infrastructure and managed services to support scalability, flexibility, and cost efficiency across private and public cloud … cloud-native technologies. Experience developing frontend applications using React and TypeScript. Familiarity with event-driven architecture, asynchronous processing, and data pipelines. Experience implementing platform observability using logs, metrics, traces, dashboards, and alerting. Familiarity with OpenTelemetry or comparable observability standards and tooling. Experience with API specifications and tools such as OpenAPI ...

Head of Cloud

Location
Norwich, England, United Kingdom
operational efficiency. Documentation & Continuous Improvement Champion high standards for technical documentation and process optimization across cloud teams. Promote a culture of knowledge sharing, observability, and proactive problem-solving. Success Metrics (KPIs) Team Growth: High-performing cloud engineers recruited, onboarded, and retained. System Evolution: Delivery excellence in cross-team technical initiatives … Stack Cloud: AWS (expert), Kubernetes (EKS) IaC & Orchestration: Terraform, Helm, Terragrunt Languages: Go, Node.js, Python (automation/tooling) CI/CD: GitHub Actions, ArgoCD Observability: Prometheus, Grafana, OpenTelemetry What We’re Looking For Proven Leadership: Experience managing and scaling high-performing engineering teams. Cloud Expertise: Deep hands-on experience architecting ...

Senior Full Stack/Native Cloud Engineer (Vice President)

Hiring Organisation
Jefferies Financial Group
Location
London, UK
Employment Type
Full-time
technical design and architecture through to production deployment, monitoring, support, and continuous improvement. Champion engineering best practices across code quality, testing, CI/CD, observability, maintainability, and operational resilience. Mentor and guide junior and mid-level engineers, setting high technical standards through code reviews, design reviews, and hands-on technical … reliable, maintainable, and responsive UI delivery. Strong understanding of software engineering best practices, including automated testing, clean code, code reviews, CI/CD, observability, and secure development practices. Experience deploying and operating production-grade systems in cloud or hybrid-cloud environments. Demonstrated ability to lead technical delivery, mentor engineers ...

Senior DevOps Engineer

Location
Belfast City District, Northern Ireland, United Kingdom
involved in software releases by fully automating delivery pipelines that includes testing Improve production stability and resiliency through process improvement and adoption of better observability tools Support application teams in adopting a consistent set of tools, practices and technology to support an agile delivery model, including automation and incremental delivery … principles Knowledge of third-party automation tools and technologies with relative pros and cons for all stages of SDLC, CI/CD pipelines and observability Platforms; Windows Server, Amazon Linux, RHEL, Ubuntu Proficiency in at least one of the following scripting languages; Python, GO, PowerShell, Bash, Groovy Programming language with ...

Sr. Database Architect

Location
United Kingdom
system performance across deployments Production Operations & Data Reliability Experience defining and maintaining production database processes, including monitoring, alerting, and incident response Familiarity with observability tools and practices (logging, metrics, tracing) Strong understanding of SLAs, SLOs, and data reliability best practices Tools & Platforms Experience with AWS data technologies (Glue, Kinesis, Lambda … data systems Experience with containerization (Docker, Kubernetes) Knowledge of encryption, anonymization, and tokenization Experience with open table formats and data catalogs Familiarity with data observability tools (e.g., Monte Carlo, Datadog, Prometheus) Who You Are (Soft Skills): Detail-oriented, with a strong data quality mindset Strong problem-solving and troubleshooting skills ...

AWS Cloud Platform Engineer

Hiring Organisation
Pontoon
Location
Dublin, City of Dublin, Republic of Ireland
Employment Type
Contract
Contract Rate
£474/day
Support containerized workloads using Docker and Kubernetes/EKS. Implement and maintain cloud security controls, IAM policies, encryption solutions, and governance frameworks. Contribute to observability and monitoring strategies using AWS-native and enterprise monitoring tools. Optimize cloud performance, resilience, and cost efficiency across large-scale AWS environments. Provide technical leadership … Experience building and supporting CI/CD pipelines and GitOps environments. Proficiency in Python, Go, Shell, or other scripting languages. Experience with monitoring and observability tools such as Grafana, Prometheus, Splunk, Dynatrace, and AWS-native solutions. Excellent communication and stakeholder management skills. Preferred Experience Financial Services, Banking, FinTech, or other ...

Platform Staff Engineer - UK

Hiring Organisation
DISCO
Location
London, UK
Employment Type
Full-time
flexible, self-service artifact generation. Researches and evaluates file processing software based on fidelity, reliability, and performance criteria. Models file processing outcomes within an observability framework to provide insights for engineering teams and business leaders. Develops engineering systems that facilitate rapid and automatic evaluation of file processors. Data Management: Crafts … projects, showing an ability to work effectively both independently and as part of a team. Demonstrated expertise in designing, implementing, and maintaining (through operational observability) high-availability, high-performance, distributed data processing systems. Experience with gRPC and Protocol Buffers for efficient, language-agnostic service-to-service communication. Proven ability ...

Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President

Location
Greater London, England, United Kingdom
deploying Large Language Model (LLM) orchestration frameworks (e.g., LangChain, Temporal, or custom agentic loops) to coordinate multi-step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root-cause analysis, anomaly detection, and semantic … technical decisions across teams. Experience working in regulated industries is a plus. Preferred Qualifications Experience building self-service platforms for development teams. Familiarity with observability and monitoring tools (Prometheus, Grafana, Datadog, CloudWatch). Background in financial services or other highly regulated environments. About Goldman Sachs At Goldman Sachs, we commit ...

Vice President, Site Reliability Engineering

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
implement, and continuously improve Service Level Indicators, Service Level Objectives, and service health measures aligned to operational and business priorities. Build and optimize monitoring, observability, and alerting capabilities using tools such as Prometheus, Grafana, AppDynamics, and Splunk. Apply AIOps capabilities to improve event correlation, anomaly detection, root cause analysis, predictive … enterprise or production environments. Demonstrated ability to define and operationalize SLIs, SLOs, dashboards, alerts, and health indicators. Hands-on experience with enterprise monitoring and observability platforms including Prometheus, Grafana, AppDynamics, and Splunk. Strong troubleshooting, analytical, and problem-solving skills in complex distributed or production environments. Strong verbal and written communication ...

Senior Software Engineer

Location
Reading, England, United Kingdom
solution design for complex cross‐cutting services Make pragmatic architectural decisions balancing scalability, security and maintainability Improve performance, reliability and observability across shared services Define engineering standards and patterns adopted by other teams Lead by example through high‐quality, production‐ready code Partner with product squads to understand their needs … patterns, access control and service integration Mentor engineers and raise technical capability across the organisation Drive CI/CD maturity and deployment confidence Embed observability (metrics, tracing, logging) as first‐class concerns Perform thoughtful code reviews and provide constructive feedback Continuously improve team practices and technical standards Required Strong Java ...

Senior AWS DevOps Engineer

Location
Greater London, England, United Kingdom
services such as ECS and EKS Work with architects and security specialists on cloud architecture, networking, identity, security and integration requirements Implement effective monitoring, observability, logging and alerting to support resilient services Provide technical leadership during complex troubleshooting and resolution of infrastructure and deployment issues Engage with clients and technical … operating containerised workloads in cloud environments Good understanding of DevSecOps, OWASP principles and security‐by‐design Experience designing and implementing monitoring, logging and observability solutions Strong scripting skills using Python, Bash, PowerShell or similar Experience making technical design decisions and providing engineering leadership within multidisciplinary teams Strong troubleshooting and problem ...

Senior AWS DevOps Engineer

Location
Manchester, England, United Kingdom
services such as ECS and EKS Work with architects and security specialists on cloud architecture, networking, identity, security and integration requirements Implement effective monitoring, observability, logging and alerting to support resilient services Provide technical leadership during complex troubleshooting and resolution of infrastructure and deployment issues Engage with clients and technical … operating containerised workloads in cloud environments Good understanding of DevSecOps, OWASP principles and security‐by‐design Experience designing and implementing monitoring, logging and observability solutions Strong scripting skills using Python, Bash, PowerShell or similar Experience making technical design decisions and providing engineering leadership within multidisciplinary teams Strong troubleshooting and problem ...

Vice President, Site Reliability Engineering

Hiring Organisation
The Bank of New York Mellon
Location
London, UK
Employment Type
Full-time
implement, and continuously improve Service Level Indicators, Service Level Objectives, and service health measures aligned to operational and business priorities. Build and optimize monitoring, observability, and alerting capabilities using tools such as Prometheus, Grafana, AppDynamics, and Splunk. Apply AIOps capabilities to improve event correlation, anomaly detection, root cause analysis, predictive … enterprise or production environments. Demonstrated ability to define and operationalize SLIs, SLOs, dashboards, alerts, and health indicators. Hands-on experience with enterprise monitoring and observability platforms including Prometheus, Grafana, AppDynamics, and Splunk. Strong troubleshooting, analytical, and problem-solving skills in complex distributed or production environments. Strong verbal and written communication ...

Senior Engineer

Location
Greater London, England, United Kingdom
roadmap. Design, build and operate a secure, resilient and scalable Azure infrastructure estate, including the associated CI/CD, infrastructure-as-code, container and observability capabilities that support application delivery. Establish and evolve engineering standards, reference architectures, guardrails and reusable platform services, promoting consistency, automation, security, maintainability and operational excellence. … cloud and/or on-premise infrastructure, including the development of reusable, maintainable infrastructure-as-code patterns. Experience implementing or operating monitoring, logging and observability capabilities using tools such as Splunk, Prometheus, Grafana or equivalent, with the ability to use operational data to improve reliability and performance. Strong scripting ...

Senior Software Engineer

Hiring Organisation
Reonomy
Location
Manchester, Greater Manchester, United Kingdom
Salary
£ 70 K
workflows using technologies such as Glue, Spark/PySpark, Kafka, or similar platforms.Implement and maintain Infrastructure as Code (IaC) using Terraform or equivalent tools.Drive observability, monitoring, logging, and operational readiness across software platforms and services.Key Qualifications:Bachelor's degree in Computer Science, Software Engineering, Engineering, or a related field … model performance validation.Experience leveraging AI-assisted development practices and transforming manual engineering workflows into automated solutions.Experience with infrastructure as code, CI/CD pipelines, observability tools, and operational best practices.Excellent problem-solving, collaboration, and communication skills with the ability to influence technical decisions across teams.Unlock your Altus Experience ...

Senior Software Engineer

Hiring Organisation
Reonomy
Location
Manchester, UK
Employment Type
Full-time
technologies such as Glue, Spark/PySpark, Kafka, or similar platforms. Implement and maintain Infrastructure as Code (IaC) using Terraform or equivalent tools. Drive observability, monitoring, logging, and operational readiness across software platforms and services. Key Qualifications: Bachelor's degree in Computer Science, Software Engineering, Engineering, or a related field … validation. Experience leveraging AI-assisted development practices and transforming manual engineering workflows into automated solutions. Experience with infrastructure as code, CI/CD pipelines, observability tools, and operational best practices. Excellent problem-solving, collaboration, and communication skills with the ability to influence technical decisions across teams. Unlock your Altus Experience ...

Senior Software Engineer (Full Stack & DevOps)

Location
Rotherham, England, United Kingdom
harder, not easier. This role exists to close that gap. You’ll build across the stack, and you’ll own the testing, deployment and observability infrastructure that lets us ship frequently and with confidence. You’ll work closely with stakeholders across marketing, studio, internal systems and accounts, translating commercial needs … code before it ships. Work with Infrastructure as Code tools (e.g. Terraform, CloudFormation) to build and manage cloud environments, including containerised workloads. Build the observability we need — logging, monitoring, tracing and alerting — so we spot problems before customers do. Improve release strategy: feature flags, staged rollouts, safe database migrations ...

Senior Software Engineer

Location
Manchester, England, United Kingdom
technologies such as Glue, Spark/PySpark, Kafka, or similar platforms. Implement and maintain Infrastructure as Code (IaC) using Terraform or equivalent tools. Drive observability, monitoring, logging, and operational readiness across software platforms and services. Key Qualifications Bachelor's degree in Computer Science, Software Engineering, Engineering, or a related field … validation. Experience leveraging AI‐assisted development practices and transforming manual engineering workflows into automated solutions. Experience with infrastructure as code, CI/CD pipelines, observability tools, and operational best practices. Excellent problem‐solving, collaboration, and communication skills with the ability to influence technical decisions across teams. Unlock your Altus Experience ...

Data Technical Lead

Location
Greater London, England, United Kingdom
analyticaland domain-oriented modelling Data governance:Catalogue, metadata, lineage, quality, security,privacyand access controls Data products:Reusable, discoverable and well-governed dataproductsand marketplaces DataOpsand observability:Testing, monitoring, operationalcontrolsand platform reliability Platform engineering:Terraform, CloudFormation, Azure Bicep and infrastructure-as-code CI/CD:GitHub Actions, Azure DevOps,Jenkinsand equivalent tooling … serving. Applying strong software engineering practices to data platforms, including testing, CI/CDand infrastructure-as-code. Establishing effective approaches to data quality, observability, governance,metadataand security. You can lead data platform modernisation and migration, including coexistence,cutoverand decommissioning. Making pragmatic technology choices and understanding the trade-offs between platform ...

AI Engineer

Hiring Organisation
Formula Recruitment Limited
Location
London, UK
Employment Type
Full-time
with cloud platforms (AWS, GCP or Azure) and infrastructure-as-code such as TerraformHands-on with DevOps practices (CI/CD, Docker, Kubernetes) and observability tools like Prometheus, Grafana or DatadogExperience with distributed systems, scaling, and both SQL and NoSQL datastoresAgentic workflow experience (autonomous or multi-agent systems ...

Software Engineer – Python/AWS/Terraform (SC or DV Cleared)

Location
Greater London, England, United Kingdom
event-driven architectures Desirable experience Stronger Python backend development using FastAPI or Flask. Kubernetes or Amazon EKS. DevSecOps and secure-by-design engineering. Observability, monitoring and automated alerting. < REST APIs and distributed systems. Ansible, CloudFormation or AWS CDK. AI or LLM integration. Previous delivery within government, defence or another highly ...

Senior Dev/ML Ops Engineer

Location
Greater London, England, United Kingdom
deployment to monitoring and continuous improvement. Build and maintain robust CI/CD pipelines for both software and ML workflows. Ensure reliability, scalability, observability, and security of production systems and ML infrastructure. Automate deployment, orchestration, and environment management using modern DevOps tooling. Collaborate closely with software engineers, data scientists ...

Head of Cloud

Hiring Organisation
Epos Now
Location
Norwich, Norfolk, United Kingdom
Salary
£ 70 K
improve velocity, reliability, and operational efficiency.Documentation & Continuous ImprovementChampion high standards for technical documentation and process optimization across cloud teams.Promote a culture of knowledge sharing, observability, and proactive problem-solving.Success Metrics (KPIs)Team Growth: High-performing cloud engineers recruited, onboarded, and retained.System Evolution: Delivery excellence in cross-team technical initiatives with ...

Head of Cloud

Hiring Organisation
Epos Now
Location
Norwich, Norfolk, UK
Employment Type
Full-time
reliability, and operational efficiency. Documentation & Continuous ImprovementChampion high standards for technical documentation and process optimization across cloud teams. Promote a culture of knowledge sharing, observability, and proactive problem-solving. Success Metrics (KPIs)Team Growth: High-performing cloud engineers recruited, onboarded, and retained. System Evolution: Delivery excellence in cross-team technical ...