51 to 75 of 3,940 Observability Jobs

Sr. Site Reliability Engineer/ SWE

Location
Ham, England, United Kingdom
platform strategy. In this role, you’ll ensure our development platform and tools let engineers focus on innovation instead of infrastructure. You’ll promote observability best practices and automate resolution of recurring issues, working closely with software engineering teams to support security, availability, and performance. Responsibilities include triaging issues, collaborating … implement, and maintain systems for high availability, scalability, and performance. Monitor and improve application reliability through proactive measures and incident response. Develop and maintain observability solutions (metrics, logging, tracing). Participate in on‐call rotations and drive root cause analysis for incidents. Collaboration & Continuous Improvement Partner with engineering teams ...

Sr. Site Reliability Engineer/ SWE

Hiring Organisation
Visa
Location
Basingstoke, Hampshire, UK
Employment Type
Full-time
platform strategy. In this role, you'll ensure our development platform and tools let engineers focus on innovation instead of infrastructure. You'll promote observability best practices and automate resolution of recurring issues, working closely with software engineering teams to support security, availability, and performance. Responsibilities include triaging issues, collaborating … implement, and maintain systems for high availability, scalability, and performance. Monitor and improve application reliability through proactive measures and incident response. Develop and maintain observability solutions (metrics, logging, tracing). Participate in on-call rotations and drive root cause analysis for incidents. Collaboration & Continuous Improvement Partner with engineering teams ...

Full Stack Lead Architect, Vice President

Location
Belfast City District, Northern Ireland, United Kingdom
frameworks, AI Agents, and intelligent automation solutions. Architect event-driven systems using Kafka, Solace, and distributed caching technologies. Establish engineering standards, SDLC best practices, observability, resiliency, and security controls. Collaborate with Architecture, Product, Data Engineering, Infrastructure, Security, and Business teams globally. Lead technical reviews, solution governance, and critical production issue …/Financial Services experience. Experience with Data Lake, Lakehouse, and Analytics platforms. Knowledge of OAuth2, JWT, Zero Trust, and API Security standards. Experience with observability platforms such as Grafana, Dynatrace, Splunk, Prometheus, or OpenTelemetry. Leadership Expectations Provide technical leadership across multiple teams and strategic initiatives. Drive architecture decisions and engineering ...

Site Reliability Engineer III

Location
Belfast City District, Northern Ireland, United Kingdom
Service Discovery (Consul, Vault), and Data Distribution (SFTP/JScape)—to Google Cloud Platform. Manage cluster lifecycles, data replication, RBAC, and workload placement. Observability & Monitoring Fabric: Design, scale, and maintain our observability backbone using tools like OpenTelemetry, Splunk, Prometheus, and Grafana. Establish and continuously improve metrics, logs, alerting strategies, SLIs … Strategic communication skills to translate technical requirements for cross-functional teams, coupled with an eagerness to learn independently and collaboratively. Preferred Qualifications/Desirable Observability Stack: Hands-on experience with telemetry tools such as OpenTelemetry, Splunk, Prometheus, and Grafana. Agile Integration: Comfort working within Agile frameworks and collaborative software development ...

Lead DevOps Engineer

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
cloud infrastructure, delivery platforms, and operational capabilities. You will remain hands-on across the engineering lifecycle, from architecture and infrastructure design through deployment, observability, incident response, and continuous improvement. We expect you to operate with a high degree of autonomy, make strategic and architectural decisions within your area, and resolve … teams to productionize AI solutions and ensure services are ready to operate reliably at scale. Establish engineering standards and reusable patterns for infrastructure, security, observability, resilience, documentation, and operational readiness. Lead architectural decisions and evaluate trade-offs across reliability, security, scalability, performance, cost, and maintainability. Take ownership of operational risks ...

Lead DevOps Engineer

Location
Greater London, England, United Kingdom
cloud infrastructure, delivery platforms, and operational capabilities. You will remain hands-on across the engineering lifecycle, from architecture and infrastructure design through deployment, observability, incident response, and continuous improvement.We expect you to operate with a high degree of autonomy, make strategic and architectural decisions within your area, and resolve complex … teams to productionize AI solutions and ensure services are ready to operate reliably at scale.* Establish engineering standards and reusable patterns for infrastructure, security, observability, resilience, documentation, and operational readiness.* Lead architectural decisions and evaluate trade-offs across reliability, security, scalability, performance, cost, and maintainability.* Take ownership of operational risks ...

Lead Cloud Platform Engineer (Kubernetes) - Remote

Location
United Kingdom
managed platform services, including capacity planning, performance tuning, cost optimisation, patching, and lifecycle management Partner with software engineering teams to support application deployment, troubleshooting, observability, and platform adoption Monitor platform health and respond to incidents, conducting root cause analysis and implementing preventative improvements Identify technical debt and contribute to platform …/CD principles and experience building and maintaining automated delivery pipelines Experience working with container technologies and cloud‐native architectures Knowledge of observability, monitoring, logging, and incident management practices Strong troubleshooting and problem‐solving skills across infrastructure, platform, and application layers Experience supporting software development teams in deploying and operating ...

DevOps Engineer, Blockchain Infra (Fully Remote)

Hiring Organisation
Binance
Location
London, United Kingdom
Salary
£ 70 K
/CD pipelines for application and infrastructure deployment.Deploy and operate middleware platforms such as Kafka, Redis and NGINX.Automate operational tasks using Golang and Python.Improve observability through monitoring, logging, alerting, and distributed tracing.Ensure platform reliability, scalability, security, and disaster recovery.Troubleshoot production incidents and perform root cause analysis.Work closely with software engineers … Kafka/Redis/NGINXStrong scripting and programming skills in: Golang/PythonExperience with Git, GitOps workflows, and CI/CD platforms.Strong understanding of observability tools such as Prometheus, OpenTelemetry.Familiarity with container technologies including Docker and Kubernetes. Strong troubleshooting and problem-solving skills.Preferred QualificationsExperience operating blockchain infrastructure or Web3 platforms.Experience ...

Senior DevOps Engineer

Location
United Kingdom
image scanning. Implement and monitor infrastructure and application security controls. Support the organisation's ongoing compliance and certification requirements. Reliability & SRE Establish and maintain observability across distributed systems. Develop proactive monitoring, alerting and performance-tuning strategies. Help maintain service-level objectives and platform availability. Investigate and resolve infrastructure and application … advantageous: MLOps or LLMOps experience. Experience with platforms such as SageMaker, Kubeflow or ZenML . Extensive on-premises Kubernetes deployment experience. Prometheus or comparable observability platforms. AWS Karpenter. AWS Compute Optimizer. Experience operating highly distributed systems. Familiarity with ISO 27001, NIST SSDF, OWASP SAMM or similar security frameworks. Understanding ...

Senior DevOps Engineer

Hiring Organisation
MarkIT Placements
Location
Didcot, Oxfordshire, South East, United Kingdom
Employment Type
Permanent
image scanning. Implement and monitor infrastructure and application security controls. Support the organisation's ongoing compliance and certification requirements. Reliability & SRE Establish and maintain observability across distributed systems. Develop proactive monitoring, alerting and performance-tuning strategies. Help maintain service-level objectives and platform availability. Investigate and resolve infrastructure and application … advantageous: MLOps or LLMOps experience. Experience with platforms such as SageMaker, Kubeflow or ZenML . Extensive on-premises Kubernetes deployment experience. Prometheus or comparable observability platforms. AWS Karpenter. AWS Compute Optimizer. Experience operating highly distributed systems. Familiarity with ISO 27001, NIST SSDF, OWASP SAMM or similar security frameworks. Understanding ...

AI Platform Engineer- Senior Consultant-AI and Digital Factory

Location
Greater London, England, United Kingdom
Implement MLOps and LLMOps pipelines (model deployment, monitoring, retraining, and fine-tuning where relevant) using Infrastructure-as-Code, GitOps, and CI/CD• Establish observability, security, and governance frameworks specific to AI systems, including cost attribution and lifecycle management• Work with clients and internal teams to develop new opportunities … equivalent), including familiarity with the Model Context Protocol (MCP)• Evaluation engineering (golden datasets, regression gates in CI, LLM-judge calibration)• Guardrail and AI-observability tooling (e.g. NeMo Guardrails, OpenTelemetry GenAI conventions, LangSmith, Braintrust)MLOps & LLMOps• Hands-on with MLOps platforms (Azure ML, Databricks, SageMaker) and vector/retrieval databases (Pinecone ...

Data Technical Lead

Hiring Organisation
PA Consulting
Location
London, UK
Employment Type
Full-time
Data governance: Catalogue, metadata, lineage, quality, security, privacy and access controls Data products: Reusable, discoverable and well-governed data products and marketplaces DataOps and observability: Testing, monitoring, operational controls and platform reliability Platform engineering: Terraform, CloudFormation, Azure Bicep and infrastructure-as-code CI/CD: GitHub Actions, Azure DevOps, Jenkins … serving. Applying strong software engineering practices to data platforms, including testing, CI/CD and infrastructure-as-code. Establishing effective approaches to data quality, observability, governance, metadata and security. You can lead data platform modernisation and migration, including coexistence, cutover and decommissioning. Making pragmatic technology choices and understanding the trade ...

Senior DevOps Engineer | London, Hybrid | up to £125k

Location
Greater London, England, United Kingdom
/or CloudFormation), ensuring repeatable environments. Drive containerisation and orchestration using Docker and Kubernetes (including managed services such as EKS/AKS). Improve observability with monitoring, logging and alerting (e.g. Prometheus, Grafana, ELK/EFK, CloudWatch). Embed security best practice: IAM, secrets management, patching, vulnerability remediation and secure ...

Platform Engineer

Location
Greater London, England, United Kingdom
Experience with Infrastructure as Code tools such as Terraform or CloudFormation. Knowledge of CI/CD platforms and deployment automation. Experience with monitoring and observability tools such as CloudWatch, Prometheus, Grafana, ELK, or similar. Good understanding of Linux systems administration. Experience troubleshooting complex application and infrastructure issues. Knowledge of networking ...

Senior AI Engineer

Hiring Organisation
INFUSED SOLUTIONS LIMITED
Location
London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£75,000
Data Scientists to productionise Machine Learning and NLP models. Develop high-performance RESTful APIs and microservices for AI model serving. Drive system reliability, monitoring, observability and performance optimisation. Champion engineering best practices including CI/CD, automated testing and Infrastructure as Code. Improve integration, interoperability and data exchange across enterprise ...

Platform Engineer

Location
Greater London, England, United Kingdom
maintainability. Prompt Engineering for Engineering Workflows: Create clear prompts for troubleshooting, documentation, and operational runbooks to improve speed and consistency. AIOps and Intelligent Observability: Apply AI to analyze logs, metrics, traces, alerts, and incident history to detect anomalies, identify root causes, reduce noise, and improve mean time to resolution. ...

AWS DevOps Platform Engineer - eSC/eDV Clearance

Location
Leicester, England, United Kingdom
Developed Vetting (DV). Preferred technical and professional experience Experience working within multi‐account AWS environments (e.g. Organizations, landing zones) Familiarity with observability and monitoring tooling (CloudWatch, Prometheus, Grafana, ELK stack) Exposure to security best practices in cloud and container environments (IAM, secrets management, Zero Trust concepts) Experience supporting ...

Lead Site Reliability Engineer (Kubernetes Required) - Hybrid

Hiring Organisation
FactSet Research Systems
Location
London, UK
Employment Type
Full-time
fluent in English both verbal and writtenundefinedAdditional Technical SkillsCloud Platforms:(e.g. AWS, GCP, Azure)CI/CD Tooling:(e.g. GitHub Actions, ArgoCD, Harness)Monitoring & Observability:(e.g. Prometheus, Grafana, Coralogix, OpenTelemetry)Infrastructure as Code:(e.g. Terraform, Pulumi)Config Management: (e.g. Ansible, Puppet, Chef)Programming/Scripting:(e.g. Python, Go, Bash)Soft ...

Site Reliability Software Engineer (Hybrid)

Location
Greater London, England, United Kingdom
help ensure that systems used by clinical, operational, and administrative teams remain stable, secure, and available. This role is responsible for building automation, improving observability, reducing manual operational work, supporting integrations, and helping maintain reliable systems that directly impact patient care and business operations. This is a hybrid position with ...

Senior AI Engineer

Hiring Organisation
Genesis10
Location
Columbus, Ohio, United States
Employment Type
Permanent
Salary
USD Hourly
front-end technologies Experience with secure API integrations and OAuth 2.0/OpenID Connect authorization patterns Understanding of software development lifecycle, CI/CD, observability, reliability, security, privacy, and responsible AI practices Strong problem-solving skills, attention to detail, and technical communication Desired skills: Hands-on experience using Claude Code ...

Senior Lead Software Engineer - Python / Go

Location
Glasgow, Scotland, United Kingdom
eliminate platform bottlenecks. Define and promote paved paths and self-service workflows for developers. Implement real-time telemetry pipelines and workflows for platform observability and analytics. Champion adoption of productivity tools through documentation, training, and engagement with the developer community. Standardize use of AI-assisted coding tools and AI-powered ...

Lead Site Reliability Engineer (Kubernetes Required) - Hybrid

Location
Greater London, England, United Kingdom
both verbal and written undefined Additional Technical Skills Cloud Platforms: (e.g. AWS, GCP, Azure) CI/CD Tooling: (e.g. GitHub Actions, ArgoCD, Harness) Monitoring & Observability: (e.g. Prometheus, Grafana, Coralogix, OpenTelemetry) Infrastructure as Code: (e.g. Terraform, Pulumi) Config Management: (e.g. Ansible, Puppet, Chef) Programming/Scripting: (e.g. Python, Go, Bash) Soft ...

DevOps Engineer, Studios

Location
Greater London, England, United Kingdom
across multiple teams. Automate infrastructure provisioning using Infrastructure as Code tools such as Terraform, CloudFormation, or similar. Monitor system performance, availability, and reliability using observability tools such as Prometheus, Grafana, and ELK stack. Ensure high availability and disaster recovery strategies are in place and tested regularly. Collaborate closely with development ...

Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham

Hiring Organisation
Goldman Sachs
Location
Birmingham, West Midlands (County), United Kingdom
Salary
£ 70 K
service meshes and ingress controllers.Advanced experience with major cloud providers (AWS, GCP, or Azure), specifically building and operating highly resilient cloud-native architectures.Proficiency with Observability stacks, including distributed tracing, logging, and metrics (e.g., Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, or CloudWatch)Experience with automated testing and SDLC concepts, developing applications ...

Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham

Location
West Midlands, England, United Kingdom
ingress controllers. Advanced experience with major cloud providers (AWS, GCP, or Azure), specifically building and operating highly resilient cloud-native architectures. Proficiency with Observability stacks, including distributed tracing, logging, and metrics (e.g., Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, or CloudWatch) Experience with automated testing and SDLC concepts, developing applications ...