776 to 800 of 4,833 Observability Jobs

DevOps Manager

Hiring Organisation
Stott & May Professional Search Limited
Location
London, UK
live service. * Lead incident management, root cause analysis and service improvement activities, ensuring issues are resolved effectively and lessons are applied. * Manage monitoring and observability, using tools such as Dynatrace or similar platforms to improve visibility, alerting and platform performance. * Oversee cloud infrastructure, environments and costs, identifying opportunities for optimisation … networking, including DNS, routing, firewalls and load balancing. * Experience managing enterprise-scale web applications, microservices and high-traffic digital platforms. * Experience with monitoring and observability platforms such as Dynatrace, Catchpoint or similar. * Experience managing incidents, problem management, root cause analysis and live production environments. * Understanding of application and cloud security ...

Lead Software Engineer - LLM Ops Platform Reliability

Location
Paisley, Scotland, United Kingdom
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … Amazon EKS and Amazon SageMaker, as well as on-prem and local GPU clusters, using reproducible infrastructure as code and continuous delivery pipelines Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads Tune GPU and accelerator ...

Infrastructure / DevOps Engineer

Location
Birmingham, England, United Kingdom
environments Define and implement infrastructure-as-code using Terraform, CDK, or equivalent Monitor platform health, define SLOs/SLAs, and build alerting and observability tooling Respond to and lead resolution of infrastructure-level incidents; drive post-mortems Harden infrastructure for HIPAA compliance — encryption, access controls, audit logging, and network security … Actions, Jenkins, or equivalent) Solid understanding of networking, security groups, load balancing, and DNS Experience with container orchestration (Docker, Kubernetes, or ECS) Familiarity with observability tooling (Datadog, Grafana, CloudWatch, or equivalent) Understanding of HIPAA infrastructure requirements (encryption at rest/in transit, audit trails, access controls) Nice to have Site ...

Senior Java Developer

Hiring Organisation
Clarify Consultancy Ltd
Location
Leyland, Lancashire, United Kingdom
Employment Type
Full-Time
Salary
£45,000 - £65,000 per annum
such as Kafka. Develop and maintain automated unit and integration tests. Work with Docker and Kubernetes within modern cloud environments. Contribute to monitoring, logging, observability and application performance. Participate in Agile ceremonies including sprint planning, refinement, reviews and retrospectives. Work closely with Product, Architecture and other technical teams to understand … performance within distributed systems. Experience with CI/CD pipelines and modern development practices is essential, as is familiarity with application monitoring, logging and observability tools. The position sits within an Agile/Scrum environment, so strong problem-solving skills, analytical thinking and excellent communication abilities are key to working ...

Corporate KYC Principle Software Engineer - Executive Director

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
regulated financial services environments Establishes engineering standards for LLM-based applications RAG pipelines, embedding workflows, vector store integrations, and model serving ensuring safety, observability, and reproducibility at scale Drives adoption of advanced technical methods and practices aligned with the latest industry standards and product development methodologies Serves as the function … more disciplines (e.g., cloud, AI/ML, data engineering) Experience in large-scale data processing, microservices, API design, Kafka, Redis, MemCached, observability tools (Dynatrace, Splunk, Grafana), and orchestration frameworks (Airflow, Temporal) Advanced working knowledge of relational and NoSQL databases, vector stores, data lake architectures, and data governance Practical cloud-native ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
infrastructure Troubleshoot and resolve infrastructure issues and fix root causes of issues at the source Implement monitoring, logging and alerting solutions to ensure system observability Collaborate with engineers and architects to improve platform documentation, standards and adoption Requirements Strong hands-on experience with cloud platforms (AWS & Azure ideally) Hands …/VNets, load balancers, DNS, security groups/NSGs) Experience with secrets management and identity/access control (IAM, OIDC, Azure AD) Familiarity with observability tooling (Prometheus, Grafana, CloudWatch, Azure Monitor) Risk Benefit Statement Learn more about the LexisNexis Risk team and how we work here We know your well ...

Platform Engineer

Location
Greater London, England, United Kingdom
leading high-performance, distributed computing products to run reliably, securely and efficiently at scale. This includes infrastructure automation, runtime orchestration, CI/CD enablement, observability and performance optimisation. You will work closely with software engineers, data scientists, product owners, delivery leads, and central IT to define and deliver the platform … scheduling. Implement Infrastructure as Code (IaC) and automated environment provisioning. Build and maintain CI/CD pipelines supporting distributed systems and shared components. Implement observability through logging, metrics and alerting, improving platform reliability and debuggability. Monitor and optimise system performance, throughput and resource utilisation. Ensure platform security, access control ...

Advanced Engineer, Investment Technology

Location
United Kingdom
subject matter experts. Their primary focus is on execution within defined parameters, applying modern engineering practices including CI/CD, automated testing, code quality, observability and data quality controls. They may also be accountable for regular reporting or process administration within the squad. Working in partnership with more experienced staff … business requirements into practical, well-engineered technical solutions. Build production-ready solutions using modern engineering practices, including CI/CD, automated testing, code quality, observability and data quality controls. Support data quality, monitoring, reporting and process administration activities to help ensure accurate and consistent outcomes. Identify and resolve technical problems ...

Advanced Engineer, Investment Technology

Location
Henley-on-Thames, England, United Kingdom
subject matter experts. Their primary focus is on execution within defined parameters, applying modern engineering practices including CI/CD, automated testing, code quality, observability and data quality controls. They may also be accountable for regular reporting or process administration within the squad. Working in partnership with more experienced staff … business requirements into practical, well-engineered technical solutions. Build production-ready solutions using modern engineering practices, including CI/CD, automated testing, code quality, observability and data quality controls. Support data quality, monitoring, reporting and process administration activities to help ensure accurate and consistent outcomes. Identify and resolve technical problems ...

Senior MLOps Engineer

Hiring Organisation
Genesys
Location
Portola Valley, California, United States
Employment Type
Any
Salary
USD 115,000 Annual
SageMaker, Azure Machine Learning, Airflow, and Kubeflow. Deploy scalable inference services using FastAPI, SageMaker Endpoints, and Azure ML Endpoints. Implement model monitoring, drift detection, observability, alerting, and production reliability. Support LLMOps, RAG, Amazon Bedrock, Azure OpenAI, LangChain, and LLM evaluation. Work with engineering, data science, security, and DevOps teams ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
development, validation, and optimization of configuration-as-code, improving delivery speed and reducing deployment risk. Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality … applications without these: Hands-on with Helm or Kustomize Experience with GitOps (e.g., Argo CD) Knowledge of secrets management (e.g., HashiCorp Vault) Experience with observability (metrics/logs/tracing) Why Cisco? At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations ...

AI Consulting –Sr. AI Architect & Client Partner

Location
Greater London, England, United Kingdom
services that underpin enterprise AI ecosystems. Guide teams on software architecture, performance, scalability, security, and maintainability. AI governance & LLMOps Architect governance frameworks for auditability, observability, explainability, and compliance. Design guardrails for hallucination, prompt injection, toxicity, and model safety. Establish LLMOps: evaluation pipelines, automated testing, CI/CD, monitoring, and production … with one of Azure, OpenAI, AWS Bedrock, Claude; Kubernetes and cloud-native deployment. LLMOps & evaluation : CI/CD for AI, automated evals, experiment tracking, observability, model lifecycle management. Responsible AI : governance frameworks, guardrails, model safety, compliance, and auditability. Preferred but not required Orchestration frameworks : LangChain/LangGraph, LlamaIndex, CrewAI, AutoGen ...

Senior Site Reliability Engineer

Hiring Organisation
CISCO Systems
Location
London, United Kingdom
Salary
£ 70 K
accelerate development, validation, and optimization of configuration-as-code, improving delivery speed and reducing deployment risk.Adaptable & Problem-Solver: Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance.Ownership & Quality: Own end-to-end configuration quality, enforcing … encourage applications without these:Hands-on with Helm or KustomizeExperience with GitOps (e.g., Argo CD)Knowledge of secrets management (e.g., HashiCorp Vault)Experience with observability (metrics/logs/tracing)CollabHiringWhy Cisco? At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations ...

Platform Engineer - Glasgow

Location
Glasgow, Scotland, United Kingdom
code generation, testing, documentation, and analysis, while understanding model limitations, protecting client data, and improving delivery quality and speed through pragmatic automation. SRE & Observability You’ll bring a reliability mindset to delivery, designing services that are operable by default and measured through meaningful SLIs/SLOs. You’ll help teams … implement pragmatic observability—logging, metrics, and distributed tracing—with actionable alerting, and you’ll contribute to (or lead) incident response and post‐incident reviews that drive learning and measurable improvements. We are looking for experience in the following skills: Strong experience with the AWS cloud platform and core services. Hands ...

Platform Engineer

Location
Greater London, England, United Kingdom
scalable solutions to thousands of customers every day. Our mission is to make deploying and operating software effortless and safe. We focus on automation, observability, and reliability, ensuring every engineering team at Clio can move faster and with confidence. You’ll be part of a globally distributed team, collaborating closely … Code using Terraform, ensuring reproducibility and compliance. Collaborate with developers to improve CI/CD pipelines, deployment strategies, and overall developer experience. Enhance observability and reliability, refining alerting, monitoring, and incident response. Collaborate on cloud optimization projects, improving performance, cost efficiency, and security posture. Mentor and guide team members, fostering ...

Senior .NET Backend Developer

Location
York and North Yorkshire, England, United Kingdom
design, implementation, testing, review, refactoring, documentation and migration, while retaining clear ownership of the outcome. Set a high standard for maintainability, automated testing, security, observability and pull-request review. Investigate complex technical and production issues and drive them through to resolution. What we are looking for Strong experience designing … clinically important data. PostgreSQL, Redis, Elasticsearch or other data and caching technologies. GraphQL, including schema design and gateway patterns. Grafana, OpenTelemetry, Prometheus or equivalent observability tooling. CI/CD, production services, microservices and message-driven systems such as RabbitMQ. How we work We value engineers who own outcomes and remain ...

Senior Software Engineer II, Developer Experience / Operational Excellence

Location
Greater London, England, United Kingdom
confidently. Within DevEx, the Operational Excellence (OPX) team is the group that keeps production healthy at scale. We provide engineering teams the platform capabilities, observability tooling, automated safeguards, incident management tooling, and safe feature release systems they need to deliver highly available systems, ship features with confidence, and investigate … health. Reduce alert noise, surface actionable signals, and empower engineering teams to operate their services confidently with minimal operational burden Develop and evolve our observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection, to give teams real-time, actionable visibility into system health and latency Contribute to AI-driven ...

Senior Full Stack Engineer

Location
Greater London, England, United Kingdom
technical designs and code produced by delivery partners, identifying risks and opportunities for improvement Champion engineering excellence across code quality, testing, continuous delivery, security, observability and platform reliability Collaborate closely with AI, Machine Learning, Product, Infrastructure and DevOps teams to deliver integrated platform capabilities and share knowledge across the engineering … Familiarity with AI orchestration frameworks, intelligent workflows and emerging AI technologies Experience with data platforms and infrastructure that support AI workloads Knowledge of monitoring, observability and continuous delivery practices Experience within telecommunications or other large scale customer facing digital platforms What's in it for you? Competitive salary and bonus ...

Platform Engineer

Location
Greater London, England, United Kingdom
cost‐efficient infrastructure that empowers engineering teams to deliver quickly without sacrificing reliability or compliance. You will be the driving force behind our DevOps, observability, and compliance readiness , ensuring our systems are audit‐ready, highly available, and optimized for both performance and cost. Key Responsibilities Architect, implement, and maintain cloud … incredible journey and learning a lot along the way. Requirements Technical stack : Azure (also AWS is a plus), Terraform, AKS (Kubernetes), Docker, GitHub Actions. Observability : Experience implementing logging, metrics, and tracing frameworks. Security : Familiarity with best practices, secrets management, and security scanning tools. Networking : Solid understanding of VPCs, private networking ...

Senior Test Automation Engineer

Location
Greater London, England, United Kingdom
Delivery Visual Studio Git Azure DevOps Repositories Azure DevOps YAML Pipelines Event Streaming & Messaging Apache Kafka (Aiven) Schema Registry Avro JSON Messaging RabbitMQ Data & Observability Microsoft SQL Server OpenTelemetry ClickStack/HyperDX ClickHouse ELK Stack (Elasticsearch, Logstash, Kibana) Excellent analytical and problem-solving capabilities. Strong stakeholder communication and collaboration skills. … Familiarity with Azure DevOps, TFS, Jira or similar Application Lifecycle Management tools. Exposure to observability and monitoring platforms. Knowledge of event-driven and microservices architectures. Experience Proven experience in test automation engineering within Agile environments. Strong experience in C# test automation development. Experience creating and maintaining BDD frameworks and behaviour ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
City of London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Cardiff, South Glamorgan, Wales, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Darlington, County Durham, North East, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...