751 to 775 of 4,719 Observability Jobs

Intelligent Automation Engineering Manager

Location
Greater London, England, United Kingdom
deployment and operational support of AI agents and AI-powered solutions. Establish engineering standards and best practices for AI architecture, orchestration, retrieval, tool invocation, observability, governance, privacy, security and cost management. Review technical designs and architecture documentation to ensure solutions align with engineering, security and governance standards. Translate emerging … with MCP (Model Context Protocol), MCP Servers, MCP Clients or enterprise AI integration frameworks. Knowledge of LLMOps, AI evaluation frameworks, model routing and AI observability tooling. Experience with enterprise integration technologies including APIs, Middleware, ESB or iPaaS platforms. Hands‐on experience integrating internal and third-party systems. Azure cloud experience ...

Senior Site Reliability Engineer (Azure)

Hiring Organisation
GTN Technical Staffing & Consulting
Location
Plano, Texas, United States
Employment Type
Any
Salary
USD Annual
Site Reliability Engineer (SRE) to lead reliability engineering initiatives across our Azure estate and Command Center operations. This role focuses on scripting, automation, and observability to ensure uptime, performance, and rapid incident response. The Senior SRE will design and implement monitoring-as-code, optimize alerting, and build self-healing automation … professional manner at all times. Team Members are required to observe the companys standards, work requirements and rules of conduct. Essential Duties & Responsibilities: Observability & Monitoring o Architect end-to-end monitoring using Azure Monitor, Log Analytics, Application Insights, and ITRS Geneos. o Implement monitoring-as-code with Terraform/Bicep ...

DevOps Manager

Hiring Organisation
Stott & May Professional Search Limited
Location
United Kingdom
Employment Type
Permanent
live service. * Lead incident management, root cause analysis and service improvement activities, ensuring issues are resolved effectively and lessons are applied. * Manage monitoring and observability, using tools such as Dynatrace or similar platforms to improve visibility, alerting and platform performance. * Oversee cloud infrastructure, environments and costs, identifying opportunities for optimisation … networking, including DNS, routing, firewalls and load balancing. * Experience managing enterprise-scale web applications, microservices and high-traffic digital platforms. * Experience with monitoring and observability platforms such as Dynatrace, Catchpoint or similar. * Experience managing incidents, problem management, root cause analysis and live production environments. * Understanding of application and cloud security ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … Amazon EKS and Amazon SageMaker, as well as on-prem and local GPU clusters, using reproducible infrastructure as code and continuous delivery pipelines Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads Tune GPU and accelerator ...

Software Development Engineer in Test

Location
Greater London, England, United Kingdom
accuracy, reliability, latency, and cost. Work with Engineering, Product, Data Science, QA, and Operations teams to deliver production-ready solutions. Design for scalability, reliability, observability, and operational excellence across the platform. Troubleshoot complex production issues across application services, data pipelines, AI workflows, and telemetry. Participate in architecture, technical design, code … C++. Strong experience working with large-scale data, telemetry, logs, events, and operational datasets. Strong understanding of cloud platforms, system scalability, reliability, performance, observability, and production operations. Strong analytical, problem-solving, communication, and collaboration skills, with the ability to work effectively across technical and cross-functional teams. Desirable skills ...

DevOps Manager

Hiring Organisation
Stott & May Professional Search Limited
Location
London, UK
live service. * Lead incident management, root cause analysis and service improvement activities, ensuring issues are resolved effectively and lessons are applied. * Manage monitoring and observability, using tools such as Dynatrace or similar platforms to improve visibility, alerting and platform performance. * Oversee cloud infrastructure, environments and costs, identifying opportunities for optimisation … networking, including DNS, routing, firewalls and load balancing. * Experience managing enterprise-scale web applications, microservices and high-traffic digital platforms. * Experience with monitoring and observability platforms such as Dynatrace, Catchpoint or similar. * Experience managing incidents, problem management, root cause analysis and live production environments. * Understanding of application and cloud security ...

Lead Software Engineer - LLM Ops Platform Reliability

Location
Paisley, Scotland, United Kingdom
strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. … Amazon EKS and Amazon SageMaker, as well as on-prem and local GPU clusters, using reproducible infrastructure as code and continuous delivery pipelines Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads Tune GPU and accelerator ...

Infrastructure / DevOps Engineer

Location
Birmingham, England, United Kingdom
environments Define and implement infrastructure-as-code using Terraform, CDK, or equivalent Monitor platform health, define SLOs/SLAs, and build alerting and observability tooling Respond to and lead resolution of infrastructure-level incidents; drive post-mortems Harden infrastructure for HIPAA compliance — encryption, access controls, audit logging, and network security … Actions, Jenkins, or equivalent) Solid understanding of networking, security groups, load balancing, and DNS Experience with container orchestration (Docker, Kubernetes, or ECS) Familiarity with observability tooling (Datadog, Grafana, CloudWatch, or equivalent) Understanding of HIPAA infrastructure requirements (encryption at rest/in transit, audit trails, access controls) Nice to have Site ...

Senior Java Developer

Hiring Organisation
Clarify Consultancy Ltd
Location
Leyland, Lancashire, United Kingdom
Employment Type
Full-Time
Salary
£45,000 - £65,000 per annum
such as Kafka. Develop and maintain automated unit and integration tests. Work with Docker and Kubernetes within modern cloud environments. Contribute to monitoring, logging, observability and application performance. Participate in Agile ceremonies including sprint planning, refinement, reviews and retrospectives. Work closely with Product, Architecture and other technical teams to understand … performance within distributed systems. Experience with CI/CD pipelines and modern development practices is essential, as is familiarity with application monitoring, logging and observability tools. The position sits within an Agile/Scrum environment, so strong problem-solving skills, analytical thinking and excellent communication abilities are key to working ...

Corporate KYC Principle Software Engineer - Executive Director

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
regulated financial services environments Establishes engineering standards for LLM-based applications RAG pipelines, embedding workflows, vector store integrations, and model serving ensuring safety, observability, and reproducibility at scale Drives adoption of advanced technical methods and practices aligned with the latest industry standards and product development methodologies Serves as the function … more disciplines (e.g., cloud, AI/ML, data engineering) Experience in large-scale data processing, microservices, API design, Kafka, Redis, MemCached, observability tools (Dynatrace, Splunk, Grafana), and orchestration frameworks (Airflow, Temporal) Advanced working knowledge of relational and NoSQL databases, vector stores, data lake architectures, and data governance Practical cloud-native ...

Platform Engineer

Location
Greater London, England, United Kingdom
leading high-performance, distributed computing products to run reliably, securely and efficiently at scale. This includes infrastructure automation, runtime orchestration, CI/CD enablement, observability and performance optimisation. You will work closely with software engineers, data scientists, product owners, delivery leads, and central IT to define and deliver the platform … scheduling. Implement Infrastructure as Code (IaC) and automated environment provisioning. Build and maintain CI/CD pipelines supporting distributed systems and shared components. Implement observability through logging, metrics and alerting, improving platform reliability and debuggability. Monitor and optimise system performance, throughput and resource utilisation. Ensure platform security, access control ...

Advanced Engineer, Investment Technology

Location
United Kingdom
subject matter experts. Their primary focus is on execution within defined parameters, applying modern engineering practices including CI/CD, automated testing, code quality, observability and data quality controls. They may also be accountable for regular reporting or process administration within the squad. Working in partnership with more experienced staff … business requirements into practical, well-engineered technical solutions. Build production-ready solutions using modern engineering practices, including CI/CD, automated testing, code quality, observability and data quality controls. Support data quality, monitoring, reporting and process administration activities to help ensure accurate and consistent outcomes. Identify and resolve technical problems ...

Advanced Engineer, Investment Technology

Location
Henley-on-Thames, England, United Kingdom
subject matter experts. Their primary focus is on execution within defined parameters, applying modern engineering practices including CI/CD, automated testing, code quality, observability and data quality controls. They may also be accountable for regular reporting or process administration within the squad. Working in partnership with more experienced staff … business requirements into practical, well-engineered technical solutions. Build production-ready solutions using modern engineering practices, including CI/CD, automated testing, code quality, observability and data quality controls. Support data quality, monitoring, reporting and process administration activities to help ensure accurate and consistent outcomes. Identify and resolve technical problems ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
development, validation, and optimization of configuration-as-code, improving delivery speed and reducing deployment risk. Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality … applications without these: Hands-on with Helm or Kustomize Experience with GitOps (e.g., Argo CD) Knowledge of secrets management (e.g., HashiCorp Vault) Experience with observability (metrics/logs/tracing) Why Cisco? At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations ...

AI Consulting –Sr. AI Architect & Client Partner

Location
Greater London, England, United Kingdom
services that underpin enterprise AI ecosystems. Guide teams on software architecture, performance, scalability, security, and maintainability. AI governance & LLMOps Architect governance frameworks for auditability, observability, explainability, and compliance. Design guardrails for hallucination, prompt injection, toxicity, and model safety. Establish LLMOps: evaluation pipelines, automated testing, CI/CD, monitoring, and production … with one of Azure, OpenAI, AWS Bedrock, Claude; Kubernetes and cloud-native deployment. LLMOps & evaluation : CI/CD for AI, automated evals, experiment tracking, observability, model lifecycle management. Responsible AI : governance frameworks, guardrails, model safety, compliance, and auditability. Preferred but not required Orchestration frameworks : LangChain/LangGraph, LlamaIndex, CrewAI, AutoGen ...

Senior Site Reliability Engineer

Hiring Organisation
CISCO Systems
Location
London, United Kingdom
Salary
£ 70 K
accelerate development, validation, and optimization of configuration-as-code, improving delivery speed and reducing deployment risk.Adaptable & Problem-Solver: Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance.Ownership & Quality: Own end-to-end configuration quality, enforcing … encourage applications without these:Hands-on with Helm or KustomizeExperience with GitOps (e.g., Argo CD)Knowledge of secrets management (e.g., HashiCorp Vault)Experience with observability (metrics/logs/tracing)CollabHiringWhy Cisco? At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations ...

Platform Engineer - Glasgow

Location
Glasgow, Scotland, United Kingdom
code generation, testing, documentation, and analysis, while understanding model limitations, protecting client data, and improving delivery quality and speed through pragmatic automation. SRE & Observability You’ll bring a reliability mindset to delivery, designing services that are operable by default and measured through meaningful SLIs/SLOs. You’ll help teams … implement pragmatic observability—logging, metrics, and distributed tracing—with actionable alerting, and you’ll contribute to (or lead) incident response and post‐incident reviews that drive learning and measurable improvements. We are looking for experience in the following skills: Strong experience with the AWS cloud platform and core services. Hands ...

Platform Engineer

Location
Greater London, England, United Kingdom
scalable solutions to thousands of customers every day. Our mission is to make deploying and operating software effortless and safe. We focus on automation, observability, and reliability, ensuring every engineering team at Clio can move faster and with confidence. You’ll be part of a globally distributed team, collaborating closely … Code using Terraform, ensuring reproducibility and compliance. Collaborate with developers to improve CI/CD pipelines, deployment strategies, and overall developer experience. Enhance observability and reliability, refining alerting, monitoring, and incident response. Collaborate on cloud optimization projects, improving performance, cost efficiency, and security posture. Mentor and guide team members, fostering ...

Senior .NET Backend Developer

Location
York and North Yorkshire, England, United Kingdom
design, implementation, testing, review, refactoring, documentation and migration, while retaining clear ownership of the outcome. Set a high standard for maintainability, automated testing, security, observability and pull-request review. Investigate complex technical and production issues and drive them through to resolution. What we are looking for Strong experience designing … clinically important data. PostgreSQL, Redis, Elasticsearch or other data and caching technologies. GraphQL, including schema design and gateway patterns. Grafana, OpenTelemetry, Prometheus or equivalent observability tooling. CI/CD, production services, microservices and message-driven systems such as RabbitMQ. How we work We value engineers who own outcomes and remain ...

Senior Software Engineer II, Developer Experience / Operational Excellence

Location
Greater London, England, United Kingdom
confidently. Within DevEx, the Operational Excellence (OPX) team is the group that keeps production healthy at scale. We provide engineering teams the platform capabilities, observability tooling, automated safeguards, incident management tooling, and safe feature release systems they need to deliver highly available systems, ship features with confidence, and investigate … health. Reduce alert noise, surface actionable signals, and empower engineering teams to operate their services confidently with minimal operational burden Develop and evolve our observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection, to give teams real-time, actionable visibility into system health and latency Contribute to AI-driven ...

Senior Full Stack Engineer

Location
Greater London, England, United Kingdom
technical designs and code produced by delivery partners, identifying risks and opportunities for improvement Champion engineering excellence across code quality, testing, continuous delivery, security, observability and platform reliability Collaborate closely with AI, Machine Learning, Product, Infrastructure and DevOps teams to deliver integrated platform capabilities and share knowledge across the engineering … Familiarity with AI orchestration frameworks, intelligent workflows and emerging AI technologies Experience with data platforms and infrastructure that support AI workloads Knowledge of monitoring, observability and continuous delivery practices Experience within telecommunications or other large scale customer facing digital platforms What's in it for you? Competitive salary and bonus ...

Platform Engineer

Location
Greater London, England, United Kingdom
cost‐efficient infrastructure that empowers engineering teams to deliver quickly without sacrificing reliability or compliance. You will be the driving force behind our DevOps, observability, and compliance readiness , ensuring our systems are audit‐ready, highly available, and optimized for both performance and cost. Key Responsibilities Architect, implement, and maintain cloud … incredible journey and learning a lot along the way. Requirements Technical stack : Azure (also AWS is a plus), Terraform, AKS (Kubernetes), Docker, GitHub Actions. Observability : Experience implementing logging, metrics, and tracing frameworks. Security : Familiarity with best practices, secrets management, and security scanning tools. Networking : Solid understanding of VPCs, private networking ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
City of London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...