801 to 825 of 1,407 Permanent Observability Jobs

GenAI Engineer - SRE

Hiring Organisation
Dimension Consulting
Location
Phoenix, Arizona, United States
Employment Type
Permanent
Salary
USD 50 Annual
Health, and Engineering Productivity initiatives. The candidate will leverage Generative AI technologies, automation frameworks, and cloud-native tooling to improve operational efficiency, incident reduction, observability, code quality, and developer productivity across enterprise platforms. The role requires close collaboration with SRE teams, platform engineering teams, application development teams, and business stakeholders … Code Monitoring tools such as Splunk, Dynatrace, Prometheus, Grafana, Datadog, New Relic SRE Knowledge Incident Management Problem Management Service Reliability Availability Management Operational Excellence Observability Principles Production Support ...

Senior Consultant Snowflake

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
ecosystem. The ideal candidate will combine deep technical expertise with strong delivery and people leadership capabilities to drive platform reliability, operational excellence, governance, automation, observability, and stakeholder management. The role requires hands‐on expertise in at least two platform technologies, with mandatory expertise in either Snowflake or Confluent Kafka. … KPIs, operational metrics, and customer commitments. Proactively manage risks, dependencies, and technical blockers. Drive compliance, security, audit readiness, and cost optimization. Implement effective monitoring, observability, and service management processes. Hands‐on experience with Snowflake, Kafka, Cloud, DevOps, and Platform engineering. Mentor and guide Snowflake, Kafka, Cloud, DevOps, and Platform engineers. ...

Senior DevOps Engineer

Hiring Organisation
Anson Mccade
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£75,000
support cloud-native platforms across AWS and Azure Implement DevSecOps and Infrastructure as Code best practices Drive platform reliability using SRE principles and observability tools Support incident management and continuous improvement initiatives Implement Terraform-based infrastructure solutions Leverage automation and AI-assisted engineering tools to improve delivery efficiency Support … delivery teams What We're Looking For in a Senior DevOps Engineer Hands-on DevSecOps and platform engineering expertise Advanced Terraform knowledge Experience with observability tools such as Dynatrace, Grafana or similar Understanding of Site Reliability Engineering principles Experience supporting production environments and incident management Strong stakeholder management and communication ...

DevSecOps Automation Lead London, United Kingdom SMA DevSecOps Posted 3 hours ago

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
automation, and deployment issues, and contribute to root cause analysis and service improvement actions.* Contribute to documentation, engineering standards, and best practices for automation, observability, and platform operations.## **Who you are*** Hands-on experience with Elastic Stack (Elasticsearch, Logstash, Kibana, Beats) or similar observability platforms.* Strong scripting or development skills ...

Senior Site Reliability Engineer

Hiring Organisation
VIQU IT
Location
Yorkshire, United Kingdom
Employment Type
Permanent
Salary
GBP 65,000 - 75,000 Annual
improvements across platform reliability, automation and infrastructure as code. Lead the implementation of CI/CD best practices to improve software delivery. Enhance monitoring, observability and incident management across cloud environments. Collaborate with engineering teams to improve performance, resilience and operational efficiency. Mentor and coach engineers, promoting SRE and DevOps … . Experience building and maintaining CI/CD pipelines. Knowledge of containerisation technologies such as Kubernetes, Amazon EKS or ECS. Experience with monitoring and observability tooling such as Grafana, Prometheus, OpenSearch or similar. Strong understanding of cloud security, resilience and infrastructure automation. Previous experience mentoring engineers or providing technical leadership. ...

Site Reliability Engineer (DV Security Clearance)

Hiring Organisation
CGI
Location
Gloucestershire, United Kingdom
Employment Type
Full Time
mission-critical services supporting national security programmes. You will work closely with software engineers, platform teams and stakeholders to automate operational processes, improve observability and enhance system reliability. In this role, you will take ownership of service health, contribute to platform evolution and help create scalable solutions that enable teams … effectiveness. Key responsibilities: ~Improve & Enhance service reliability, availability and operational resilience ~Automate & Optimise infrastructure, operational workflows and platform processes ~Design & Implement monitoring, alerting and observability solutions ~Support & Scale Kubernetes and containerised environments ~Develop & Deliver CI/CD pipelines and deployment automation capabilities ~Investigate & Resolve incidents, conducting root cause analysis ...

Sr Lead AI Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
Owns the design and build of the team\'s platform: deployment pipelines, model serving, containerisation, orchestration, and environment management Sets the standard for reliability, observability, and operational excellence across the team\'s production AI/ML services Builds the tooling and paved paths that let AI engineers ship agentic … record of building deployment and release automation Experience serving, scaling, and monitoring ML models or data-intensive services in production (MLOps) Practical experience with observability tooling (metrics, logging, tracing) and production incident response Experience operating ML/LLM workloads in production (LLMOps, inference reliability, cost/performance management) Strong communication ...

Engineering Lead

Hiring Organisation
17918
Location
London, United Kingdom
external systems. Collaborate with architects, product teams, vendors, and business stakeholders to ensure successful solution delivery. Drive best practices across software engineering, DevOps, observability, resiliency, and operational excellence. Conduct architecture reviews, code reviews, and technical design assessments. Provide technical mentoring and hands-on guidance to engineering teams. Contribute directly … while leading multiple engineering teams. Desirable: Experience within Banking, Financial Services, or large-scale enterprise transformation programmes. Experience with Docker and Kubernetes. Knowledge of observability, monitoring, and site reliability practices. Experience supporting geographically distributed engineering teams. Exposure to AI-enabled engineering tools, automation frameworks, or developer productivity tooling. What ...

Senior Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Salary
£65,000
delivery challenges. Conduct code reviews and design reviews to promote high-quality engineering standards and knowledge sharing. Define and improve best practices for testing, observability, performance, reliability, and operational excellence. Investigate production issues and drive continuous improvement through root-cause analysis. Ensure solutions meet security, compliance, reliability, and performance requirements. … Experience with C# and .NET development. Knowledge of SQL Server and PostgreSQL. Experience with Infrastructure as Code (IaC) and GitOps practices. Experience implementing monitoring, observability, and operational tooling. What Success Looks Like Delivering impactful customer-facing features and projects successfully. Driving sound architectural decisions and engineering excellence. Building secure, scalable ...

GCP Cloud Architect

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
rehost, replatform, refactor, repurchase, retire, retain, and relocate. Broad hands‐on knowledge of core Google Cloud services across compute, containers, storage, databases, networking, security, observability, DevOps, and cost management. In‐depth experience designing Google Cloud landing zones and cloud foundations, including resource hierarchy, projects, billing, IAM, shared … methodologies as the market evolves. Ability to identify opportunities to use automation, AI‐assisted engineering, AIOps, and cloud-native tooling to improve reliability, observability, delivery speed, and operational efficiency. Excellent English verbal and written communication skills, with the ability to engage effectively with both technical and non‐technical stakeholders. Ability ...

Site Reliability Engineer (SRE)

Hiring Organisation
BC Forward
Location
Chandler, Arizona, United States
Employment Type
Permanent
Salary
USD 7,023 Annual
seeking a Site Reliability Engineer III to join our team. The ideal candidate will have strong experience in cloud computing, infrastructure automation, and observability tooling and a proven ability to implement reliable, automated, and measurable service operations across complex environments. Responsibilities: Establish and maintain partnerships with Application Development and Production … focus on compute, storage, network, and security services. Develop and maintain code and automation using Python, Golang, and shell scripting. Implement monitoring and observability with Prometheus, Dynatrace, Azure Monitor, and Log Analytics. Contribute to CI/CD pipelines using Git, Jenkins, and GitOps practices. Decompose complex objectives into units ...

Principal Java Engineer

Hiring Organisation
Jobleads-UK
Location
Wallingford, England, United Kingdom
continuous improvement. Production systems are reliable, observable and operationally excellent Lead root cause analysis and resolution of complex production issues. Drive improvements in system observability, monitoring and operational performance. Ensure applications are designed and operated to meet reliability, availability and performance targets. Partner with Operations, DevOps and QA teams … Claude, Codex, Gitlab Duo, etc) REST APIs, OpenAPI, Microservices, Event-driven architecture (RabbitMQ) Containers, Docker, AWS, Linux CI/CD with GitLab Pipelines & Jenkins Observability: logging, metrics and monitoring MySQL, Apache Solr Front-end UI (e.g. Angular) Person Specification Strategic and systems-thinking mindset Excellent communication and stakeholder management skills ...

Sr Lead AI Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
Owns the design and build of the team's platform: deployment pipelines, model serving, containerisation, orchestration, and environment management Sets the standard for reliability, observability, and operational excellence across the team's production AI/ML services Builds the tooling and paved paths that let AI engineers ship agentic … record of building deployment and release automation Experience serving, scaling, and monitoring ML models or data-intensive services in production (MLOps) Practical experience with observability tooling (metrics, logging, tracing) and production incident response Experience operating ML/LLM workloads in production (LLMOps, inference reliability, cost/performance management) Strong communication ...

Automation & Platform Engineer

Hiring Organisation
Capgemini
Location
Oxfordshire, United Kingdom
Employment Type
Full Time
services, agent services, APIs, and microservices. Implement infrastructure-as-code for platform environments and network automation resources. Ensure automation platforms meet security, compliance, availability, observability, and operational resilience requirements. Your Profile Experience in network automation, platform engineering, DevOps, cloud engineering, or telecom automation. Strong hands-on experience with Ansible, Terraform … vendor APIs. API management and orchestration. Intent-to-configuration workflows. Data pipelines. Vector databases and graph APIs. MCP integration. Security frameworks and compliance controls. Observability and logging. Preferred Certifications Kubernetes CKA/CKAD. Terraform Associate. Red Hat Ansible certification. Google Cloud, AWS, or Azure certification. Cisco, Juniper, Nokia, or Ericsson ...

Senior DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Manchester, England, United Kingdom
quickly while remaining safe, resilient and compliant. Partnering closely with feature teams, providing hands-on support, coaching and mentoring on CI/CD, automation, observability and DevOps best practices, and supporting the growth of junior engineers. Continuously improving platform standards and engineering quality, through design reviews, code reviews, knowledge sharing … approach. Working knowledge of infrastructure-as-code (e.g. Terraform) and how it supports automated delivery, rather than being the primary focus. Experience with observability and monitoring tooling such as Dynatrace, Prometheus, Splunk or similar. And any experience of these would be really useful Application development and testing ecosystems (e.g. Java ...

Staff Software Engineer - AI

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event‐driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency Establish engineering best practices through hands‐on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement Partner with ...

Software Engineer - AI Platform & Agents

Hiring Organisation
Moody's
Location
Greater London, United Kingdom
Employment Type
Full Time
distributed systems, and event-driven architectures in modern programming languages such as Python, TypeScript, Java, Go, or similar Familiarity with MLOps practices, model monitoring, observability, versioning, and automated deployment pipelines preferred Strong problem-solving skills with the ability to navigate ambiguity, experiment rapidly, and deliver impactful solutions that create measurable … business challenges autonomously Optimize applications for performance, scalability, reliability, and cost efficiency while supporting increasing AI adoption and usage Implement engineering best practices around observability, monitoring, testing, security, and operational excellence Establish and champion MLOps practices including model lifecycle management, prompt versioning, automated evaluation, and continuous improvement Build reusable frameworks ...

Sr. Software Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
while architecting systems that handle high-throughput data pipelines. Participate in building robust export pipelines, streaming architectures, webhook integrations and MCP servers. Maintain high observability and reliability standards using tools like Coralogix, CloudWatch, and Grafana. Participate in on-call rotation and incident response for owned services. What You'll Bring … static site generators). Familiarity with authentication, API gateways, and rate limiting strategies. Experience in compliance standards for APIs and data handling. Experience with observability tools and practices. Languages: Golang (primary) with some TypeScript Monitoring: Coralogix, Grafana, CloudWatch CI/CD & IaC: GitHub Actions, Terraform What We Offer Generous paid ...

Staff Software Engineer - AI

Hiring Organisation
Moody's
Location
Greater London, United Kingdom
Employment Type
Full Time
technologies in production environments • Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines • Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency • Establish engineering best practices through hands-on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence • Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement • Partner with ...

AI Engineering Enablement Director

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
/ML, software, or platform engineering, with exposure to automated testing and infrastructure‐as‐code or policy‐as‐code.* Working knowledge of AI observability (logs, metrics, traces, behavioural signals) and practical methods to evaluate or improve AI system behaviour.* Familiarity with AI risk and governance frameworks (e.g., NIST … FinOps, such as cost‐aware model selection, unit economics, or prompt‐efficiency practices.* Experience with MLOps or AI delivery tooling, or with AI‐specific observability systems.* Participation in industry communities or standards bodies, with the ability to translate external practice into internal adoption.* Experience facilitating workshops or engineering enablement events. ...

Solution Architect

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Solution Architect Willis Re is building its global technology estate from the ground up, unencumbered by legacy and designed around data, analytics and modern cloud platforms. We are looking for a Solution Architect to shape ...

Solution Architect (Contract Role)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Solution Architect – 6-month Contract Willis Re is building its global technology estate from the ground up, centered on data, analytics, and modern cloud platforms. We are looking for a Solution Architect to shape and ...

Python Technical Lead FinTech

Hiring Organisation
Run-Time Group Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent
Python Technical Lead Salary: £90-120K work model: Hybrid Were looking for a Python Technical Lead to drive the architecture, development, and delivery of high-performance financial systems. Youll lead a team of engineers ...

Network Analytics & Automation Leader with AI Platforms

Hiring Organisation
Jobleads-UK
Location
Chester, England, United Kingdom
Overview Automation Technologies and AI/ML-Driven Platforms and Analytics Tools; in the realm of automation technologies and AI/ML-driven observability platforms and analytics tools, the following are essential: Terraform Itential NetDevOps Splunk Python React JS Django Database Technologies Proficiency with database technologies is crucial, including: MySQL ...

Senior Director, Data & AI Platform Engineering

Hiring Organisation
Jobleads-UK
Location
United Kingdom
workflows. You will partner with Platform Product Management to ensure scalability, reliability, security, and broad adoption across product domains. You will drive production-grade observability, governance, and AI safeguards while aligning with #J-18808-Ljbffr ...