1,126 to 1,150 of 1,968 Observability Jobs

GenAI Engineer - SRE

Hiring Organisation
Dimension Consulting
Location
Phoenix, Arizona, United States
Employment Type
Permanent
Salary
USD 50 Annual
Health, and Engineering Productivity initiatives. The candidate will leverage Generative AI technologies, automation frameworks, and cloud-native tooling to improve operational efficiency, incident reduction, observability, code quality, and developer productivity across enterprise platforms. The role requires close collaboration with SRE teams, platform engineering teams, application development teams, and business stakeholders … Code Monitoring tools such as Splunk, Dynatrace, Prometheus, Grafana, Datadog, New Relic SRE Knowledge Incident Management Problem Management Service Reliability Availability Management Operational Excellence Observability Principles Production Support ...

Devops SRE

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
shared Kubernetes services such as CoreDNS , cert‐manager , Dynatrace , Cloudability , and Infoblox . Familiarity with OPA Gatekeeper for policy enforcement and tenant isolation. Security, Observability & Performance Strong security mindset with a proven track record of designing secure, resilient cloud‐native systems. Experience implementing observability stacks including Prometheus , Dynatrace , and OpenTelemetry ...

Observability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
best ideas take time to evolve. Together we’re building a world‐class platform to amplify our teams’ most powerful ideas. Role The Observability Engineering Team manages the doors – both entry and exit – to the telemetry backends at G‐Research, ensuring our engineers can effectively produce and consume telemetry … their services. As an Observability Engineer, you’ll help make observability seamless for developers and platform teams by building pipelines to ingest and route data in predictable, composable ways, as well as visualising that data after the fact. You’ll have deep experience across observability stacks, a clear understanding ...

Senior Consultant Snowflake

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
ecosystem. The ideal candidate will combine deep technical expertise with strong delivery and people leadership capabilities to drive platform reliability, operational excellence, governance, automation, observability, and stakeholder management. The role requires hands‐on expertise in at least two platform technologies, with mandatory expertise in either Snowflake or Confluent Kafka. … KPIs, operational metrics, and customer commitments. Proactively manage risks, dependencies, and technical blockers. Drive compliance, security, audit readiness, and cost optimization. Implement effective monitoring, observability, and service management processes. Hands‐on experience with Snowflake, Kafka, Cloud, DevOps, and Platform engineering. Mentor and guide Snowflake, Kafka, Cloud, DevOps, and Platform engineers. ...

Senior DevOps Engineer

Hiring Organisation
Anson Mccade
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£75,000
support cloud-native platforms across AWS and Azure Implement DevSecOps and Infrastructure as Code best practices Drive platform reliability using SRE principles and observability tools Support incident management and continuous improvement initiatives Implement Terraform-based infrastructure solutions Leverage automation and AI-assisted engineering tools to improve delivery efficiency Support … delivery teams What We're Looking For in a Senior DevOps Engineer Hands-on DevSecOps and platform engineering expertise Advanced Terraform knowledge Experience with observability tools such as Dynatrace, Grafana or similar Understanding of Site Reliability Engineering principles Experience supporting production environments and incident management Strong stakeholder management and communication ...

SRE Security engineer

Hiring Organisation
FBI &TMT
Location
London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
£500 - £594 per day
pipelines. Support incident management, root cause analysis, and continuous improvement activities. Collaborate with stakeholders to implement security controls and compliance requirements. Drive improvements in observability, monitoring, logging, and alerting. Conduct vulnerability remediation and support security reviews and audits. Essential Skills & Experience Active SC Clearance (mandatory). Proven experience … including Kubernetes. Experience implementing and maintaining CI/CD pipelines. Strong understanding of security principles, vulnerability management, and security monitoring. Experience with monitoring and observability tools such as Prometheus, Grafana, Splunk, ELK, or similar. Excellent troubleshooting and stakeholder engagement skills. Desirable Experience working within government, defence, or highly regulated environments. ...

SRE Security engineer

Hiring Organisation
Matchtech
Location
London, South East, England, United Kingdom
Employment Type
Contractor
Contract Rate
£500 - £594 per day
pipelines. Support incident management, root cause analysis, and continuous improvement activities. Collaborate with stakeholders to implement security controls and compliance requirements. Drive improvements in observability, monitoring, logging, and alerting. Conduct vulnerability remediation and support security reviews and audits. Essential Skills & Experience Active SC Clearance (mandatory). Proven experience … including Kubernetes. Experience implementing and maintaining CI/CD pipelines. Strong understanding of security principles, vulnerability management, and security monitoring. Experience with monitoring and observability tools such as Prometheus, Grafana, Splunk, ELK, or similar. Excellent troubleshooting and stakeholder engagement skills. Desirable Experience working within government, defence, or highly regulated environments. ...

DevSecOps Automation Lead London, United Kingdom SMA DevSecOps Posted 3 hours ago

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
automation, and deployment issues, and contribute to root cause analysis and service improvement actions.* Contribute to documentation, engineering standards, and best practices for automation, observability, and platform operations.## **Who you are*** Hands-on experience with Elastic Stack (Elasticsearch, Logstash, Kibana, Beats) or similar observability platforms.* Strong scripting or development skills ...

Site Reliability Engineer (DV Security Clearance)

Hiring Organisation
CGI
Location
Gloucestershire, United Kingdom
Employment Type
Full Time
mission-critical services supporting national security programmes. You will work closely with software engineers, platform teams and stakeholders to automate operational processes, improve observability and enhance system reliability. In this role, you will take ownership of service health, contribute to platform evolution and help create scalable solutions that enable teams … effectiveness. Key responsibilities: ~Improve & Enhance service reliability, availability and operational resilience ~Automate & Optimise infrastructure, operational workflows and platform processes ~Design & Implement monitoring, alerting and observability solutions ~Support & Scale Kubernetes and containerised environments ~Develop & Deliver CI/CD pipelines and deployment automation capabilities ~Investigate & Resolve incidents, conducting root cause analysis ...

Java Engineer

Hiring Organisation
Response Informatics
Location
Belfast, County Antrim, Northern Ireland, United Kingdom
Employment Type
Contract
Contract Rate
From £250 to £300 per day
delivering APIs using Java, Go, or both. Strong understanding of cloud-native application architecture, including containerized services, service-to-service communication, configuration management, observability, and resilience patterns. Experience delivering production services deployed on Kubernetes. Experience building highly available, scalable, and reliable distributed systems. Strong knowledge of API design principles, including … Helm, Operators, Custom Resource Definitions, or GitOps workflows. Experience with distributed systems, stateful services, or high-volume transactional platforms. Experience with cloud-native observability tools such as Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, or similar. Experience implementing security best practices for APIs and cloud-native services, including OAuth2, mTLS, secrets ...

Sr Lead AI Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
Owns the design and build of the team\'s platform: deployment pipelines, model serving, containerisation, orchestration, and environment management Sets the standard for reliability, observability, and operational excellence across the team\'s production AI/ML services Builds the tooling and paved paths that let AI engineers ship agentic … record of building deployment and release automation Experience serving, scaling, and monitoring ML models or data-intensive services in production (MLOps) Practical experience with observability tooling (metrics, logging, tracing) and production incident response Experience operating ML/LLM workloads in production (LLMOps, inference reliability, cost/performance management) Strong communication ...

Camunda Architect / Lead Architect (Camunda 8 Preferred)

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
strategy across Camunda SaaS and self-managed models Collaborate with engineering and DevOps teams on: CI/CD pipelines containerised deployments Kubernetes-based environments observability and platform monitoring Technical Leadership Provide technical leadership and architectural guidance across engineering teams and delivery squads Mentor engineers and support the uplift of workflow … Camunda 8 Experience with Camunda SaaS and self-managed deployments at scale Exposure to cloud platforms such as: AWS Azure GCP Experience with observability tooling such as: Prometheus Grafana OpenTelemetry Experience with: enterprise integration patterns API management workflow/task UI customisation Experience in regulated industries such as: Banking Insurance ...

Engineering Lead

Hiring Organisation
Hays
Location
Cheshire, North West, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
Up to £500.0 per day + Inside IR35
external systems. Collaborate with architects, product teams, vendors, and business stakeholders to ensure successful solution delivery. Drive best practices across software engineering, DevOps, observability, resiliency, and operational excellence. Conduct architecture reviews, code reviews, and technical design assessments. Provide technical mentoring and hands-on guidance to engineering teams. Contribute directly … while leading multiple engineering teams. Desirable: Experience within Banking, Financial Services, or large-scale enterprise transformation programmes. Experience with Docker and Kubernetes. Knowledge of observability, monitoring, and site reliability practices. Experience supporting geographically distributed engineering teams. Exposure to AI-enabled engineering tools, automation frameworks, or developer productivity tooling. What ...

Engineering Lead

Hiring Organisation
17918
Location
London, United Kingdom
external systems. Collaborate with architects, product teams, vendors, and business stakeholders to ensure successful solution delivery. Drive best practices across software engineering, DevOps, observability, resiliency, and operational excellence. Conduct architecture reviews, code reviews, and technical design assessments. Provide technical mentoring and hands-on guidance to engineering teams. Contribute directly … while leading multiple engineering teams. Desirable: Experience within Banking, Financial Services, or large-scale enterprise transformation programmes. Experience with Docker and Kubernetes. Knowledge of observability, monitoring, and site reliability practices. Experience supporting geographically distributed engineering teams. Exposure to AI-enabled engineering tools, automation frameworks, or developer productivity tooling. What ...

Principal Cloud Engineer

Hiring Organisation
Jobleads-UK
Location
Bristol, England, United Kingdom
storage solutions. This is a hand‐on technical role requiring a solid background in the use of cloud infrastructure, deployment using Infrastructure‐as‐Code, observability, high‐performance networking and storage systems. You may have been working in an IT organisation, a datacentre, a cloud provider or as a developer … users in their use. Turn end‐user and product requirements into deployed services. Help to build automation to collect and analyse metrics and other observability data from the cloud services to support clear identification and reporting of any issues. Work with users to provide information of any product‐related issues ...

Senior Solutions Architect

Hiring Organisation
Jobleads-UK
Location
Reading, England, United Kingdom
scaling strategies Integrate GameLift with AWS services (Lambda, DynamoDB, API Gateway, etc.) Optimise cost, performance, and multi‐region deployments Implement monitoring, logging, and observability solutions Qualifications AWS Certified Solutions Architect – Professional (required) Additional AWS certifications (e.g., DevOps, Security) desirable SKILLS AND EXPERIENCE Proven experience in solutions architecture or senior technical … Code: Terraform, CloudFormation, AWS CDK Game Development: Unreal/Unity basics, backend integration patterns Multiplayer Systems: Backend architecture, session management, real‐time systems Observability: CloudWatch, Prometheus, Grafana, or similar Soft Skills Strong communication and presentation skills Ability to engage both technical and non‐technical stakeholders Strategic thinking and problem‐solving ...

Senior Software Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Salary
£65,000
delivery challenges. Conduct code reviews and design reviews to promote high-quality engineering standards and knowledge sharing. Define and improve best practices for testing, observability, performance, reliability, and operational excellence. Investigate production issues and drive continuous improvement through root-cause analysis. Ensure solutions meet security, compliance, reliability, and performance requirements. … Experience with C# and .NET development. Knowledge of SQL Server and PostgreSQL. Experience with Infrastructure as Code (IaC) and GitOps practices. Experience implementing monitoring, observability, and operational tooling. What Success Looks Like Delivering impactful customer-facing features and projects successfully. Driving sound architectural decisions and engineering excellence. Building secure, scalable ...

GCP Cloud Architect

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
rehost, replatform, refactor, repurchase, retire, retain, and relocate. Broad hands‐on knowledge of core Google Cloud services across compute, containers, storage, databases, networking, security, observability, DevOps, and cost management. In‐depth experience designing Google Cloud landing zones and cloud foundations, including resource hierarchy, projects, billing, IAM, shared … methodologies as the market evolves. Ability to identify opportunities to use automation, AI‐assisted engineering, AIOps, and cloud-native tooling to improve reliability, observability, delivery speed, and operational efficiency. Excellent English verbal and written communication skills, with the ability to engage effectively with both technical and non‐technical stakeholders. Ability ...

Site Reliability Engineer (SRE)

Hiring Organisation
BC Forward
Location
Chandler, Arizona, United States
Employment Type
Permanent
Salary
USD 7,023 Annual
seeking a Site Reliability Engineer III to join our team. The ideal candidate will have strong experience in cloud computing, infrastructure automation, and observability tooling and a proven ability to implement reliable, automated, and measurable service operations across complex environments. Responsibilities: Establish and maintain partnerships with Application Development and Production … focus on compute, storage, network, and security services. Develop and maintain code and automation using Python, Golang, and shell scripting. Implement monitoring and observability with Prometheus, Dynatrace, Azure Monitor, and Log Analytics. Contribute to CI/CD pipelines using Git, Jenkins, and GitOps practices. Decompose complex objectives into units ...

Principal Java Engineer

Hiring Organisation
Jobleads-UK
Location
Wallingford, England, United Kingdom
continuous improvement. Production systems are reliable, observable and operationally excellent Lead root cause analysis and resolution of complex production issues. Drive improvements in system observability, monitoring and operational performance. Ensure applications are designed and operated to meet reliability, availability and performance targets. Partner with Operations, DevOps and QA teams … Claude, Codex, Gitlab Duo, etc) REST APIs, OpenAPI, Microservices, Event-driven architecture (RabbitMQ) Containers, Docker, AWS, Linux CI/CD with GitLab Pipelines & Jenkins Observability: logging, metrics and monitoring MySQL, Apache Solr Front-end UI (e.g. Angular) Person Specification Strategic and systems-thinking mindset Excellent communication and stakeholder management skills ...

Sr Lead AI Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
Owns the design and build of the team's platform: deployment pipelines, model serving, containerisation, orchestration, and environment management Sets the standard for reliability, observability, and operational excellence across the team's production AI/ML services Builds the tooling and paved paths that let AI engineers ship agentic … record of building deployment and release automation Experience serving, scaling, and monitoring ML models or data-intensive services in production (MLOps) Practical experience with observability tooling (metrics, logging, tracing) and production incident response Experience operating ML/LLM workloads in production (LLMOps, inference reliability, cost/performance management) Strong communication ...

Automation & Platform Engineer

Hiring Organisation
Capgemini
Location
Oxfordshire, United Kingdom
Employment Type
Full Time
services, agent services, APIs, and microservices. Implement infrastructure-as-code for platform environments and network automation resources. Ensure automation platforms meet security, compliance, availability, observability, and operational resilience requirements. Your Profile Experience in network automation, platform engineering, DevOps, cloud engineering, or telecom automation. Strong hands-on experience with Ansible, Terraform … vendor APIs. API management and orchestration. Intent-to-configuration workflows. Data pipelines. Vector databases and graph APIs. MCP integration. Security frameworks and compliance controls. Observability and logging. Preferred Certifications Kubernetes CKA/CKAD. Terraform Associate. Red Hat Ansible certification. Google Cloud, AWS, or Azure certification. Cisco, Juniper, Nokia, or Ericsson ...

Senior DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Manchester, England, United Kingdom
quickly while remaining safe, resilient and compliant. Partnering closely with feature teams, providing hands-on support, coaching and mentoring on CI/CD, automation, observability and DevOps best practices, and supporting the growth of junior engineers. Continuously improving platform standards and engineering quality, through design reviews, code reviews, knowledge sharing … approach. Working knowledge of infrastructure-as-code (e.g. Terraform) and how it supports automated delivery, rather than being the primary focus. Experience with observability and monitoring tooling such as Dynatrace, Prometheus, Splunk or similar. And any experience of these would be really useful Application development and testing ecosystems (e.g. Java ...

Staff Software Engineer - AI

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event‐driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency Establish engineering best practices through hands‐on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement Partner with ...

Software Engineer - AI Platform & Agents

Hiring Organisation
Moody's
Location
Greater London, United Kingdom
Employment Type
Full Time
distributed systems, and event-driven architectures in modern programming languages such as Python, TypeScript, Java, Go, or similar Familiarity with MLOps practices, model monitoring, observability, versioning, and automated deployment pipelines preferred Strong problem-solving skills with the ability to navigate ambiguity, experiment rapidly, and deliver impactful solutions that create measurable … business challenges autonomously Optimize applications for performance, scalability, reliability, and cost efficiency while supporting increasing AI adoption and usage Implement engineering best practices around observability, monitoring, testing, security, and operational excellence Establish and champion MLOps practices including model lifecycle management, prompt versioning, automated evaluation, and continuous improvement Build reusable frameworks ...