376 to 400 of 4,013 Observability Jobs

Camunda Architect / Lead Architect (Camunda 8 Preferred)

Location
Greater London, England, United Kingdom
strategy across Camunda SaaS and self-managed models Collaborate with engineering and DevOps teams on: CI/CD pipelines containerised deployments Kubernetes-based environments observability and platform monitoring Technical Leadership Provide technical leadership and architectural guidance across engineering teams and delivery squads Mentor engineers and support the uplift of workflow … Camunda 8 Experience with Camunda SaaS and self-managed deployments at scale Exposure to cloud platforms such as: AWS Azure GCP Experience with observability tooling such as: Prometheus Grafana OpenTelemetry Experience with: enterprise integration patterns API management workflow/task UI customisation Experience in regulated industries such as: Banking Insurance ...

DevOps Engineer - SC Cleared

Hiring Organisation
Opus Recruitment Solutions Ltd
Location
Martock, Somerset, United Kingdom
Employment Type
Full-Time
Salary
£350.00 - £500.00 per day
CI. Automate deployment, configuration and operational processes to improve efficiency and reliability. Support containerised workloads and cloud-native deployment patterns. Implement monitoring, logging and observability solutions to improve platform visibility and operational performance. Drive infrastructure resilience, scalability, security and cost optimisation initiatives. Support incident response, troubleshooting and root cause analysis … Code expertise. Experience building and maintaining CI/CD pipelines, ideally using GitLab CI. Strong Docker and containerisation experience. Experience with monitoring and observability tools such as Prometheus, Grafana and CloudWatch. Knowledge of AWS security best practices, IAM and Well-Architected principles. Experience implementing automation and reducing manual operational overhead. ...

Operations Engineering Lead

Hiring Organisation
Willis Towers Watson
Location
London, UK
Employment Type
Full-time
base platform operations team. You will help set direction for shared infrastructure and engineering operations - the production hosting, security perimeter, release pipelines, observability, and developer tooling that every engineering squad depends on to ship safely. This is a role that includes Line Management responsibility for your team … networking, scalability, performance, and cost-efficiency across production environmentsOversee Infrastructure as Code practices (e.g. Terraform) and ensure environments are consistent, auditable, and secureShape monitoring, observability and alerting across the engineering organisation (e.g. Datadog, CloudWatch) so issues are detected and resolved before customers are impacted4Lead incident management practices - on-call, triage ...

Devops Engineer

Hiring Organisation
ISR RECRUITMENT LIMITED
Location
Nationwide, United Kingdom
Employment Type
Contract
Contract Rate
£475 - £500/day (Outside IR35)
SAML Application Technologies: React.js | Angular.js You will work extensively with Terraform, Kubernetes, GitHub Actions, Flux and GitOps, alongside cloud networking, security, identity, API management, observability and event-driven architectures. Role and Responsibilities: Deploy and manage infrastructure across multiple AWS cloud environments using Terraform and Infrastructure as Code (IaC). Deploy … manage and maintain cluster configurations. Securely configure and deploy AWS API Gateway and associated API capabilities. Design and implement effective operational monitoring, alerting and observability solutions. Design and support cloud networking across AWS, Azure and GCP, including VPCs, VNets, subnets, routing, load balancing and DNS. Deploy and operate event-driven ...

Lead DevOps Engineer

Location
Greater London, England, United Kingdom
secure supply chain configurations. Establish IaC and GitOps standards, adding automated testing to every infrastructure change. Prototype agentic infrastructure components, including deployment and observability platforms in service meshes. Contribute to the Kong AI Gateway, including Dataplane deployments, ACM/SSL integration, and observability via DataDog Champion DevSecOps maturity by embedding … Data, and AI teams to shape DevOps and AI platform architectures with regulatory compliance. Stay ahead of CNCF and AI ecosystem innovations, from eBPF observability to agent‐aware orchestration. What We’re Looking For We value diverse backgrounds and paths to expertise. If you bring solid experience in several ...

Operations Engineering Lead

Hiring Organisation
WTW
Location
Greater London, United Kingdom
Employment Type
Full Time
base platform operations team. You will help set direction for shared infrastructure and engineering operations - the production hosting, security perimeter, release pipelines, observability, and developer tooling that every engineering squad depends on to ship safely. This is a role that includes Line Management responsibility for your team … performance, and cost-efficiency across production environments Oversee Infrastructure as Code practices (e.g. Terraform) and ensure environments are consistent, auditable, and secure Shape monitoring, observability and alerting across the engineering organisation (e.g. Datadog, CloudWatch) so issues are detected and resolved before customers are impacted4 Lead incident management practices - on-call ...

Devops SRE

Location
Greater London, England, United Kingdom
shared Kubernetes services such as CoreDNS , cert‐manager , Dynatrace , Cloudability , and Infoblox . Familiarity with OPA Gatekeeper for policy enforcement and tenant isolation. Security, Observability & Performance Strong security mindset with a proven track record of designing secure, resilient cloud‐native systems. Experience implementing observability stacks including Prometheus , Dynatrace , and OpenTelemetry ...

Staff Software Engineer - Databases SRE | UK | Remote

Hiring Organisation
Grafana Labs
Location
United Kingdom
Salary
£ 70 K
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … observability and observability to AI, giving teams (and their agents) unified visibility so they can see, understand, and act on all their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. ...

Senior Cloud Architect (all genders)

Hiring Organisation
Lam Research
Location
Villach, Kärnten, Austria
Employment Type
Permanent
Salary
EUR Annual
guardrails. Developer Platform & Experience Build an internal developer platform (IDP) that abstracts complexity and provides self-service: project scaffolding, environment creation, CI/CD, observability, and secrets management. Curate platform catalogs/registries (Terraform module registry, container base images, reusable pipeline templates). Partner with product/engineering teams … Policy, OPA, Conftest) and continuous compliance. Build secure-by-default blueprints: network segmentation, private endpoints, encryption, vulnerability mgmt, and SBOM/SLSA practices. SRE, Observability, and Operations Embed SRE practices: SLOs, error budgets, incident response, postmortems, chaos/gamedays. Standardize observability: logs, metrics, traces, dashboards, and alerts (e.g., Azure Monitor ...

Devops Engineer - SC Cleared

Location
Leeds, England, United Kingdom
/CD pipelines for a public sector client. You'll work with a reasonable degree of autonomy, contributing to infrastructure automation, deployment reliability, and observability, while collaborating closely with senior engineers and the wider delivery team. Key Responsibilities Design, build, and maintain infrastructure-as-code using Terraform and Ansible across … support security standards, change management, and release governance processes Contribute to sprint planning, refinement, and other agile ceremonies Support continuous improvement of automation, observability, and platform reliability Required Skills & Experience Experience in a DevOps, Platform Engineering, or Infrastructure Engineering role Hands-on experience with Terraform and Ansible for infrastructure ...

Lead Site Reliability / DevOps Engineer

Location
Auchentibber, Scotland, United Kingdom
practices within an application or platform Fluency in at least one programming language such as (e.g., Java, Python, Go, etc.) Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency … high-availability services Deep understanding of distributed system design principles, networking (TCP/IP, DNS, load balancing), Linux internals. Contributions to open-source observability or telemetry projects. Experience working with agent control planes and management protocols. Hands-on knowledge of OpAMP is highly desirable. #J-18808-Ljbffr ...

GenAI Platform Operations Lead

Location
United Kingdom
security, scalability, and operational excellence of the Oasis Platform. Working at the intersection of Operations and Platform Engineering, you'll drive improvements in automation, observability, DevOps practices, and AI-powered operations while leading and mentoring a high-performing team. This role is ideal for an experienced SRE, DevOps, Platform Engineering … platforms, particularly AWS, EKS/ECS and containerised workloads in production environments. Site Reliability Engineering (SRE) & Operational Excellence – Proven experience improving platform reliability, resilience, observability and incident management through SRE principles, automation and data-driven operational practices. DevOps & Automation Mindset – Experience with CI/CD pipelines (Jenkins, Azure DevOps), Infrastructure ...

Senior AWS Platform Engineer (AI Platform)

Hiring Organisation
Sanderson Recruitment
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£550 - £600 per day
platform capability on AWS yourself. Bedrock in particular - model access, guardrails, knowledge bases, agentic patterns - plus the surrounding tooling (SageMaker, vector stores, evaluation and observability). Expert-level Infrastructure as Code (Terraform strongly preferred; CloudFormation/CDK welcome) and strong scripting (Python, Bash) for complex automation and self-service tooling … grade, self-service platform: Amazon Bedrock and the surrounding tooling (model access and routing, guardrails, knowledge bases/RAG patterns, agentic workflows), with the observability, cost controls and security posture required to run it in a regulated environment (AI Experience not essential but highly beneficial) Reasonable Adjustments: Respect and equality ...

Software Engineer III - AI/ML Platform Reliability

Location
Auchentibber, Scotland, United Kingdom
enhance the reliability and scalability of AI/ML platforms and applications to accommodate fast-growing demands. Own NFRs and develop tooling for observability, security, resilience, infrastructure management and operations excellence. Build and maintain scalable infrastructure to support the deployment and operation of large-scale AI platforms and apps. Build … architecture. Experience building large scale infrastructure and and cloud-native delivery practice in Google Cloud, AWS, or Azure and Terraform. Extensive experience implementing advanced observability using tools like Open Telemetry, Dynatrace, Grafana, and/or cloud-native services. Systematic problem-solving and troubleshooting skills in a complex system. Hands ...

GCP Cloud Architect

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
rehost, replatform, refactor, repurchase, retire, retain, and relocate. Broad hands-on knowledge of core Google Cloud services across compute, containers, storage, databases, networking, security, observability, DevOps, and cost management. In-depth experience designing Google Cloud landing zones and cloud foundations, including resource hierarchy, projects, billing, IAM, shared … methodologies as the market evolves. Ability to identify opportunities to use automation, AI-assisted engineering, AIOps, and cloud-native tooling to improve reliability, observability, delivery speed, and operational efficiency. Excellent English verbal and written communication skills, with the ability to engage effectively with both technical and non-technical stakeholders. Ability ...

Lead DevSecOps Engineer

Hiring Organisation
AECOM
Location
London, United Kingdom
Salary
£ 70 K
DevSecOps capabilities that help engineers build, release, and operate services across our multi-cloud platform. You will shape CI/CD, Infrastructure as Code, observability, reliability, cloud governance, and secure engineering across Microsoft Azure and Google Cloud Platform (GCP). You will also manage cloud-vendor service reviews, security advisories … roadmap discussions.Evolve CI/CD pipelines and reusable workflows that enable fast, reliable, and repeatable testing and deployment, with automated quality and security controlsBuild observability and operational readiness through monitoring, logging, tracing, runbooks, and incident response, improving the reliability and performance of production servicesDefine and track service reliability objectives ...

Google Cloud Platform (GCP) Architect

Location
Glasgow, Scotland, United Kingdom
operational excellence by proactively identifying patterns in system failures, operational metrics, and data; designing and implementing systematic improvements to system reliability, performance, and observability Evaluate and lead vendor/technology assessments - leading sessions with external vendors, startups, and internal teams to drive outcomes-oriented evaluation of architectural designs, technical credentials … automation Frameworks on Kubernetes, including authoring reconciliation loops, admission controllers, webhooks, and custom controllers; strong understanding of Kubernetes internals, API machinery, RBAC, multi-tenancy, observability, and operational best practices for production environments. In-depth Google Cloud development experience -architecting, building, deploying, and operating production cloud-native workloads using core ...

Principal Java Engineer

Location
Wallingford, England, United Kingdom
continuous improvement. Production systems are reliable, observable and operationally excellent Lead root cause analysis and resolution of complex production issues. Drive improvements in system observability, monitoring and operational performance. Ensure applications are designed and operated to meet reliability, availability and performance targets. Partner with Operations, DevOps and QA teams … Claude, Codex, Gitlab Duo, etc) REST APIs, OpenAPI, Microservices, Event-driven architecture (RabbitMQ) Containers, Docker, AWS, Linux CI/CD with GitLab Pipelines & Jenkins Observability: logging, metrics and monitoring MySQL, Apache Solr Front-end UI (e.g. Angular) Person Specification Strategic and systems-thinking mindset Excellent communication and stakeholder management skills ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
architecture and technical decision‐making as our platform evolves Champion engineering excellence through automated testing, CI/CD, code reviews and pair programming Use observability and monitoring to understand service health, troubleshoot complex distributed systems and improve reliability Mentor and support other engineers, sharing knowledge and helping raise technical capability … modern CI/CD practices Strong understanding of automated testing, code quality, security and software engineering best practices Experience troubleshooting distributed systems and using observability to improve performance and reliability Strong communication and collaboration skills, with experience contributing to technical decisions Experience mentoring and supporting other engineers through code reviews ...

Cloud Platform Engineer

Location
Greater London, England, United Kingdom
both human and machine/agent identities; secure by default, least privilege, secrets management and continuous compliance. Apply SRE practices - SLOs/SLIs, observability, capacity planning, resilience and blameless incident management - to keep the platform reliable and cost-efficient. Partner with data engineering to design and optimise data pipelines, data … workloads and data ‐ including secrets management, least‐privilege IAM and machine/workload identity. Solid programming/scripting ability (e.g. Python, Go) and strong observability, reliability and cost‐optimisation practices. Desirable requirements: Experience working as a Site Reliability Engineer (SRE) with SLOs/SLIs, error budgets and incident management. ...

Sr Lead AI Platform Engineer

Hiring Organisation
JP Morgan Chase
Location
Glasgow, Lanarkshire, United Kingdom
Salary
£ 70 K
Tech.Job responsibilitiesOwns the design and build of the team's platform: deployment pipelines, model serving, containerisation, orchestration, and environment managementSets the standard for reliability, observability, and operational excellence across the team's production AI/ML servicesBuilds the tooling and paved paths that let AI engineers ship agentic … track record of building deployment and release automationExperience serving, scaling, and monitoring ML models or data-intensive services in production (MLOps)Practical experience with observability tooling (metrics, logging, tracing) and production incident responseExperience operating ML/LLM workloads in production (LLMOps, inference reliability, cost/performance management)Strong communication skills ...

Sr Lead AI Platform Engineer

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
Owns the design and build of the team's platform: deployment pipelines, model serving, containerisation, orchestration, and environment management Sets the standard for reliability, observability, and operational excellence across the team's production AI/ML services Builds the tooling and paved paths that let AI engineers ship agentic … record of building deployment and release automation Experience serving, scaling, and monitoring ML models or data-intensive services in production (MLOps) Practical experience with observability tooling (metrics, logging, tracing) and production incident response Experience operating ML/LLM workloads in production (LLMOps, inference reliability, cost/performance management) Strong communication ...

Technical Architect - UK Public Sector | eSC/eDV

Location
Leicester, England, United Kingdom
with cloud and platform teams to ensure solutions are optimised for AWS and modern cloud environments. Promote best practice across software architecture, security, performance, observability and resilience. Support architecture governance through design reviews, peer reviews and architecture assurance activities. Drive adoption of modern engineering practices including CI/CD, Infrastructure … architectures and messaging technologies. Knowledge of Kubernetes, EKS and cloud‐native application platforms. Familiarity with Infrastructure as Code using Terraform or CloudFormation. Experience with observability tooling such as CloudWatch, Prometheus, Grafana or ELK. Experience within large‐scale enterprise, regulated or security‐conscious environments. Exposure to architecture governance, design authorities ...

Technical Architect - UK Public Sector | eSC/eDV

Location
Hursley, England, United Kingdom
with cloud and platform teams to ensure solutions are optimised for AWS and modern cloud environments. Promote best practice across software architecture, security, performance, observability and resilience. Support architecture governance through design reviews, peer reviews and architecture assurance activities. Drive adoption of modern engineering practices including CI/CD, Infrastructure … architectures and messaging technologies. Knowledge of Kubernetes, EKS and cloud‐native application platforms. Familiarity with Infrastructure as Code using Terraform or CloudFormation. Experience with observability tooling such as CloudWatch, Prometheus, Grafana or ELK. Experience within large‐scale enterprise, regulated or security‐conscious environments. Exposure to architecture governance, design authorities ...

Lead Azure Engineer

Location
Sheffield, England, United Kingdom
building CI/CD pipelines with automated quality gates such as unit and integration tests, SAST/DAST where applicable, and dependency scanning. Drive observability and production excellence by implementing logging, metrics, and tracing, creating actionable alerts, defining SLOs and SLIs, and leading incident response and post‐incident improvement activities. … Azure DevOps for CI/CD. The role involves close collaboration with product, architecture, UX, QA, and platform teams in a setting that emphasises observability, security, and production excellence. The engineering culture promotes inclusive teamwork, high code quality, and automation across the delivery lifecycle. Location Sheffield, UK Rate/Salary ...