1,701 to 1,725 of 3,973 Observability Jobs

Senior Infrastructure Platform Engineer Veeam / VMware / Hyper-V

Hiring Organisation
100% IT Recruitment Ltd
Location
United Kingdom
Employment Type
Permanent, Work From Home
Salary
£60,000
major incident recovery, particularly around backup, restore and platform availability. Drive automation to reduce manual processes and improve operational efficiency. Develop and maintain monitoring, observability and alerting using tools such as Grafana. Coordinate the day-to-day priorities of the Platform Engineering team. Maintain engineering standards, documentation and operational procedures. … Azure DevOps-focused role. Experience supporting highly available production infrastructure. Strong troubleshooting and problem-solving skills across enterprise infrastructure. Experience with monitoring and observability platforms such as Grafana. Experience automating operational tasks using PowerShell, scripting or similar technologies. Excellent understanding of backup, disaster recovery and platform resilience. Ability to coordinate ...

SRE Technical Lead

Hiring Organisation
83zero Limited
Location
Wokingham, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
senior technical escalation point for major incidents and high-risk releases. Lead blameless post-incident reviews and ensure measurable service improvements. Define and oversee observability, monitoring and capacity management practices. Ensure SRE approaches align with security, governance and compliance requirements. Mentor and coach senior engineers, helping to improve SRE maturity … OpenShift. Experience designing and supporting hybrid and multi-cloud platforms. Experience with service mesh technologies such as Istio. Strong hands-on experience with observability tooling including Prometheus, Grafana, Loki, Tempo and OpenTelemetry. Infrastructure as Code and GitOps expertise using tools such as Helm, Kustomize, ArgoCD and Tekton. Experience building ...

Senior Microsoft Power Platform and AI Developer

Location
Greater London, England, United Kingdom
agents use trusted knowledge, approved tools, Model Context Protocol (MCP) services, approvals, human hand-offsand safe failure paths. You will ensure that evaluation, observability, identity,permissionsand data boundaries are built into delivery rather than added later. You will review code and designs, mentor developers, improve engineeringpracticesand support live services. … JavaScript,C#or Python, with experience extending low-code services appropriately. Strong software engineering practice, including testing, version control, code review, documentation, CI/CD, observability and supporting live services. Ability to lead technical decisions for medium-to-high complexity work, communicate trade-offs and elevate architecture or security risks appropriately. ...

Senior Full stack Developer - Birmingham - Perm

Location
Birmingham, England, United Kingdom
recurring technical problems and implementing long-term solutions. Improving platform reliability, resilience, and overall product quality. Performing application profiling, performance tuning, and optimisation. Enhancing observability, monitoring, alerting, and diagnostic capabilities. Working with engineering teams to improve development practices and technical standards. Reducing technical debt and identifying opportunities for platform improvement. … Strong communication skills and the ability to collaborate effectively across engineering teams. Desirable Experience working on SaaS platforms or cloud-based applications. Exposure to observability and monitoring tools. Experience with performance profiling and optimisation techniques. Knowledge of scalability, resilience, and reliability engineering principles. Familiarity with CI/CD pipelines ...

Sales Specialist (UK/I) - DevOps & DevEx

Hiring Organisation
Adaptavist
Location
London, UK
Employment Type
Full-time
secured, deployed and operated across UK/I.This role acts as a specialist advisor focused on Developer Experience, Platform Engineering, DevSecOps, Cloud Native Engineering, Observability and AI-enabled software delivery. This role requires a deep understanding of modern software engineering practices, developer productivity, platform operating models and software delivery transformation. … aligned to DevOps, Developer Experience and software delivery transformation initiatives. Identify, qualify and pursue opportunities across Developer Experience, Platform Engineering, DevSecOps, Cloud Native Engineering, Observability, and AI-assisted Software Delivery and Software Supply Chain Security. Develop customer-specific hypotheses, transformation opportunities and executive points of view that align engineering challenges ...

Data Architect / Engineering Lead

Location
Greater London, England, United Kingdom
deployment predictable through Git-based development, CI/CD for the warehouse and semantic layer, automated provisioning and consistent environments. Set standards for testing, observability, backfills, quarantine and the day-to-day operation of analytics as software. Lead the Data Engineering discipline through standards, technical direction, mentoring, hiring input … attribute-level security, PII classification and GDPR-sensitive data. Experience treating analytics as software, including Git, CI/CD, automated testing, environment parity, observability, SLOs and cost awareness. Functional or discipline leadership experience, setting standards, raising the technical bar and mentoring without relying on formal line management. The ability ...

AI Engineer IRC302970

Location
United Kingdom
products. As a Senior/Lead AI Engineer, you will own and drive the development of our core agentic frameworks, evaluation pipelines, and observability tooling to ensure they operate with safety, trust, and intelligence at scale. You won’t just be integrating AI into a product, you will lead … CrewAI is a plus. Production Python engineering : Clean, modular, testable, maintainable code — you care about system reliability as much as model output. Evaluation & observability : Practical experience instrumenting tracing (LangSmith, Arize) and building CI/CD pipelines built specifically for LLMs. Engineering foundations : Distributed systems, scalable data pipelines (Kafka ...

AWS Platform Engineer

Location
Langley Mill, England, United Kingdom
pipelines for infrastructure and platform services Implement AWS security controls, IAM patterns, secrets management and secure networking Establish monitoring, logging, alerting and operational observability across the platform Implement tagging, cost visibility, budgets and FinOps controls to help teams understand and manage cloud consumption Work closely with architecture, security, data engineering … connectivity between AWS and existing environments Embed security by design, including least-privilege access, encryption, secrets management, vulnerability management and policy enforcement Build platform observability using metrics, logs, traces, dashboards and automated alerting Establish backup, recovery, resilience and operational readiness patterns Implement cloud cost management through tagging standards, budgets, monitoring ...

Lead Quality Engineer

Location
Greater London, England, United Kingdom
Lead QA at Waracle, you will play a pivotal leadership role, shaping test strategies across multiple squads, establishing automation architectures, and championing observability and non-functional requirements. Operating at a strategic level, you’ll mentor engineers, build strong relationships with senior stakeholders, and influence the adoption of forward-thinking testing … platforms. Quality & Automation Architecture: Defining squad-level automation approaches and embedding robust quality gates within CI/CD pipelines to elevate software standards. NFR & Observability Championing: Defining comprehensive strategies for performance, security, and accessibility, ensuring systems are inherently testable and debuggable in production. Delivery & Team Wellbeing: Leading QA workstreams with ...

Senior .NET Engineer – Hybrid Payments Platform (London)

Location
Greater London, England, United Kingdom
scheme integrations Payment compliance Client-facing platforms Collaborate closely with engineering teams to define and evolve system architecture. Drive best practices in security, resilience, observability and operational excellence. Use AI-assisted development tools effectively and pragmatically to improve engineering productivity. Mentor and support other engineers through code reviews, design discussions … distributed and event-driven architectures. Hands-on experience dealing with: Idempotency Message ordering Retries and failure recovery Strong understanding of security, availability, reliability and observability principles. Experience using monitoring and telemetry tools such as Grafana, Application Insights or similar platforms. Strong SQL Server experience, including schema design, performance tuning ...

Staff Backend Engineer - Data Platform

Location
Greater London, England, United Kingdom
drive the technical vision and implementation for our foundational data platform — from experimentation, event ingestion pipelines to our data lake, governance frameworks, and data observability and real-time analytics capabilities. You will work closely with data scientists, machine learning engineers, backend teams, and product leaders, providing deep technical expertise … analytics. Champion engineering excellence, setting high technical standards and advocating for best practices in system design, maintainability, performance, and privacy. Lead efforts in data observability, governance, and privacy-by-design principles, ensuring their robust implementation across the organization. Mentor and coach engineers, elevating the technical capabilities of the team ...

Senior Platform Software Engineer - SRE

Location
Greater London, England, United Kingdom
doing: The Senior Platform Software Engineer role in our SRE team combines software engineering practices with cloud infrastructure, distributed systems patterns, storage systems and observability to deliver on a wide range of projects - ranging from tooling to core Platform services which serve production traffic. Own the availability and performance … mission-critical services and build automation to prevent problems recurrence. Improve the system’s scalability, observability, and alerting. Build tooling to improve our platform and accelerate the overall software development. Practice sustainable incident response and blameless postmortems. Collaborate with product teams to help them tackle technical issues and design ...

Software Reliability Engineer

Location
Greater London, England, United Kingdom
secrets management, authentication and authorisation Simplifying and automating application deployment processes, including automated database changes with Liquibase and installed software with Ansible. Introducing standard observability patterns. Overhauling exception handling and logging. Required Qualifications : Several years of experience with Test-Driven Development (TDD) using multiple test frameworks. Really excellent understanding … languages and/or TypeScript. Highest quality coding skills An excellent understanding of safe practices for critical systems, including deployment architecture and observability At home with multiple continuous integration and deployment systems Excellent understanding of secure coding practices Competent with Docker and Openshift Actively embracing AI coding Comfortable with build ...

Solution Architect AWS-to-GCP Migration (Hybrid 3 Days/Week)

Hiring Organisation
Bitsoft International, Inc
Location
New York, United States
Employment Type
Any
Salary
USD Annual
/Kubernetes, Containers, Cloud Run, APIs & Microservices Cloud Storage, Cloud SQL, Pub/Sub & Event-Driven Architecture Terraform, IaC, CI/CD & DevSecOps Observability, DR/Resiliency, Cost Optimization & API Management Strong technical leadership, architecture governance and distributed/offshore team management ...

ML / Backend Engineer @ Sqwish

Location
Cambridge, England, United Kingdom
matter at once. You enjoy building systems that are clean enough to reason about, but pragmatic enough to ship. You care about tests, observability, and operational safety, but you do not hide behind process. Problems you’ll tackle Building low-latency optimisation APIs that sit on the critical path … Python services across serving, workers, training workflows, and internal tooling Working with Postgres, Redis, queues/streams, migrations, and event-driven workflows Making reliability, observability, and deployment safety part of the product from the beginning Core responsibilities Write production-grade Rust and Python services Design clean domain boundaries around requests ...

Engineering Lead

Location
Greater London, England, United Kingdom
share ownership of technical direction. Shape how the team works, not just what it builds Drive improvements to developer experience, CI/CD, observability and release practices that make the whole team faster and more confident. Make pragmatic trade-offs that balance reliability, performance and cost across AWS services. Decide … modern serverless (Lambda, API Gateway, SQS, EventBridge, DynamoDB, S3, CloudWatch) or other cloud platforms. A pragmatic approach to technical decisions, balancing reliability, observability, performance and cost, and bringing engineers along on the reasoning. Track record of raising engineering standards as a force-multiplier, through code reviews, pairing, design discussions ...

Technical Lead

Location
Greater London, England, United Kingdom
Support engineers through complex technical challenges without becoming the decision-maker or implementation owner for every issue. Promote strong engineering practices including testing, automation, observability, security, documentation and sustainable software development. Encourage constructive technical challenge, knowledge sharing and continuous learning across the team. Work with the Engineering Manager to identify … sustainability of the Marketing & Commercial Data capabilities throughout their lifecycle. Ensure solutions are designed and engineered with appropriate consideration for security, resilience, scalability, observability, maintainability and supportability. Work with the Engineering Manager and Commercial Platform team to ensure new and changed capabilities are operationally ready and can be effectively supported ...

Platform Principal Engineer

Location
Greater London, England, United Kingdom
self-service capabilities. Upskill and Mentor: Transition the in-house engineering team into a high-performing internal platform team throughout the platform build process. Observability: Design and implement enterprise-grade logging, metrics, and tracing for Kubernetes at scale. IaC Leadership: Implement and manage Infrastructure as Code to a senior standard … Terraform/Open Tofu module design. (MUST) Kubernetes Engineering: GitOps (Argo CD/Flux), secrets management, ingress/mesh, and OPA/Gatekeeper. (MUST) Observability: OpenTelemetry (MUST) Tooling: Spacelift, Atlantis, or Terraform Cloud (Desired) Governance: EPAC (Enterprise Policy as Code) (Desired) What You'll Bring To Us: Recent, hands ...

AI Engineer

Location
Greater London, England, United Kingdom
full-stack AI engineering role, with applied AI product delivery at its core. You will build AI workflows, tools and integrations, evaluation and observability capabilities, together with the APIs, services and user interfaces needed to ship them reliably. This is not a research-only role. You will apply agreed enterprise … feedback loops. Manage prompts, model configuration, tool schemas and routing as tested product assets, including fallbacks and cost/latency trade‐offs. Implement AI observability through traces, logs, metrics and evaluation results, so quality, reliability, failure modes, latency and cost are visible and can be improved. Build clear APIs ...

Agent Engineer

Location
Greater London, England, United Kingdom
concept through to production. Take part in technical design reviews, planning sessions, and code reviews to continuously improve system quality. Contribute to infrastructure and observability practices alongside the Engineering team — you won't own this alone, but you'll be expected to care about how your services run in production. … design. Nice to Have Experience with Node.js frameworks like NestJS or Express. Hands-on experience with Terraform, or infrastructure-as-code tooling. Experience with observability platforms like Datadog (metrics, tracing, alerting). Exposure to DynamoDB or other NoSQL databases at scale. Experience with distributed or event-driven architectures ...

Embedded Software Engineer

Location
Greater London, England, United Kingdom
vulnerability management and secure provisioning. Edge & Cloud Integration Integrate devices with cloud IoT platforms and backend services. Define and improve telemetry, health monitoring, and observability for deployed devices. Support reliable operation of connected devices in the field, including debugging fleet issues and improving resilience. Compliance & Testing Support testing and validation …/CD workflows for embedded software. Practical debugging experience using lab and software tools such as logic analysers, protocol analysers, network sniffers, or observability platforms. Strong problem-solving skills and the ability to work across hardware and software boundaries. Nice to Have Experience with low-power wireless technologies such ...

Backend Engineers (Ruby)

Location
Greater London, England, United Kingdom
technical insight into upcoming work and helping pull the team together to ship it Deliver your work using agile methodologies and tools like tests, observability, A/B tests, and feature flags Mentor colleagues to help them grow as engineers, and actively support their development Contribute to cross-cutting concerns … Rails, building data models, APIs, and business logic services in a production monolith Comfortable delivering with agile methodologies and practices like automated testing, observability, A/B testing, and feature flags A collaborative approach — you work well with product, design, and analytics partners, and enjoy shaping work beyond just ...

Principal Network Engineer

Location
Greater London, England, United Kingdom
across firewalls, NAT, VPN, security policies, and multi-tenant segmentation. Design highly available and scalable security architectures appropriate for mission-critical AI infrastructure. Reliability, Observability & Operations Lead complex technical escalations and root-cause analysis for network performance, reliability, and stability issues. Establish measurable SLOs and operational standards for network services. … technical direction for network observability, telemetry, monitoring, and alerting. Ensure clear visibility into fabric health, traffic patterns, performance, and capacity. Develop runbooks, automation, and engineering improvements that systematically reduce operational toil. Act as a senior 3rd/4th line escalation point for complex networking issues. Network Data & Configuration Management Ensure ...

GenAI Engineer

Location
Greater London, England, United Kingdom
indexing, retrieval policies, grounding, guardrails) and agent frameworks. Take basic infra ownership on GCP (or AWS/Azure): networking, autoscaling, CI/CD, IaC, observability, and cost tuning. Participate in on‐call for your area and drive root‐cause analysis with crisp follow‐ups. 15% Collaborate Pair with back … series analysis (forecasting, change‐point, drift). Cloud & ops: Basic infra ownership on GCP (or AWS/Azure): networking, autoscaling, CI/CD, IaC, observability, and cost control. Communication: You explain results clearly, align stakeholders, and write crisp docs. Bonus points DevOps wizardry; GPU/accelerator experience. Multimodal pipelines (text ...

Database Reliability Engineer

Hiring Organisation
Starling Bank
Location
London, UK
Employment Type
Full-time
Cross-Cloud Portability: Use CNPG and cloud-native patterns to ensure our database layer remains provider-agnostic, allowing seamless deployment across AWS and GCPEvolve Observability & Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will ensure we have the visibility to detect performance regressions and health … excited by the challenge of "multi-everything"—multi-tenant, multi-region, and multi-cloud—while ensuring rigorous data integrity and mobilityA Security & Observability Mindset: You believe security is paramount. You focus on building deep observability (Prometheus/Grafana/OpenTelemetry/Humio) and automated guardrails so the fleet is secure ...