2,076 to 2,100 of 4,601 Observability Jobs

Senior .NET Engineer – Hybrid Payments Platform (London)

Location
Greater London, England, United Kingdom
scheme integrations Payment compliance Client-facing platforms Collaborate closely with engineering teams to define and evolve system architecture. Drive best practices in security, resilience, observability and operational excellence. Use AI-assisted development tools effectively and pragmatically to improve engineering productivity. Mentor and support other engineers through code reviews, design discussions … distributed and event-driven architectures. Hands-on experience dealing with: Idempotency Message ordering Retries and failure recovery Strong understanding of security, availability, reliability and observability principles. Experience using monitoring and telemetry tools such as Grafana, Application Insights or similar platforms. Strong SQL Server experience, including schema design, performance tuning ...

Staff Backend Engineer - Data Platform

Location
Greater London, England, United Kingdom
drive the technical vision and implementation for our foundational data platform — from experimentation, event ingestion pipelines to our data lake, governance frameworks, and data observability and real-time analytics capabilities. You will work closely with data scientists, machine learning engineers, backend teams, and product leaders, providing deep technical expertise … analytics. Champion engineering excellence, setting high technical standards and advocating for best practices in system design, maintainability, performance, and privacy. Lead efforts in data observability, governance, and privacy-by-design principles, ensuring their robust implementation across the organization. Mentor and coach engineers, elevating the technical capabilities of the team ...

Senior Platform Software Engineer - SRE

Location
Greater London, England, United Kingdom
doing: The Senior Platform Software Engineer role in our SRE team combines software engineering practices with cloud infrastructure, distributed systems patterns, storage systems and observability to deliver on a wide range of projects - ranging from tooling to core Platform services which serve production traffic. Own the availability and performance … mission-critical services and build automation to prevent problems recurrence. Improve the system’s scalability, observability, and alerting. Build tooling to improve our platform and accelerate the overall software development. Practice sustainable incident response and blameless postmortems. Collaborate with product teams to help them tackle technical issues and design ...

Software Reliability Engineer

Location
Greater London, England, United Kingdom
secrets management, authentication and authorisation Simplifying and automating application deployment processes, including automated database changes with Liquibase and installed software with Ansible. Introducing standard observability patterns. Overhauling exception handling and logging. Required Qualifications : Several years of experience with Test-Driven Development (TDD) using multiple test frameworks. Really excellent understanding … languages and/or TypeScript. Highest quality coding skills An excellent understanding of safe practices for critical systems, including deployment architecture and observability At home with multiple continuous integration and deployment systems Excellent understanding of secure coding practices Competent with Docker and Openshift Actively embracing AI coding Comfortable with build ...

Solution Architect AWS-to-GCP Migration (Hybrid 3 Days/Week)

Hiring Organisation
Bitsoft International, Inc
Location
New York, United States
Employment Type
Any
Salary
USD Annual
/Kubernetes, Containers, Cloud Run, APIs & Microservices Cloud Storage, Cloud SQL, Pub/Sub & Event-Driven Architecture Terraform, IaC, CI/CD & DevSecOps Observability, DR/Resiliency, Cost Optimization & API Management Strong technical leadership, architecture governance and distributed/offshore team management ...

ML / Backend Engineer @ Sqwish

Location
Cambridge, England, United Kingdom
matter at once. You enjoy building systems that are clean enough to reason about, but pragmatic enough to ship. You care about tests, observability, and operational safety, but you do not hide behind process. Problems you’ll tackle Building low-latency optimisation APIs that sit on the critical path … Python services across serving, workers, training workflows, and internal tooling Working with Postgres, Redis, queues/streams, migrations, and event-driven workflows Making reliability, observability, and deployment safety part of the product from the beginning Core responsibilities Write production-grade Rust and Python services Design clean domain boundaries around requests ...

Engineering Lead

Location
Greater London, England, United Kingdom
share ownership of technical direction. Shape how the team works, not just what it builds Drive improvements to developer experience, CI/CD, observability and release practices that make the whole team faster and more confident. Make pragmatic trade-offs that balance reliability, performance and cost across AWS services. Decide … modern serverless (Lambda, API Gateway, SQS, EventBridge, DynamoDB, S3, CloudWatch) or other cloud platforms. A pragmatic approach to technical decisions, balancing reliability, observability, performance and cost, and bringing engineers along on the reasoning. Track record of raising engineering standards as a force-multiplier, through code reviews, pairing, design discussions ...

Technical Lead

Location
Greater London, England, United Kingdom
Support engineers through complex technical challenges without becoming the decision-maker or implementation owner for every issue. Promote strong engineering practices including testing, automation, observability, security, documentation and sustainable software development. Encourage constructive technical challenge, knowledge sharing and continuous learning across the team. Work with the Engineering Manager to identify … sustainability of the Marketing & Commercial Data capabilities throughout their lifecycle. Ensure solutions are designed and engineered with appropriate consideration for security, resilience, scalability, observability, maintainability and supportability. Work with the Engineering Manager and Commercial Platform team to ensure new and changed capabilities are operationally ready and can be effectively supported ...

Platform Principal Engineer

Location
City Of London, England, United Kingdom
self-service capabilities. Upskill and Mentor: Transition the in-house engineering team into a high-performing internal platform team throughout the platform build process. Observability: Design and implement enterprise-grade logging, metrics, and tracing for Kubernetes at scale. IaC Leadership: Implement and manage Infrastructure as Code to a senior standard … Terraform/Open Tofu module design. (MUST) Kubernetes Engineering: GitOps (Argo CD/Flux), secrets management, ingress/mesh, and OPA/Gatekeeper. (MUST) Observability: OpenTelemetry (MUST) Tooling: Spacelift, Atlantis, or Terraform Cloud (Desired) Governance: EPAC (Enterprise Policy as Code) (Desired) What You'll Bring To Us: Recent, hands ...

AI Engineer

Location
Greater London, England, United Kingdom
full-stack AI engineering role, with applied AI product delivery at its core. You will build AI workflows, tools and integrations, evaluation and observability capabilities, together with the APIs, services and user interfaces needed to ship them reliably. This is not a research-only role. You will apply agreed enterprise … feedback loops. Manage prompts, model configuration, tool schemas and routing as tested product assets, including fallbacks and cost/latency trade‐offs. Implement AI observability through traces, logs, metrics and evaluation results, so quality, reliability, failure modes, latency and cost are visible and can be improved. Build clear APIs ...

Agent Engineer

Location
Greater London, England, United Kingdom
concept through to production. Take part in technical design reviews, planning sessions, and code reviews to continuously improve system quality. Contribute to infrastructure and observability practices alongside the Engineering team — you won't own this alone, but you'll be expected to care about how your services run in production. … design. Nice to Have Experience with Node.js frameworks like NestJS or Express. Hands-on experience with Terraform, or infrastructure-as-code tooling. Experience with observability platforms like Datadog (metrics, tracing, alerting). Exposure to DynamoDB or other NoSQL databases at scale. Experience with distributed or event-driven architectures ...

Embedded Software Engineer

Location
Greater London, England, United Kingdom
vulnerability management and secure provisioning. Edge & Cloud Integration Integrate devices with cloud IoT platforms and backend services. Define and improve telemetry, health monitoring, and observability for deployed devices. Support reliable operation of connected devices in the field, including debugging fleet issues and improving resilience. Compliance & Testing Support testing and validation …/CD workflows for embedded software. Practical debugging experience using lab and software tools such as logic analysers, protocol analysers, network sniffers, or observability platforms. Strong problem-solving skills and the ability to work across hardware and software boundaries. Nice to Have Experience with low-power wireless technologies such ...

Backend Engineers (Ruby)

Location
Greater London, England, United Kingdom
technical insight into upcoming work and helping pull the team together to ship it Deliver your work using agile methodologies and tools like tests, observability, A/B tests, and feature flags Mentor colleagues to help them grow as engineers, and actively support their development Contribute to cross-cutting concerns … Rails, building data models, APIs, and business logic services in a production monolith Comfortable delivering with agile methodologies and practices like automated testing, observability, A/B testing, and feature flags A collaborative approach — you work well with product, design, and analytics partners, and enjoy shaping work beyond just ...

Principal Network Engineer

Location
Greater London, England, United Kingdom
across firewalls, NAT, VPN, security policies, and multi-tenant segmentation. Design highly available and scalable security architectures appropriate for mission-critical AI infrastructure. Reliability, Observability & Operations Lead complex technical escalations and root-cause analysis for network performance, reliability, and stability issues. Establish measurable SLOs and operational standards for network services. … technical direction for network observability, telemetry, monitoring, and alerting. Ensure clear visibility into fabric health, traffic patterns, performance, and capacity. Develop runbooks, automation, and engineering improvements that systematically reduce operational toil. Act as a senior 3rd/4th line escalation point for complex networking issues. Network Data & Configuration Management Ensure ...

GenAI Engineer

Location
Greater London, England, United Kingdom
indexing, retrieval policies, grounding, guardrails) and agent frameworks. Take basic infra ownership on GCP (or AWS/Azure): networking, autoscaling, CI/CD, IaC, observability, and cost tuning. Participate in on‐call for your area and drive root‐cause analysis with crisp follow‐ups. 15% Collaborate Pair with back … series analysis (forecasting, change‐point, drift). Cloud & ops: Basic infra ownership on GCP (or AWS/Azure): networking, autoscaling, CI/CD, IaC, observability, and cost control. Communication: You explain results clearly, align stakeholders, and write crisp docs. Bonus points DevOps wizardry; GPU/accelerator experience. Multimodal pipelines (text ...

Database Reliability Engineer

Hiring Organisation
Starling Bank
Location
London, UK
Employment Type
Full-time
Cross-Cloud Portability: Use CNPG and cloud-native patterns to ensure our database layer remains provider-agnostic, allowing seamless deployment across AWS and GCPEvolve Observability & Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will ensure we have the visibility to detect performance regressions and health … excited by the challenge of "multi-everything"—multi-tenant, multi-region, and multi-cloud—while ensuring rigorous data integrity and mobilityA Security & Observability Mindset: You believe security is paramount. You focus on building deep observability (Prometheus/Grafana/OpenTelemetry/Humio) and automated guardrails so the fleet is secure ...

Software Engineer, Agentic AI

Location
Cambridge, England, United Kingdom
product and platform capabilities for Roku TV. You will own the full lifecycle of agent development from prototyping and architecture through orchestration, evaluation, deployment, observability, and continuous improvement. You will contribute directly to Roku's AI strategy by engineering reusable components, optimizing agent workflows, and ensuring strong real-world performance … systems around them. Create reusable agent templates, modular components, and paved-path patterns that accelerate adoption across teams and use cases. Establish strong evaluation, observability, and monitoring for conversation quality, task success rate, latency, cost, and overall system performance. Build safeguards that improve production readiness and reliability, including testing pipelines ...

Data Platform DevOps Analyst

Location
Reading, England, United Kingdom
journey supporting the migration from Teradata to Databricks, validating production readiness, improving monitoring and automation capabilities and helping embed modern DataOps practices across deployment, observability and support processes. What You’ll Bring Here at Primark, we want everyone to feel valued – so please bring your authentic self to work … platforms and DataOps practices with exposure to Azure DevOps, Databricks, dbt Cloud, Azure data services, SQL, Python, ETL/ELT processes, job scheduling, monitoring, observability and deployment automation. Experience supporting production environments and service operations including monitoring, alerting, incident management, ticketing systems, release management, environment management, change governance and enterprise ...

Senior Developer Experience Security Engineer

Location
Greater London, England, United Kingdom
experience across Motorway.. We have recently built and rolled out a new container platform on top of AWS Fargate, and are currently enhancing our observability, reliability, and developer-focused tooling. We will continue to build and evolve secure, standardised platform capabilities that reduce cognitive load and help teams ship faster … experience across Motorway.. We have recently built and rolled out a new container platform on top of AWS Fargate, and are currently enhancing our observability, reliability, and developer-focused tooling. We will continue to build and evolve secure, standardised platform capabilities that reduce cognitive load and help teams ship faster ...

Technical Support Engineer

Location
Paignton, England, United Kingdom
major escalations and transformation initiatives. Lead or contribute to the design and deployment of scalable support workflows, including automated log parsing, triage tooling, and observability pipelines. Identify trends in support data to drive systemic product or process improvements. Collaborate cross-functionally to evolve support readiness standards for new product releases … lead technical investigations, influence engineering teams, and manage high-pressure escalations independently. Experience with support and engineering tools such as Salesforce, Jira, Git, and observability stacks. Desirable Skills & Knowledge Knowledge of GNSS simulation systems, timing/sync architectures, or satellite navigation protocols. Familiarity with DevOps, CI/CD, or site ...

Software Engineer III - Fullstack (Java, React, Python and AI) Engineer

Location
Greater London, England, United Kingdom
with product, marketing, and partners to translate requirements into well-designed technical solutions Improve engineering excellence through code reviews, test automation, CI/CD, observability, performance tuning, and operational best practices Ensure solutions meet security, privacy, and compliance expectations including consent management, data minimization, and access controls Required Qualifications, Capabilities … with campaign management and attribution systems Background in AI-enabled content generation, experimentation, or workflow automation Skills in test automation, CI/CD, and observability tools Experience optimizing performance and scalability Familiarity with consent management and data minimization practices Ability to drive innovation in personalization and measurement Experience working ...

Managed Service Operations - Head of Practice

Location
Greater London, England, United Kingdom
service outcomes. Key responsibilities You will lead the development and maturity of Made Tech’s operational capabilities incident, problem, and change management; monitoring and observability; automation and AIOps; governance; operational playbooks; runbooks; service health metrics; and 24/7/365 support patterns. You will ensure our teams have … ITIL practices blended with modern DevOps, SRE, Agile and platform‐engineering approaches. Broad technical awareness across cloud platforms, application architectures, data platforms, networks, observability tooling, security‐by‐design, and automation. Ability to create and evolve operational standards, playbooks, governance models, templates, and frameworks that drive consistency, stability, and efficiency. Skilled ...

Lead Mobile Engineer

Location
Greater London, England, United Kingdom
deployment processes, including CI/CD pipelines and app store submissions for Google Play and Apple App Store.* Champion automated testing, app stability, observability, and the responsible use of AI to improve developer productivity and software quality.**Knowledge, Skills and Behaviours**Essential* 7+ years of professional software engineering experience, with … experience using mobile testing frameworks and automated testing approaches.* Experience owning mobile CI/CD pipelines, build automation and app store release processes.* An observability mindset, with experience using crash reporting, performance monitoring and analytics tools to improve app reliability.* Comfortable working with Git, Jira, Confluence and modern agile engineering ...

Senior Data Engineer

Location
Greater London, England, United Kingdom
Insight teams to maintain reliable, accurate and meaningful data models as our products evolve Establish and promote good engineering practices across data pipelines, modelling, observability, performance and reliability Ensure data quality and integrity throughout our pipelines, identifying and addressing issues before they impact downstream users Provide technical guidance to engineers … data transformation skills, particularly using technologies such as Python, Spark and SQL Experience working with cloud-based data platforms Experience designing systems for reliability, observability, performance and data quality Strong understanding of data modelling and the principles behind well-structured analytical data Experience making technical design decisions and taking ownership ...

Software Engineer III - Fullstack (Java, React, Python and AI) Engineer

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
automationCollaborate with product, marketing, and partners to translate requirements into well-designed technical solutionsImprove engineering excellence through code reviews, test automation, CI/CD, observability, performance tuning, and operational best practicesEnsure solutions meet security, privacy, and compliance expectations including consent management, data minimization, and access controlsRequired Qualifications, Capabilities, and Skills … Skills: Experience with campaign management and attribution systemsBackground in AI-enabled content generation, experimentation, or workflow automationSkills in test automation, CI/CD, and observability toolsExperience optimizing performance and scalabilityFamiliarity with consent management and data minimization practicesAbility to drive innovation in personalization and measurementExperience working in cross-functional teamsJ.P. Morgan ...