2,426 to 2,450 of 5,504 Permanent Observability Jobs

Software Reliability Engineer

Location
Greater London, England, United Kingdom
secrets management, authentication and authorisation Simplifying and automating application deployment processes, including automated database changes with Liquibase and installed software with Ansible. Introducing standard observability patterns. Overhauling exception handling and logging. Required Qualifications : Several years of experience with Test-Driven Development (TDD) using multiple test frameworks. Really excellent understanding … languages and/or TypeScript. Highest quality coding skills An excellent understanding of safe practices for critical systems, including deployment architecture and observability At home with multiple continuous integration and deployment systems Excellent understanding of secure coding practices Competent with Docker and Openshift Actively embracing AI coding Comfortable with build ...

Solution Architect AWS-to-GCP Migration (Hybrid 3 Days/Week)

Hiring Organisation
Bitsoft International, Inc
Location
New York, United States
Employment Type
Any
Salary
USD Annual
/Kubernetes, Containers, Cloud Run, APIs & Microservices Cloud Storage, Cloud SQL, Pub/Sub & Event-Driven Architecture Terraform, IaC, CI/CD & DevSecOps Observability, DR/Resiliency, Cost Optimization & API Management Strong technical leadership, architecture governance and distributed/offshore team management ...

Software Engineer-OSM

Hiring Organisation
Dynamic Systems, Inc
Location
Buda, Texas, United States
Employment Type
Permanent
Salary
USD Annual
databases, APIs, and MCPs provided by the core team to power those applications. • Own OSM application delivery and maintenance: design, implementation, testing, deployment, observability, documentation, user support, and issue triage. Coordinate with the help desk and core platform teams when issues cross system boundaries. • Work with the corporate Dev team … code, test behavior, explain their decisions, and own what ships. Software that acts on company systems follows real governance: appropriate access, managed secrets, testing, observability, and review before release. We value demonstrated work over titles or buzzwords. What we offer • A stable, established company with real commitment behind this work. ...

ML / Backend Engineer @ Sqwish

Location
Cambridge, England, United Kingdom
matter at once. You enjoy building systems that are clean enough to reason about, but pragmatic enough to ship. You care about tests, observability, and operational safety, but you do not hide behind process. Problems you’ll tackle Building low-latency optimisation APIs that sit on the critical path … Python services across serving, workers, training workflows, and internal tooling Working with Postgres, Redis, queues/streams, migrations, and event-driven workflows Making reliability, observability, and deployment safety part of the product from the beginning Core responsibilities Write production-grade Rust and Python services Design clean domain boundaries around requests ...

Engineering Lead

Location
Greater London, England, United Kingdom
share ownership of technical direction. Shape how the team works, not just what it builds Drive improvements to developer experience, CI/CD, observability and release practices that make the whole team faster and more confident. Make pragmatic trade-offs that balance reliability, performance and cost across AWS services. Decide … modern serverless (Lambda, API Gateway, SQS, EventBridge, DynamoDB, S3, CloudWatch) or other cloud platforms. A pragmatic approach to technical decisions, balancing reliability, observability, performance and cost, and bringing engineers along on the reasoning. Track record of raising engineering standards as a force-multiplier, through code reviews, pairing, design discussions ...

Technical Lead

Location
Greater London, England, United Kingdom
Support engineers through complex technical challenges without becoming the decision-maker or implementation owner for every issue. Promote strong engineering practices including testing, automation, observability, security, documentation and sustainable software development. Encourage constructive technical challenge, knowledge sharing and continuous learning across the team. Work with the Engineering Manager to identify … sustainability of the Marketing & Commercial Data capabilities throughout their lifecycle. Ensure solutions are designed and engineered with appropriate consideration for security, resilience, scalability, observability, maintainability and supportability. Work with the Engineering Manager and Commercial Platform team to ensure new and changed capabilities are operationally ready and can be effectively supported ...

Platform Principal Engineer

Location
City Of London, England, United Kingdom
self-service capabilities. Upskill and Mentor: Transition the in-house engineering team into a high-performing internal platform team throughout the platform build process. Observability: Design and implement enterprise-grade logging, metrics, and tracing for Kubernetes at scale. IaC Leadership: Implement and manage Infrastructure as Code to a senior standard … Terraform/Open Tofu module design. (MUST) Kubernetes Engineering: GitOps (Argo CD/Flux), secrets management, ingress/mesh, and OPA/Gatekeeper. (MUST) Observability: OpenTelemetry (MUST) Tooling: Spacelift, Atlantis, or Terraform Cloud (Desired) Governance: EPAC (Enterprise Policy as Code) (Desired) What You'll Bring To Us: Recent, hands ...

AI Engineer

Location
Greater London, England, United Kingdom
full-stack AI engineering role, with applied AI product delivery at its core. You will build AI workflows, tools and integrations, evaluation and observability capabilities, together with the APIs, services and user interfaces needed to ship them reliably. This is not a research-only role. You will apply agreed enterprise … feedback loops. Manage prompts, model configuration, tool schemas and routing as tested product assets, including fallbacks and cost/latency trade‐offs. Implement AI observability through traces, logs, metrics and evaluation results, so quality, reliability, failure modes, latency and cost are visible and can be improved. Build clear APIs ...

Agent Engineer

Location
Greater London, England, United Kingdom
concept through to production. Take part in technical design reviews, planning sessions, and code reviews to continuously improve system quality. Contribute to infrastructure and observability practices alongside the Engineering team — you won't own this alone, but you'll be expected to care about how your services run in production. … design. Nice to Have Experience with Node.js frameworks like NestJS or Express. Hands-on experience with Terraform, or infrastructure-as-code tooling. Experience with observability platforms like Datadog (metrics, tracing, alerting). Exposure to DynamoDB or other NoSQL databases at scale. Experience with distributed or event-driven architectures ...

Embedded Software Engineer

Location
Greater London, England, United Kingdom
vulnerability management and secure provisioning. Edge & Cloud Integration Integrate devices with cloud IoT platforms and backend services. Define and improve telemetry, health monitoring, and observability for deployed devices. Support reliable operation of connected devices in the field, including debugging fleet issues and improving resilience. Compliance & Testing Support testing and validation …/CD workflows for embedded software. Practical debugging experience using lab and software tools such as logic analysers, protocol analysers, network sniffers, or observability platforms. Strong problem-solving skills and the ability to work across hardware and software boundaries. Nice to Have Experience with low-power wireless technologies such ...

Backend Engineers (Ruby)

Location
Greater London, England, United Kingdom
technical insight into upcoming work and helping pull the team together to ship it Deliver your work using agile methodologies and tools like tests, observability, A/B tests, and feature flags Mentor colleagues to help them grow as engineers, and actively support their development Contribute to cross-cutting concerns … Rails, building data models, APIs, and business logic services in a production monolith Comfortable delivering with agile methodologies and practices like automated testing, observability, A/B testing, and feature flags A collaborative approach — you work well with product, design, and analytics partners, and enjoy shaping work beyond just ...

Principal Network Engineer

Location
Greater London, England, United Kingdom
across firewalls, NAT, VPN, security policies, and multi-tenant segmentation. Design highly available and scalable security architectures appropriate for mission-critical AI infrastructure. Reliability, Observability & Operations Lead complex technical escalations and root-cause analysis for network performance, reliability, and stability issues. Establish measurable SLOs and operational standards for network services. … technical direction for network observability, telemetry, monitoring, and alerting. Ensure clear visibility into fabric health, traffic patterns, performance, and capacity. Develop runbooks, automation, and engineering improvements that systematically reduce operational toil. Act as a senior 3rd/4th line escalation point for complex networking issues. Network Data & Configuration Management Ensure ...

GenAI Engineer

Location
Greater London, England, United Kingdom
indexing, retrieval policies, grounding, guardrails) and agent frameworks. Take basic infra ownership on GCP (or AWS/Azure): networking, autoscaling, CI/CD, IaC, observability, and cost tuning. Participate in on‐call for your area and drive root‐cause analysis with crisp follow‐ups. 15% Collaborate Pair with back … series analysis (forecasting, change‐point, drift). Cloud & ops: Basic infra ownership on GCP (or AWS/Azure): networking, autoscaling, CI/CD, IaC, observability, and cost control. Communication: You explain results clearly, align stakeholders, and write crisp docs. Bonus points DevOps wizardry; GPU/accelerator experience. Multimodal pipelines (text ...

Software Engineer, Agentic AI

Location
Cambridge, England, United Kingdom
product and platform capabilities for Roku TV. You will own the full lifecycle of agent development from prototyping and architecture through orchestration, evaluation, deployment, observability, and continuous improvement. You will contribute directly to Roku's AI strategy by engineering reusable components, optimizing agent workflows, and ensuring strong real-world performance … systems around them. Create reusable agent templates, modular components, and paved-path patterns that accelerate adoption across teams and use cases. Establish strong evaluation, observability, and monitoring for conversation quality, task success rate, latency, cost, and overall system performance. Build safeguards that improve production readiness and reliability, including testing pipelines ...

Data Platform DevOps Analyst

Location
Reading, England, United Kingdom
journey supporting the migration from Teradata to Databricks, validating production readiness, improving monitoring and automation capabilities and helping embed modern DataOps practices across deployment, observability and support processes. What You’ll Bring Here at Primark, we want everyone to feel valued – so please bring your authentic self to work … platforms and DataOps practices with exposure to Azure DevOps, Databricks, dbt Cloud, Azure data services, SQL, Python, ETL/ELT processes, job scheduling, monitoring, observability and deployment automation. Experience supporting production environments and service operations including monitoring, alerting, incident management, ticketing systems, release management, environment management, change governance and enterprise ...

Senior Cloud Infrastructure and Network Operations Solutions Architect

Location
Reading, England, United Kingdom
network OS internalsand host networkingand OS‐level security, together with an understandingof the storage and compute traffic patternsof HPC/AI clusters. Automation, GitOps & Observability: Proficiency in Pythonand Bash scripting, configuration management and Infrastructure‐as‐Code tools (e.g. Ansible, Terraform), GitOps‐based network configuration and firmware/upgrade management … large fleets, and observability stacks (Grafana, Loki, Prometheus) applied to fabric telemetry. Fleet Reliability& Customer Engagement: Demonstrated abilityto measure and improve MTBI and goodput on large GPU clusters—fault detection, drain and remediation workflows, SLO/error‐budget definition and post‐incident review–combining strong consultative background leading architectural reviews ...

Technical Support Engineer

Location
Paignton, England, United Kingdom
major escalations and transformation initiatives. Lead or contribute to the design and deployment of scalable support workflows, including automated log parsing, triage tooling, and observability pipelines. Identify trends in support data to drive systemic product or process improvements. Collaborate cross-functionally to evolve support readiness standards for new product releases … lead technical investigations, influence engineering teams, and manage high-pressure escalations independently. Experience with support and engineering tools such as Salesforce, Jira, Git, and observability stacks. Desirable Skills & Knowledge Knowledge of GNSS simulation systems, timing/sync architectures, or satellite navigation protocols. Familiarity with DevOps, CI/CD, or site ...

Senior Developer Experience Security Engineer

Location
Greater London, England, United Kingdom
experience across Motorway.. We have recently built and rolled out a new container platform on top of AWS Fargate, and are currently enhancing our observability, reliability, and developer-focused tooling. We will continue to build and evolve secure, standardised platform capabilities that reduce cognitive load and help teams ship faster … experience across Motorway.. We have recently built and rolled out a new container platform on top of AWS Fargate, and are currently enhancing our observability, reliability, and developer-focused tooling. We will continue to build and evolve secure, standardised platform capabilities that reduce cognitive load and help teams ship faster ...

Software Engineer III - Fullstack (Java, React, Python and AI) Engineer

Location
Greater London, England, United Kingdom
with product, marketing, and partners to translate requirements into well-designed technical solutions Improve engineering excellence through code reviews, test automation, CI/CD, observability, performance tuning, and operational best practices Ensure solutions meet security, privacy, and compliance expectations including consent management, data minimization, and access controls Required Qualifications, Capabilities … with campaign management and attribution systems Background in AI-enabled content generation, experimentation, or workflow automation Skills in test automation, CI/CD, and observability tools Experience optimizing performance and scalability Familiarity with consent management and data minimization practices Ability to drive innovation in personalization and measurement Experience working ...

Managed Service Operations - Head of Practice

Location
Greater London, England, United Kingdom
service outcomes. Key responsibilities You will lead the development and maturity of Made Tech’s operational capabilities incident, problem, and change management; monitoring and observability; automation and AIOps; governance; operational playbooks; runbooks; service health metrics; and 24/7/365 support patterns. You will ensure our teams have … ITIL practices blended with modern DevOps, SRE, Agile and platform‐engineering approaches. Broad technical awareness across cloud platforms, application architectures, data platforms, networks, observability tooling, security‐by‐design, and automation. Ability to create and evolve operational standards, playbooks, governance models, templates, and frameworks that drive consistency, stability, and efficiency. Skilled ...

Senior Data Engineer

Location
Greater London, England, United Kingdom
Insight teams to maintain reliable, accurate and meaningful data models as our products evolve Establish and promote good engineering practices across data pipelines, modelling, observability, performance and reliability Ensure data quality and integrity throughout our pipelines, identifying and addressing issues before they impact downstream users Provide technical guidance to engineers … data transformation skills, particularly using technologies such as Python, Spark and SQL Experience working with cloud-based data platforms Experience designing systems for reliability, observability, performance and data quality Strong understanding of data modelling and the principles behind well-structured analytical data Experience making technical design decisions and taking ownership ...

Software Engineer III - Fullstack (Java, React, Python and AI) Engineer

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
automationCollaborate with product, marketing, and partners to translate requirements into well-designed technical solutionsImprove engineering excellence through code reviews, test automation, CI/CD, observability, performance tuning, and operational best practicesEnsure solutions meet security, privacy, and compliance expectations including consent management, data minimization, and access controlsRequired Qualifications, Capabilities, and Skills … Skills: Experience with campaign management and attribution systemsBackground in AI-enabled content generation, experimentation, or workflow automationSkills in test automation, CI/CD, and observability toolsExperience optimizing performance and scalabilityFamiliarity with consent management and data minimization practicesAbility to drive innovation in personalization and measurementExperience working in cross-functional teamsJ.P. Morgan ...

Senior Data Engineer

Location
Maidenhead, England, United Kingdom
data solutions are reliable, scalable, performant, secure, and production‐ready Monitor, troubleshoot, and continuously improve pipeline performance, data quality, and platform stability Drive automation, observability, and supportability across data, analytics, and AI/ML solutions Our Ideal Candidate Strong data engineering experience with hands‐on delivery of scalable data pipelines … productionization, LLM‐based applications, or agentic AI patterns will be an added advantage Experience with DevOps and DataOps practices, including CI/CD, monitoring, observability, and incident support Maersk is committed to a diverse and inclusive workplace, and we embrace different styles of thinking. Maersk is an equal opportunities employer ...

Data Architect

Location
Greater London, England, United Kingdom
duties, isolation, and alignment to standards such as ISO 27001. Provide architectural direction for cloud data platforms, infrastructure as code, CI/CD, and observability, ensuring designs are cost‐aware and operable. Define the data lineage and controls that support responsible production AI — monitoring, confidence scoring, and human … approaches, and how each changes data architecture requirements, including context window economics and prompt caching. Familiarity with infrastructure as code, CI/CD, and observability for data platforms. Awareness of MCP and tool‐orchestration concepts and their implications for composable, data‐driven AI systems. Fluent use of AI‐assisted development ...

Client Service Delivery, Sr Manager

Location
Birmingham, England, United Kingdom
Service Delivery Management Own full lifecycle service delivery across infrastructure and cloud environments, ensuring alignment to SLAs, KPIs, scope, and cost. Leverage AIOps and observability tools (e.g.Dynatrace, Datadog, New Relic, Elastic) to proactivelymonitorservice health and performance. Utilisepredictive alerting and anomaly detection to prevent incidents andoptimisedelivery priorities. Coordinate across internal teams … infrastructure and cloud environments Strong understanding of IT Managed Services frameworks Hands-on experience with AIOps tools such as Dynatrace and ServiceNow Familiarity with observability tools (e.g.Datadog, New Relic, Elastic) Knowledge of event analytics tools such as Splunk IT Service Intelligence andMoogsoft Experience in stakeholder and client management Financial management ...