2,601 to 2,625 of 2,830 Remote/Hybrid Observability Jobs

Sr. Software Engineer

Hiring Organisation
Bixal
Location
Salt Lake City, Utah, United States
Employment Type
Permanent
Salary
USD Annual
developer-facing tooling for a federal agency's platform. Design and automate cross-platform workflows spanning source control, CI/CD, issue tracking, observability, and collaboration tools to reduce manual developer toil. Build reusable automation patterns (APIs, bots, webhooks, event-driven services, IaC, internal tooling) that standardize engineering workflows. Integrate … combination of education and experience. Demonstrated hands-on experience designing and automating cross-platform workflows spanning source control, CI/CD pipelines, issue tracking, observability, and collaboration tools. Proven ability to build reusable automation patterns (APIs, bots, webhooks, event-driven services, IaC, internal tooling). Experience integrating identity, access management ...

Sr. Software Engineer

Hiring Organisation
Bixal
Location
Sioux Falls, South Dakota, United States
Employment Type
Permanent
Salary
USD Annual
developer-facing tooling for a federal agency's platform. Design and automate cross-platform workflows spanning source control, CI/CD, issue tracking, observability, and collaboration tools to reduce manual developer toil. Build reusable automation patterns (APIs, bots, webhooks, event-driven services, IaC, internal tooling) that standardize engineering workflows. Integrate … combination of education and experience. Demonstrated hands-on experience designing and automating cross-platform workflows spanning source control, CI/CD pipelines, issue tracking, observability, and collaboration tools. Proven ability to build reusable automation patterns (APIs, bots, webhooks, event-driven services, IaC, internal tooling). Experience integrating identity, access management ...

Sr. Software Engineer

Hiring Organisation
Bixal
Location
Rapid City, South Dakota, United States
Employment Type
Permanent
Salary
USD Annual
developer-facing tooling for a federal agency's platform. Design and automate cross-platform workflows spanning source control, CI/CD, issue tracking, observability, and collaboration tools to reduce manual developer toil. Build reusable automation patterns (APIs, bots, webhooks, event-driven services, IaC, internal tooling) that standardize engineering workflows. Integrate … combination of education and experience. Demonstrated hands-on experience designing and automating cross-platform workflows spanning source control, CI/CD pipelines, issue tracking, observability, and collaboration tools. Proven ability to build reusable automation patterns (APIs, bots, webhooks, event-driven services, IaC, internal tooling). Experience integrating identity, access management ...

Senior Software Engineer, Backend

Hiring Organisation
Cognitiv
Location
Seattle, Washington, United States
Employment Type
Permanent
Salary
USD Annual
accounting for scalability, performance, failure modes, and security as the platform grows across the ad-tech ecosystem. Elevate Reliability: Identify bottlenecks and improve system observability (metrics, logging, monitoring). You will lead blameless post-mortems and implement long-term systemic fixes to prevent incident recurrence. Drive Engineering Excellence: Strengthen … strongly typed backend language such as Java or Kotlin, with experience building APIs, services, or data pipelines. Operational Mindset: Deep understanding of system reliability, observability, and "up-front" design for maintainability. Curiosity: You're naturally curious and eager to understand how things work. You ask thoughtful questions, dig into unfamiliar ...

Senior Software Engineer, CLV Monitoring

Hiring Organisation
SimpliSafe
Location
Quincy, Massachusetts, United States
Employment Type
Permanent
Salary
USD Annual
maintain testing strategies to ensure services stay performant and maintainable as they scale. Own services end-to-end, from design through production, including observability and resolving performance bottlenecks and technical debt. Collaborate with product, design, and other engineering teams to turn ambiguous problems into well-scoped technical solutions. … language (TypeScript, C#, etc.) with willingness to learn new tools and languages. Proficiency in a modern, cloud-scale NoSQL database. Proficiency in implementing modern observability frameworks and logging, and using them to diagnose real issues in production systems. Fluency with modern AI tools (e.g., Claude Code, Cursor) across research, design ...

Software Engineer II

Location
Greater London, England, United Kingdom
right production partner and fulfilled without a hitch. As a Software Engineer, you'll work closely with a small team to improve the resilience, observability, and scalability of our systems. From refining service architecture to setting up robust monitoring and alerting, you’ll play a hands‐on role in keeping … stakeholders across teams to support new product launches and vendor integrations Continuously improve how we work — from refining CI/CD pipelines to strengthening observability and developer experience About You Java (21) AWS suite Spring boot You care about great user experience, improving internal tooling, and want to collaborate with ...

Senior Applied AI Engineer (Private Sector Contractor)

Location
United Kingdom
data ingestion through to inference, owning the whole path rather than a slice of it. Make confidence earned, not asserted. You build the evaluation, observability and guardrails that show how a system actually behaves, its agent behaviour, model performance and failure modes. Set the technical bar. … reasoning Experience with edge or offline AI deployments Familiarity with Kubernetes (EKS/OpenShift) for managing deployed applications MLOps experience: model evaluation, monitoring, reproducibility Observability tooling for agentic systems (model drift, agent behaviour, performance monitoring) Experience with agent orchestration patterns and inter‐agent communication protocols (e.g. A2A) Familiarity with ...

Senior AI Engineer

Location
Harwell, England, United Kingdom
data ingestion through to inference, owning the whole path rather than a slice of it. Make confidence earned, not asserted. You build the evaluation, observability and guardrails that show how a system actually behaves, its agent behaviour, model performance and failure modes. Set the technical bar. … reasoning Experience with edge or offline AI deployments Familiarity with Kubernetes (EKS/OpenShift) for managing deployed applications MLOps experience: model evaluation, monitoring, reproducibility Observability tooling for agentic systems (model drift, agent behaviour, performance monitoring) Experience with agent orchestration patterns and inter‐agent communication protocols (e.g. A2A) Familiarity with ...

Staff Software Engineer, ML Infrastructure

Hiring Organisation
SimpliSafe
Location
Cambridge, Massachusetts, United States
Employment Type
Permanent
Salary
USD Annual
highest-stakes ML systems at SimpliSafe. Identify and remove the systemic bottlenecks in our ML deployment infrastructure - whether that's serving reliability, deployment friction, observability gaps, scaling, or cost. Build and operate real-time CV inference at scale Own the design and evolution of cloud-side inference systems that process … durable. Own reliability and operational excellence Lead incident response and postmortems for critical ML systems; turn lessons learned into platform-level improvements. Define SLOs, observability standards, and on-call practices for ML services in production. Qualifications 8+ years of software engineering experience, with a clear track record of building ...

Cloud SRE Lead: Scale Reliability & Observability

Location
Manchester, England, United Kingdom
Lloyds Banking Group is seeking a Lead Site Reliability Engineer to strengthen reliability, observability and operational excellence across Azure and Google Cloud Platform. You will lead a team of SREs, set engineering standards and drive improvements across cloud environments, with collaboration across product and platform teams. The role involves shaping ...

Senior Salesforce Engineer

Location
Greater London, England, United Kingdom
Some key info for you about Liberis: We were founded in 2007 We have provided over $3bn of funding to small businesses so far We have been named in CNBC & Statista Top 150 UK Fintechs ...

Lead Software Engineer

Hiring Organisation
Oscar Associates (UK) Limited
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£65,000
Lead Software Engineer Location: Manchester City Centre (Hybrid) Salary: £55k - £65k + Benefits The Role We're partnering with an established digital technology business looking for a Lead Software Engineer to join their growing engineering ...

Platform Engineer - Hybrid CI/CD & Observability

Location
United Kingdom
environments, enabling software and AI workloads for car development and race operations. You will design and implement infrastructure as code, CI/CD, and observability strategies, collaborating with software and AI engineers to improve developer experience and reliability. Hybrid working options are available. #J-18808-Ljbffr ...

Software Engineer, Observability

Location
Greater London, England, United Kingdom
shaping our story, you’ll help define what comes next. About the Role: We are looking for a Software Engineer to join our Observability team. Vercel users rely on Observability to monitor and understand their applications’ health and behavior. In this role, you will design, implement, and maintain Observability products … large-scale data ingestion, storage, and processing from distributed systems. Develop cutting-edge visualization tools to provide insights into application behavior and performance. Integrate observability features with popular frontend tools, frameworks, and build systems to enhance developer experience. Write clean, efficient, and well-documented code, ensuring platform reliability through thorough ...

Sr Director, Enterprise AI

Location
Greater London, England, United Kingdom
SaaS vendors, platform integrations such as ServiceNow, Salesforce, and SAP, and custom-build trade-offs; bringing well-reasoned recommendations to executive leadership.Use Dynatrace’s observability platform as a strategic advantage to monitor, measure, and continuously improve the reliability, performance, and business impact of internally deployed AI solutions.AI Governance and Responsible …/buy/partner decisions at the portfolio level, including vendor evaluation, contract negotiation support, and post-implementation performance tracking.Bonus: hands-on experience with observability or monitoring platforms — Dynatrace or similar — applied to AI system performance and reliability.Why you will love being a DynatracerDynatrace is a leader in unified observability ...

Sr. Azure Infrastructure Engineer (Ref: 18543)

Hiring Organisation
Professional Technology Integration, Inc
Location
Mechanicsville, Virginia, United States
Employment Type
Any
Salary
USD Annual
governance of Azure resources, networking, and security policies. Operations & Monitoring: Implement, maintain, and monitor mission-critical cloud resources in a large enterprise environment. Observability: Build meaningful observability solutions using Azure Monitor, Workbooks, Alerts, and related tooling. Migration & Deployment: Plan, lead, and execute cloud migration initiatives, including moving workloads from … network protocols, and network-level troubleshooting. Proven track record of executing on-premises to cloud migration projects in an enterprise setting. Demonstrated experience with observability tools (Azure Monitor, Alerts, Workbooks) and vulnerability remediation tools (e.g., Tenable). Strong sense of ownership with the ability to manage complex technical challenges from ...

Principal Platform Engineer

Location
Wales, United Kingdom
working on Mondays and Fridays Take technical ownership of the platforms that support Centerprise services and customer operations. You will lead improvements in reliability, observability and automation while remaining closely involved in complex engineering, major incidents and service recovery. Role Summary As Principal Platform Engineer, you will … hours escalation when required. Identify and address technical debt, operational risk and platform weaknesses. Ensure services remain supportable, recoverable and operationally efficient. Observability and service health Own monitoring and observability tooling, standards and operational dashboards. Develop service health metrics that provide clear and useful operational insight. Improve the quality ...

Staff Backend Engineer - Grafana Second Horizon | UK | Remote

Location
United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … remote opportunity, and we would be interested in applicants located in Spain, Sweden, UK, Ireland or Germany. The Opportunity: At Grafana Labs, we build observability tools that help users understand, respond to, and improve their systems – regardless of scale, complexity, or tech stack. We recently started a skunkworks initiative with ...

Cloud Platform Architect (Private Cloud)

Hiring Organisation
Ixceed Solutions Ltd
Location
Handsworth, West Midlands, UK
Employment Type
Full-time
environments Translate business and technology requirements into actionable platform engineering roadmaps Own Kubernetes architecture decisions including cluster design multitenancy ingress service mesh secrets management observability upgrade strategies and workload onboarding Establish GitOps CICD automation and configuration management standards Provide technical leadership for Kafka and EventDriven Architecture adoption Define Kafka governance … following VMware or Virtualization Technologies Kubernetes and Container Platforms KafkaEvent Streaming Platforms Strong Kubernetes experience including cluster operations networking upgrades multitenancy policy management observability and workload onboarding Strong Kafka architecture operations governance security and capacity planning experience Deep understanding of Distributed Systems Networking Fault Tolerance Performance Tuning and Capacity Planning ...

Cloud Platform Architect (Private Cloud)

Hiring Organisation
Ixceed Solutions Ltd
Location
Sheffield, South Yorkshire, Yorkshire, United Kingdom
Employment Type
Contract
environments Translate business and technology requirements into actionable platform engineering roadmaps Own Kubernetes architecture decisions including cluster design multitenancy ingress service mesh secrets management observability upgrade strategies and workload onboarding Establish GitOps CICD automation and configuration management standards Provide technical leadership for Kafka and EventDriven Architecture adoption Define Kafka governance … following VMware or Virtualization Technologies Kubernetes and Container Platforms KafkaEvent Streaming Platforms Strong Kubernetes experience including cluster operations networking upgrades multitenancy policy management observability and workload onboarding Strong Kafka architecture operations governance security and capacity planning experience Deep understanding of Distributed Systems Networking Fault Tolerance Performance Tuning and Capacity Planning ...

CommercePlatform Technical Lead

Hiring Organisation
Stott & May Professional Search Limited
Location
Hampshire, South East, United Kingdom
Employment Type
Contract
cause analysis and permanent fixes. - Review code, architecture and engineering practices. - Troubleshoot issues across APIs, integrations, frontend applications and Azure services. - Improve reliability, automation, observability, testing and deployment processes. - Support knowledge transfer and establish strong technical ownership of the platform. Essential Experience - Strong commercial experience with Commercetools and composable commerce … across: - Commercetools - TypeScript/Node.js - React/Next.js - Azure Functions, App Services, Cosmos DB - REST APIs and integrations - CI/CD pipelines - Automated testing - Observability and monitoring tools - Cloud-native and infrastructure automation practices Ideal Candidate A technically credible leader who can move between strategy and hands-on engineering, confidently ...

Cloud SRE

Location
Greater London, England, United Kingdom
databases. Day to day, you will monitor service health, resolve incidents, execute controlled changes, maintain patching and hardening standards, and improve reliability through automation, observability, and clear operational documentation. This is a 6-month contract, to work remotely (outside IR35) Key Responsibilities Infrastructure discovery and assessment - Build and maintain … AIOps and self-healing to detect and resolve issues earlier. Cloud administration - Administer compute, storage, and OS-layer services in Azure; support application infrastructure. Observability & monitoring - Implement and tune monitoring, alerting, and dashboards; improve signal quality and reduce noise. Security & compliance - Apply hardening, patching, and access controls in line with ...

Lead AI Engineer

Hiring Organisation
Capco
Location
London, UK
Employment Type
Full-time
experience deploying LLMs and multi-modal models at scaleStrong engineering background in Python with proven backend and API development skillsSolid understanding of scalable MLOps, observability, and cloud-native AI deploymentExcellent communication, problem-solving, and project management skills in agile environmentsBonus Points ForExperience with agentic frameworks (e.g., LangChain, LlamaIndex)Experience … deep learning frameworks and front-end developmentFamiliarity with Langfuse, Langsmith, or other LLM observability toolsUnderstanding of Model Context Protocol and bias/hallucination mitigation techniquesPrevious success in integrating GenAI solutions into enterprise-scale systemsWhy Join CapcoDeliver high-impact technology solutions for Tier 1 financial institutionsWork in a collaborative, flat ...

Principal Engineer, CSRE Provisioning (Remote, United Kingdom)

Location
Greater London, England, United Kingdom
Engineering, operating, stabilising, standardising, upgrading, migrating and, where appropriate, retiring them. The estate grew through years of pragmatic local decisions: inconsistent deployment patterns, patchy observability, ageing infrastructure. Our job is to reverse that. We are not measured by how much infrastructure we build - we are measured by how much complexity … technical position in cross-team architecture and design reviews with Platform Engineering, application teams, and senior leadership. Build and evolve the reliability, DR and observability practices used across the estate, going deeper than any single platform to ensure they generalize. Lead service acceptance for the most complex or highest-risk ...

Site Reliability Engineer

Location
Leeds, England, United Kingdom
exceptional customer experiences through reliable, scalable and high-performing technology. Operating at the heart of our betting and gaming platforms, our SRE team combines observability, automation and engineering excellence to ensure our systems perform when it matters most. Working across operational, proactive and engineering initiatives, you'll collaborate with development … coordinate responses and restore services efficiently. Conduct detailed postmortem investigations, identifying root causes, challenging assumptions and driving meaningful improvements to prevent recurrence. Utilise observability tools such as Splunk, New Relic and CloudWatch to monitor platform health, investigate issues and uncover performance insights. Design, execute and analyse performance testing activities ...