451 to 475 of 865 Remote Observability Jobs

Site Reliability Engineer, Big Data (Remote, International)

Hiring Organisation
PulsePoint
Location
United Kingdom, UK
Employment Type
Full-time
messaging layerHadoop and Ceph as distributed storage layerSQL Server backup and recoveryTerraform, Ansible, Puppet, ArgoCD for operational automationPrometheus, Grafana, Icinga and PagerDuty as observability layerBare-metal servers and hybrid cloud/on-prem infrastructureWhat matters: understand distributed systems, failure recovery, and operational patterns at scale. Apache Kafka and Ceph expertise ...

Lead Software Engineer

Location
United Kingdom
reliability, cost and maintainability. Turning ambiguous product problems into clear technical solutions and execution plans. Raising engineering quality through stronger system design, testing, security, observability and development practices. Leading technically critical initiatives and mentoring engineers across teams. What We're Looking For 7+ years of software engineering experience building ...

Staff Software Engineer

Location
United Kingdom
reliability, cost and maintainability. Turning ambiguous product problems into clear technical solutions and execution plans. Raising engineering quality through stronger system design, testing, security, observability and development practices. Leading technically critical initiatives and mentoring engineers across teams. What We're Looking For 7+ years of software engineering experience building ...

Principal Software Engineer

Location
United Kingdom
reliability, cost and maintainability. Turning ambiguous product problems into clear technical solutions and execution plans. Raising engineering quality through stronger system design, testing, security, observability and development practices. Leading technically critical initiatives and mentoring engineers across teams. What We're Looking For 7+ years of software engineering experience building ...

Lead Applications Support

Hiring Organisation
Experis
Location
London, United Kingdom
Employment Type
Permanent
Salary
£65000 - £70000/annum bonus + bens
resilience. Desirable Experience Experience supporting investment management, wealth management, trading, or financial platforms. ITIL certification or strong ITIL working knowledge. Experience with monitoring and observability tools. Scripting or automation experience. Knowledge of operational resilience, disaster recovery, and security controls. Experience contributing to platform upgrades, migrations, and technology transformation initiatives. About ...

Technical Product Manager

Hiring Organisation
Creditsafe
Location
Cardiff, Wales, United Kingdom
product capabilities are exposed through the API surface. API design fluency. REST principles, OpenAPI/Swagger, pagination, auth patterns, error handling API observability including usage analytics, rate limiting, deprecation telemetry Benefits Competitive Salary. Company Laptop supplied. Bonus Scheme. 25 Days Annual Leave (plus bank holidays). Hybrid working model. Healthcare ...

Senior Software Engineer C Linux Networking - Remote

Hiring Organisation
Saxon Recruitment Solutions
Location
Edinburgh, UK
Employment Type
Full-time
router leveraging a software dataplane. It provides a broad set of routing protocols and other network features, as well as cutting edge configuration and observability capabilities. What you'll needTechnical leadership for a significant development or project, with some demonstrable experience operating as a Subject Matter ExpertDemonstrable experience designing ...

Senior Commercial Counsel - EMEA (Remote in UK) (977-SLS)

Location
United Kingdom
client, a mission-driven, fast-growing cloud company and leader in real-time analytics, data warehousing, and data observability, has exclusively engaged Solutus Legal Search to assist its executives in hiring an attorney to serve as Senior Commercial Counsel. Reporting to the Associate General Counsel, this attorney will ...

Staff Software Developer (IAM)

Location
Greater London, England, United Kingdom
organizations. This shift brings challenges that cannot be solved at the level of individual tools alone, challenges that involve governance, security, cost control, observability, and coordination between humans and autonomous agents. The JetBrains Platform team (JCP) is building the foundation that connects developer workflows, team-level collaboration, and organizational control ...

Lead Engineer - United Kingdom

Location
Greater London, England, United Kingdom
Break ambiguous product and technical problems into clear engineering solutions and execution plans. Set a high bar for code quality, system design, testing, security, observability and production reliability. Review technical designs and code from other engineers and provide clear, actionable feedback. Mentor engineers and help strengthen technical capability across ...

Engineering Manager - AI Payments App

Location
Greater London, England, United Kingdom
Drive technical architecture and engineering decisions across APIs, databases, integrations, infrastructure and customer-facing systems. Set engineering standards around code quality, system design, testing, observability, security and production reliability. Stay close to the technical work through architecture reviews, code reviews, debugging and critical engineering decisions. Develop engineers through regular feedback ...

Staff Machine Learning Engineer - Ops

Location
Greater London, England, United Kingdom
Actions experience Strong communications skills with a collaborative mindset Desirable Experience with Pytorch, TensorRT, quantisation and model deployment Experience with Grafana monitoring and production observability This is a full-time role based in our office in London. At Wayve we want the best of all worlds so we operate ...

Engineering Manager - Customer London, UK

Location
Greater London, England, United Kingdom
team owns. Collaborate with tech leads, architects, and other EMs to share context, resolve dependencies, and improve system design. Encourage good practices around testing, observability, performance, and resilience – ensuring your team owns their systems in production. Hiring and Talent Development Take a leading role in hiring, onboarding, and growing diverse ...

Senior Software Engineer - UI Developer Tooling (Hybrid, London)

Location
Greater London, England, United Kingdom
push hooks so engineers catch issues in seconds, not in a failed pipeline twenty minutes later. Work with UI Quality, UI Release and Observability so the platform feels like one thing through a single CLI. Pair with Product Group leads to migrate legacy build and tooling patterns onto the paved ...

Lead AI Experience Engineer (Remote, United Kingdom)

Location
United Kingdom
Experience defining AI evaluation frameworks and using interaction data, testing, and analytics to improve performance. Knowledge of RAG, semantic search, enterprise knowledge systems, grounding, observability, and regression testing. Working knowledge of APIs, cloud services, software architecture, integration patterns, and modern engineering practices. Ability to lead complex initiatives, navigate ambiguity ...

Senior DevOps & Infrastructure Engineer

Location
United Kingdom
monitoring, release management, and platform operations. Help define practical, secure, and scalable ways of using AI in DevOps processes, whilemaintainingstrong engineering standards and governance. Observability, Reliability & Continuous Improvement Improve monitoring, alerting, logging, and observability practices to help teams detect issues earlier and resolve incidents faster. Analyze platform performance, deployment quality ...

Principal Technology Architect Hybrid Cloud Platforms -Germany, UK, Netherlands

Hiring Organisation
Infosys Technologies
Location
London, UK
Employment Type
Full-time
across cloud and on‐prem: IAM, PAM, KMS, certificates, secrets management. Partner with security architects to ensure platform designs meet enterprise security requirements.5. Resilience, Observability & Performance EngineeringArchitect high‐availability and disaster recovery models across hybrid environments. Define availability zones, failover strategies, multi‐region patterns, and RTO/RPO targets. Implement … full‐stack observability: logging, metrics, tracing, synthetic testing, SLO/SLI models. Ensure platform performance aligns with business workloads, including real-time and latency-sensitive applications.6. Governance, Risk & Compliance (GRC)Establish enterprise policies for segmentation, tagging, lifecycle management, patching, and resource standards. Govern platform consistency through enterprise architecture boards ...

Principal Java Engineer

Location
Wallingford, England, United Kingdom
continuous improvement. Production systems are reliable, observable and operationally excellent Lead root cause analysis and resolution of complex production issues. Drive improvements in system observability, monitoring and operational performance. Ensure applications are designed and operated to meet reliability, availability and performance targets. Partner with Operations, DevOps and QA teams … Claude, Codex, Gitlab Duo, etc) REST APIs, OpenAPI, Microservices, Event-driven architecture (RabbitMQ) Containers, Docker, AWS, Linux CI/CD with GitLab Pipelines & Jenkins Observability: logging, metrics and monitoring MySQL, Apache Solr Front-end UI (e.g. Angular) Person Specification Strategic and systems-thinking mindset Excellent communication and stakeholder management skills ...

Senior Reliability Engineer

Hiring Organisation
Fitch Ratings
Location
London, UK
Employment Type
Full-time
someone who is curious about the evolving role of AI in infrastructure engineering, someone who actively explores how AI-assisted tooling, automation, and intelligent observability can raise the bar for reliability and developer experience. You will collaborate closely with global development and engineering teams to deliver reliable, resilient, and high … reliability, security, and efficiencyIdentify, contain, and mitigate risk across all cloud environments, maintaining a robust security posture for infrastructure and applicationsImplement proactive monitoring and observability practices to detect and prevent issues before they impact usersDevelop and maintain automation and tooling solutions, including AI-assisted approaches to reduce toil and accelerate ...

AI Engineer – LLMs, NLP & Market Intelligence

Location
Greater London, England, United Kingdom
automated research workflows Retrieval, embeddings and semantic search Narrative detection and clustering Multilingual text intelligence Large-scale data and model pipelines Model evaluation, observability and monitoring Production ML infrastructure on AWS The problems are often open-ended. You might be evaluating how reliably different models identify changes in a market … models RAG and vector databases Apache Airflow AWS S3 ECS/EKS Redshift Docker Kubernetes GitHub Actions Pulumi or Terraform SQL Model monitoring and observability Distributed processing Financial markets, economics or commodities Financial-market experience is not required. Curiosity about how markets, economics and global events interact is more important. ...

Data Engineer, Vice President

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent, Work From Home
datasets, and extensible pipelines that support multiple Company Intelligence products and advanced analytics use cases. Ensure high standards of data quality, governance, lineage, and observability, proactively managing operational and compliance risks across enterprise-grade data products. Partner effectively with product, analytics, and business leaders, developing a deep understanding of strategic … control, testing, CI/CD). Experience building and operating data pipelines using workflow orchestration frameworks (e.g. Apache Airflow), with a focus on reliability, observability, dependency management, and operational resilience. Experience designing and operating cloud-native data platforms (AWS or Azure preferred) and enterprise data warehouses (Snowflake preferred), including performance ...

Data Engineer, Vice President

Hiring Organisation
Hackajob Ltd
Location
Slough, Berkshire, UK
Employment Type
Full-time
datasets, and extensible pipelines that support multiple Company Intelligence products and advanced analytics use cases. Ensure high standards of data quality, governance, lineage, and observability, proactively managing operational and compliance risks across enterprise-grade data products. Partner effectively with product, analytics, and business leaders, developing a deep understanding of strategic … control, testing, CI/CD). Experience building and operating data pipelines using workflow orchestration frameworks (e.g. Apache Airflow), with a focus on reliability, observability, dependency management, and operational resilience. Experience designing and operating cloud-native data platforms (AWS or Azure preferred) and enterprise data warehouses (Snowflake preferred), including performance ...

Production AI Engineer - Vice President

Location
Greater London, England, United Kingdom
engineering techniques to integrate large language models (LLMs) into operational tooling, incident response pipelines, and developer productivity platforms. Leads the development of AI-native observability solutions — leveraging intelligent agents to detect anomalies, predict failures, and automate remediation before issues impact end users. Writes clean, well-tested, and well-documented code … . Operational experience of using middleware technologies (MQ, Apache Kafka, etc.) to run services at scale is desirable. Strong experience with end-to-end observability stacks (Datadog, AppDynamics, Dynatrace, etc.) is desirable. Degree in Computer Science, Mathematics, Physics, or a related technical subject is desirable. Experience of senior stakeholder management. ...

Site Reliability Engineer- Spacetime UK

Location
Greater London, England, United Kingdom
system for a platform that transforms how networks of satellites, ground stations, and fleets are interconnected and orchestrated. You will be building the core observability stack that ensures the reliability of systems critical to the operation of satellite megaconstellations and missions to deep space. This is a greenfield/brownfield … expert, helping to define and implement the strategy and building the tools that empower our engineers. You will support the roadmap to mature our observability stack, moving from cloud-native tools to a robust, scalable, and insightful platform built on best-in-class technologies (Prometheus, OpenTelemetry, etc.). ...

Senior Backend Engineer — Commercial Planning (Hybrid)

Location
City of Westminster, England, United Kingdom
thousands of colleagues. You’ll work on Commercial Planning and Fashion, Home & Beauty transformation initiatives, collaborating with architecture, product managers and partners to raise observability, security and performance on cloud-native platforms. #J-18808-Ljbffr ...