3,301 to 3,325 of 4,288 Permanent Observability Jobs

Senior Network Engineer — Remote/Hybrid (Multi-Region DC)

Location
Greater London, England, United Kingdom
regions while modernizing automation with Ansible and Terraform. The role offers flexible remote/hybrid work with relocation support and a strong focus on observability using Prometheus and Grafana. Remote-friendly with international scope. #J-18808-Ljbffr ...

Lead AI Platform Engineer - Secure, Scalable Data Infra

Location
Greater London, England, United Kingdom
safe and scalable AI adoption. You will collaborate with AI/ML research teams, replace manual operations with self-service workflows, and implement robust observability and security by default. This role offers ownership of production systems in a regulated financial setting. #J-18808-Ljbffr ...

Technical Account Manager

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
team rewriting the rules of observability. A platform at the bleeding edge, empowering businesses to understand and act on their data through better observability, all in real time. Innovation removes the need for unnecessary indexing, cutting costs and complexity. Truly end to end, the platform encompasses everything; from logs … This is a high-impact role that demands serious technical firepower. You'll need deep experience with Cloud native tooling, hands-on knowledge of observability tools e.g. Grafana, DataDog or Splunk, and the ability to troubleshoot containerised environments like a pro. You have the technical knowledge and the confidence ...

Senior Software Engineer, AI-Powered Expense Platform

Location
Greater London, England, United Kingdom
interfaces with AI components, focusing on accuracy, latency, and scalable data flows. You will work across international teams, mentor others, and ensure robust observability, auditing, and compliance in financial contexts. Based in London or willing to relocate, you will join a fast-paced, tech-driven environment. #J-18808-Ljbffr ...

Applied AI Engineer New

Location
Greater London, England, United Kingdom
built quickly and as separate systems. The next challenge is to bring these approaches together: build reusable agentic infrastructure, establish a robust evaluation and observability layer, and create systems that allow us to automate new workflows quickly and reliably as Dwelly scales. This is not an AI research role. … loops. Move us from one-off AI solutions toward reusable infrastructure where new workflows can be introduced quickly and with predictable reliability. 2. Evaluation & observability Build the evaluation framework that allows us to understand how our agents perform and why they succeed or fail. Make testing, tracing, debugging, and evaluating ...

Director, Technical Account Management

Hiring Organisation
Datadog
Location
London, UK
Employment Type
Full-time
growing services revenue; you position TAM as a value driver, not a cost center. Technically fluent across infrastructure, cloud platforms (AWS, Azure, GCP), observability, and monitoring, credible with practitioners and C-level executives alike. A strategic operator who pairs vision with execution: you set direction, build the systems to support … region. Bonus Points: Experience redesigning organizational structures mid-growth: building new team shapes or specializations as a business scales. Track record evangelizing cloud adoption, observability, or security to C-level and board-level stakeholders across EMEA.Hands-on experience with Datadog or other leading cloud monitoring and observability platforms. Multilingual ...

Senior Connectivity Engineer / Network Engineer

Location
Wallingford, England, United Kingdom
WireGuard/Tailscale or equivalent): access‐as‐code, policy patterns, posture/health automation, and resilience/disaster recovery planning. Deliver fleet‐wide connectivity observability: monitoring, alerting, reporting, and actionable signals that help teams diagnose end‐to‐end issues quickly. Improve cellular/SIM lifecycle management: provisioning automation, usage/…/PMTUD, conntrack, nftables/iptables) and diagnosing kernel‐level networking behaviour. Proficient in Go and/or Python and experienced with modern observability tooling; bonus points for containers/IoT OS, ACL‐as‐code patterns, and carrier/router API integrations. #J-18808-Ljbffr ...

Principal GenAI Platform Architect (Full-Stack)

Location
York and North Yorkshire, England, United Kingdom
back-end work, designing scalable GenAI features and integrating MCP services. You will lead architecture and rollout of autonomous AI agents, ensure robust observability, and collaborate with product and design teams to turn vision into production-ready systems. The role demands hands-on experience with React/TypeScript ...

Senior/Staff Software Engineer (Nova Core)

Location
Greater London, England, United Kingdom
operations across Nova Cloud deployments. This role focuses on the Nova Core “inner loop”: service architecture, APIs, data models, persistence, authn/authz, observability, and developer experience that other Nova modules and product teams depend on. What you’ll do Own and ship critical Nova Core backend services (e.g., common … engineering and product teams. What success looks like Core Nova services are delivered, adopted, and operated reliably with clear SLIs/SLOs and runbooks. Observability is strong enough that incidents are detected quickly and resolved faster over time (improving MTTD/MTTR). API versioning and compatibility practices reduce integration ...

Sr Director, Platform Engineering – Data Platform & Agentic Platform

Location
Greater London, England, United Kingdom
operate agent workflow platform capabilities aligned to product‐defined standards and interfaces, including traceability, state handling, and convergence patterns Implement production‐grade evaluation, observability, auditability, and guardrail mechanisms required for safe AI workflows Implement security controls, access governance, encryption, and audit requirements in partnership with InfoSec while ensuring enterprise SDLC … large‐scale SaaS systems with production operations accountability Demonstrated success building and operating platforms adopted by multiple product teams, including reliability discipline (SLOs), observability, and incident management Strong hands‐on technical leadership background in distributed systems and platform engineering Deep experience with data platform engineering at scale, including ingestion ...

Senior/Staff Software Engineer

Location
Greater London, England, United Kingdom
tooling, optimizing for reliability, latency, cost, and debuggability in production. Build and maintain the surrounding infrastructure: data pipelines, evaluation harnesses, prompt and model management, observability, and safety/guardrails. Work across the stack—from backend integrations and APIs to simple UI hooks—to deliver complete AI features, not just model … workflows quickly, then harden what works. Own, downscope, ship, iterate: one clear owner per feature, from prototype to production. Fundamentals done well: evaluation, observability, and safety are part of the first version, not an afterthought. Competitive salary and meaningful equity. Health, dental, and vision coverage. Flexible time off and support ...

Operations and SRE Manager

Location
Greater London, England, United Kingdom
internal and external customers. You will be responsible for driving reliability improvements, advancing automation and AI-Ops capabilities, and leading a team focused on observability, incident response, operational excellence, and continuous improvement. Responsibilities Lead the transformation of the Operations function towards an AI-Ops operating model, driving the adoption … improvement actions are owned, tracked and completed. Strengthen operational process adherence, ensuring responsibilities are clear and delegation is effective. Drive SRE practices across observability, automation, disaster recovery, design for reliability, on-call readiness and production support. Protect service levels by ensuring engineering effort is balanced across InfoSec commitments, operational tickets ...

Front Office Equities Trading Technology Support

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
firm's systems to ensure operational stability and availabilityAssist in the monitoring of production environments for anomalies and address issues utilizing standard observability toolsIdentity issues for escalation and communication, and provide solutions to the business and technology stakeholdersAnalyze complex situations and trends to anticipate and solve incident, problem, and change … meaningful relationships to achieve common goalsDemonstrates knowledge of applications or infrastructure in a large-scale technology environment both on premises and public cloudExperience in observability and monitoring tools and techniquesExposure to processes in scope of the Information Technology Infrastructure Library (ITIL) framework Preferred qualifications, capabilities, and skillsExperience with ...

Front Office Equities Trading Technology Support

Location
Greater London, England, United Kingdom
firm’s systems to ensure operational stability and availability Assist in the monitoring of production environments for anomalies and address issues utilizing standard observability tools Identity issues for escalation and communication, and provide solutions to the business and technology stakeholders Analyze complex situations and trends to anticipate and solve incident … achieve common goals Demonstrates knowledge of applications or infrastructure in a large-scale technology environment both on premises and public cloud Experience in observability and monitoring tools and techniques Exposure to processes in scope of the Information Technology Infrastructure Library (ITIL) framework Preferred qualifications, capabilities, and skills Experience with ...

Software Engineer (Simulation, Evaluation, Validation)

Location
Greater London, England, United Kingdom
differences between on-road and simulated execution, identifying issues across data, inference and simulated components Improve simulation reproducibility, reliability and debuggability through automated testing, observability and better developer tooling Profile and improve simulator performance, helping us run increasingly large evaluation workloads efficiently Work with internal users and adjacent engineering teams … sensor data such as camera, radar, lidar or GNSS, including modelling uncertainty or noiseExperience integrating machine-learning inference into production systemsExperience with performance profiling, observability or debugging distributed systemsFamiliarity with large-scale batch processing, cloud infrastructure or GPU-based workloads #J-18808-Ljbffr ...

Professional Services Consultant

Location
Greater London, England, United Kingdom
ITRS, we make society's critical technology work. Our mission is to deliver automated and holistic IT observability solutions that safeguard critical applications and enable innovation. We are the only monitoring and observability platform designed for the most demanding and regulated industries — trusted by 90% of Tier 1 capital markets ...

Senior Firewall / Security Network Engineer

Location
Greater London, England, United Kingdom
traffic, TLS handshake failures, VPN outages, firewall failovers. Own the out-of-band (OOB) management network for firewalls: reachability, console access, and monitoring enablement. Observability, logging & compliance Integrate firewalls with logging and monitoring — FortiAnalyzer, SIEM log forwarding, netflow, SNMP. Help close misconfiguration and risk findings Required qualifications Hands‐on FortiGate … around network/security systems (e.g. API‐driven agents, MCP servers) Network automation — Ansible, AWX, or similar — for config and policy management. SIEM/observability tooling (FortiAnalyzer, netflow, SNMP, log pipelines). Multi‐site WAN and data‐center networking experience. Scripting (Python or similar) for tooling and validation. Exposure ...

Customer Solutions Engineer

Location
Greater London, England, United Kingdom
About ITRS At ITRS, we make society's critical technology work. Our mission is to deliver automated and holistic IT observability solutions that safeguard critical applications and enable innovation. We are the only monitoring and observability platform designed for the most demanding and regulated industries — trusted by 90% of Tier ...

AI Product Analyst

Hiring Organisation
The Portfolio Group
Location
London, United Kingdom
Employment Type
Permanent
Salary
£80000 - £85000/annum
fact. Day-to-Day Responsibilities Product Performance & Analytics: Measure how our AI products perform across retrieval quality, correctness of output, and user engagement. Observability Management: Own production quality observability, tracking metrics like thumbs-down rates, regeneration rates, task abandonment, and usage drift. Dashboard Engineering: Build and maintain the daily analysis ...

python developer market data

Location
Greater London, England, United Kingdom
data. Задачи Design and build data platform components using Python; Develop and optimise ETL pipelines, Data Lakes/Lakehouses and distributed systems; Improve infrastructure, observability and CI/CD practices; Maintain and evolve existing systems, including microservices, ETL and Excel add-ins; Support platform operations to ensure reliability and performance. ...

Site Reliability Engineer (DV Security Clearance)

Location
Manchester, England, United Kingdom
Engineer (SRE) to join a high-performing team supporting multiple data product and platform groups. This role is focused on improving the reliability, scalability, observability, deployment, and operational support of critical data-driven platforms and services operating within complex production environments. The successful candidate will work closely with engineering, platform ...

Senior Cloud Data Engineer - KSP

Location
Greater London, England, United Kingdom
optimise streaming data pipelines (Kafka/Flink or equivalent) to enable near real‐time data availability. Ensure data quality and reliability through validation frameworks, observability, and robust handling of late‐arriving or inconsistent data. Design data contracts and schemas that enable reliable integration between upstream event producers and downstream consumers. … medallion architecture or similar data layering approaches. Experience working with streaming technologies (Kafka, Flink, or similar). Strong understanding of data quality, testing, and observability practices. Experience designing schemas and handling data consistency challenges in distributed systems. Ability to work closely with stakeholders to translate business needs into scalable data ...

Data Architect

Hiring Organisation
Tiro Partners
Location
London, United Kingdom
Employment Type
Permanent
connect Design data models, pipelines, storage and governance frameworks Build reusable data capabilities, services and technical components Set standards across APIs, integration, deployment and observability Work closely with engineering, infrastructure, security and data teams Help ensure AI solutions are scalable, secure and maintainable Requirements 6+ years experience Data Architecture/ ...

Service Engineer

Hiring Organisation
Hackajob Ltd
Location
Knutsford, Cheshire, North West, United Kingdom
Employment Type
Permanent
line with ITIL processes, and working with engineering teams to improve application reliability, performance, and operational resilience. You will also leverage monitoring and observability tools to proactively identify issues and minimise service disruption. To be successful as an Application Support Engineer, you should have: Strong AWS knowledge from an application … such as GitLab pipelines, with an understanding of release and deployment practices Experience using monitoring and automation tools such as Kibana, AppDynamics and other observability platforms to support service performance and operational excellence You may be assessed on the key critical skills relevant for success in the role, such ...

Lead DevSecOps

Location
Greater London, England, United Kingdom
technical direction Genuine autonomy and influence from day one Build and mentor the DevSecOps function as the company scales Build security, compliance and observability into the development lifecycle What they're looking for Experience building internal developer platforms or sophisticated CI/CD ecosystems Deep expertise across Kubernetes and cloud ...