26 to 41 of 41 Observability Jobs in West London

Staff Software Engineer - AI

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency Establish engineering best practices through hands-on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement Partner with ...

Staff Software Engineer-AI

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … maintainability, system performance, reliability, security, scalability, and cost efficiency Establish engineering best practices through hands-on contribution, code reviews, technical design reviews, automated testing, observability, monitoring, and operational excellence Champion machine learning operations practices including model lifecycle management, prompt versioning, automated evaluation, deployment pipelines, monitoring, and continuous improvement Partner with ...

C# Developer

Hiring Organisation
CEI England Limited T/A CEI GB
Location
Isleworth, London, United Kingdom
Employment Type
Contract
Contract Rate
£380 per day
Oracle and Couchbase . Implement and support event-driven architectures using messaging platforms such as Kafka , ActiveMQ , or AmazonMQ . Monitor application health using observability tools, including logs, metrics, dashboards, and alerting. Build and maintain CI/CD pipelines to enable automated testing and deployments. Troubleshoot, debug, and resolve production … NoSQL databases, particularly Oracle and Couchbase . Experience with messaging and event-driven technologies such as Kafka , ActiveMQ , or AmazonMQ . Good understanding of observability and monitoring tools, including Prometheus , Grafana , and AWS CloudWatch . Experience with CI/CD pipelines , Concourse , and Git-based development workflows. Strong troubleshooting ...

Lead SRE - Chase UK

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … knowledge of microservice infrastructure components, including service discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes ...

Site Reliability Engineer — Cloud & Live Ops (Hybrid)

Location
Uxbridge, England, United Kingdom
resilient, secure, and highly available platforms underpinning live, broadcast‐adjacent services. Based at Stockley Park in Uxbridge with hybrid options, the role involves improving observability, incident response, automation, disaster recovery, and collaborating with engineering, operations and project stakeholders. #J-18808-Ljbffr ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
undergoing a multi year convergence and modernization journey. You will play a pivotal role in shaping our next generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives in fast paced front office environments, enjoys … Deep knowledge of reliability engineering principles: SLIs/SLOs, real-time telemetry, disaster recovery planning, capacity planning, and performance tuning. Experience designing and implementing observability frameworks for mission critical systems. Proven ability to lead incident response and drive long term remediation. Solid programming skills in Python, Java, or Kotlin, with ...

Senior Associate, Full-Stack Engineer

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
ways: Design, build, and maintain backend services, batches and APIs, contributing to UI components as needed. Own end-to-end delivery: implementation, testing, deployment, observability, and reliability. Write clean, well-tested code; participate in code reviews and continuous improvement. Collaborate with product, design, and operations to translate business needs into … microservices Proficiency in Java with Spring. Experience with CI/CD, automated testing (JUnit/Spock), and containers (Docker). Familiarity with microservices, observability/telemetry (e.g., Splunk, AppDynamics), and cloud deployments. Curiosity to understand the business domain and translate product strategy into technical solutions. How we work: Agile (Scrum ...

Advanced Solutions Architect

Location
Chiswick, England, United Kingdom
concept development when needed to de-risk architectural decisions## ## **Engineering Quality & Standards*** Define and champion non-functional standards: performance baselines, resilience patterns, observability requirements, and security controls* Lead or contribute to performance and load testing design, interpreting results in the context of SLAs and platform growth targets* Establish … delivery leads, product managers, and executive stakeholders* Drive alignment between iGaming and wider L&W engineering teams on shared concerns: API strategy, data platforms, observability, and shared services* Participate in hiring and technical assessment processes for engineering and architecture candidates**Qualifications**Degree in Computer Science, Software Engineering, or a related ...

Senior Software Engineer-AI

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
operation with limited supervision Hands-on experience with cloud-native technologies, serverless applications, event-driven architectures, data pipelines, relational and NoSQL databases, vector databases, observability tooling, and automated deployment pipelines Solid understanding of algorithms, data structures, scalability, reliability, performance optimization, security best practices, and engineering trade-offs Experience mentoring engineers … technical designs, participate in design reviews, and identify risks, constraints, trade-offs, and alternative approaches Maintain engineering excellence through automated testing, code reviews, observability, monitoring, alerting, operational readiness, and participation in on-call support Apply machine learning operations practices, including prompt versioning, automated evaluation, deployment pipelines, monitoring, and production issue ...

Senior Associate, Full-Stack Engineer Opportunities

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
ways: Design, build, and maintain backend services, batches and APIs, contributing to UI components as needed. Own end-to-end delivery: implementation, testing, deployment, observability, and reliability. Write clean, well-tested code; participate in code reviews and continuous improvement. Collaborate with product, design, and operations to translate business needs into … microservices Proficiency in Java with Spring. Experience with CI/CD, automated testing (JUnit/Spock), and containers (Docker). Familiarity with microservices, observability/telemetry (e.g., Splunk, AppDynamics), and cloud deployments. Curiosity to understand the business domain and translate product strategy into technical solutions. Working knowledge of Groovy with ...

Principal Platform Engineer (12 Month FTC)

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...

Principal Software Engineer-AI

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
production or in platforming (LLM or MCP gateway, agentic runtime, auth, data retrieval, eval tooling) Experience running AI systems in production at scale, including observability, cost and capacity planning, regression detection, and incident response for AI-powered applications Experience operating production distributed systems on AWS/Azure, with a strong … grasp of reliability, observability, and incident response at scale Deep knowledge of cloud-native technologies, serverless applications, event-driven architectures, data and inference pipelines, relational, NoSQL, and vector databases, and modern software architecture patterns Proven track record of owning multi-year technical strategy and architectural roadmaps, guiding teams from ...

Product Associate - SRE Team - Chase UK

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with engineering teams to ensure … services are designed, delivered, and operated with reliability in mind. Job responsibilities Support the product strategy and delivery of reliability capabilities, including standards, observability, incident practices, automation, and developer experience improvements. Partner with engineers, site reliability engineers, and cross-functional teams to understand problems, gather requirements, and translate ideas into ...

Vice President, Production Services Application Support

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
risk while improving platform stability. Drive initiatives through to completion with strong ownership, accountability, urgency, and quality. Champion an automation-first mindset, leveraging AI, observability, and tooling to reduce manual effort and improve service quality. Identify systemic issues and drive sustainable remediation through process simplification, platform improvements, and close partnership … incident, problem, and change management with measurable improvements in stability and service recovery. Demonstrated automation-first and AI-enabled mindset, with experience driving tooling, observability, and process automation. Strong ownership mentality and execution focus, with the ability to take initiatives from concept through delivery and embed sustainable outcomes. Deep technical ...

Lead Software Engineer - Proxy/SSE Network Security

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
resilience outcomes. Drive operational excellence at scale for perimeter, proxy, and SSE services in the US, including incident, change, and problem management rigor, observability and resiliency validation practices, automation to improve repeatability and evidence quality, reduction of client and partner impact, and execution of Technology Lifecycle Management (TLM) and modernization … design, exception frameworks, audit-ready traceability, and measurable risk reduction reporting. Experience with large-scale operations for externally facing or security enforcement services, including observability strategy, resilience testing, incident response alignment, and reduction of repeat incidents and client-impacting events. Experience designing and operating hybrid edge architectures and cloud interconnect ...

Front Office Equities Trading Technology Support

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
firm's systems to ensure operational stability and availability Assist in the monitoring of production environments for anomalies and address issues utilizing standard observability tools Identity issues for escalation and communication, and provide solutions to the business and technology stakeholders Analyze complex situations and trends to anticipate and solve incident … achieve common goals Demonstrates knowledge of applications or infrastructure in a large-scale technology environment both on premises and public cloud Experience in observability and monitoring tools and techniques Exposure to processes in scope of the Information Technology Infrastructure Library (ITIL) framework Preferred qualifications, capabilities, and skills Experience with ...