1,476 to 1,500 of 1,810 Permanent Observability Jobs

Site Reliability Engineer (SRE) - Cloud & Automation

Hiring Organisation
Spencer Rose Ltd
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP 60,000 - 70,000 Annual
implementation of SRE practices across the organisation, working closely with infrastructure teams to optimise deployment processes and embed automation and operational excellence. Enhance observability and reliability , defining and implementing SLAs, SLOs and SLIs to improve alerting, monitoring, and capacity planning. Identify and eliminate toil , developing frameworks to analyse recurring issues … beneficial). Experience supporting and building multi-environment, multi-region cloud platforms (AWS or GCP), using IaC and GitOps workflows. Hands-on experience with observability/APM tooling such as Grafana, Datadog or Dynatrace. Background working in regulated financial services or banking environments. Excellent troubleshooting, analytical and communication skills, able ...

Vice President, DevOps Production Services

Hiring Organisation
Jobleads-UK
Location
Manchester, England, United Kingdom
enterprise applications and ensure platform stability, resiliency, and availability. Monitor application health, system performance, batch jobs, interfaces, and alerts using enterprise monitoring and observability tools. Investigate, troubleshoot, and resolve production incidents within defined SLAs. Perform root cause analysis (RCA) for recurring issues and drive permanent fixes. Analyze production logs, identify … Cloud experience preferred. Knowledge of automation/scripting using Python, Shell, or PowerShell. Exposure to DevOps/SRE practices, CI/CD pipelines, and observability tooling. Strong communication skills with the ability to provide concise incident and executive status updates. #J-18808-Ljbffr ...

Full Stack Engineer (Contract) – Leeds

Hiring Organisation
Jobleads-UK
Location
Leeds, England, United Kingdom
secure, scalable and maintainable applications Create automated unit and integration tests Contribute to CI/CD pipelines and continuous delivery Implement logging, monitoring and observability best practices Support production issues and continuous improvement initiatives Participate in peer reviews and Agile ceremonies Produce and maintain technical documentation The Team … testing Strong troubleshooting and problem-solving skills Excellent communication and collaborative approach Nice to Have Cloud-native development experience Contract testing and UI automation Observability (logging, metrics and tracing) Financial Services experience JIRA or similar Agile tooling To Be Considered... Please either apply by clicking online or emailing your ...

Azure DevOps Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
modules aligned to business and security requirements Implement and enforce Azure security controls and policies Lead platform security initiatives across Azure environments Contribute to observability, alerting, and Site Reliability Engineering practices Collaborate with engineering teams to deliver resilient and scalable solutions Security & Platform Focus Areas Implement perimeter security using Azure … Terraform Kubernetes certification and experience with AKS Deep understanding of DevOps and platform engineering principles Strong knowledge of cloud security best practices Experience with observability, monitoring, and SRE concepts Why Apply Fully remote within the UK Work with a mission‐driven, highly respected health tech organisation Modern cloud environment with ...

Senior Backend Java Developer (Java/Spring Boot)

Hiring Organisation
HTC Global Services Inc
Location
Dearborn, Michigan, United States
Employment Type
Permanent
Salary
USD Annual
infrastructure and deployment processes to enhance resiliency and reliability. Support application security practices, including data protection through encryption and anonymization. Troubleshoot production issues using observability and debugging tools. Required Qualifications Bachelor's degree. 6+ years of overall IT experience. 4+ years of software development experience. 5+ years of experience with … backend services. Experience building or maintaining test frameworks and testing tools. Strong understanding of software quality practices and continuous improvement processes. Knowledge of observability, debugging, and production operations. Experience with Docker, CI/CD pipelines, and cloud computing concepts. Experience with databases such as PostgreSQL, MySQL, or MongoDB. ...

Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Watford, England, United Kingdom
high-demand events. What you’ll be doing Objectives of the role Maintain reliable production services across digital platforms Improve monitoring, alerting, and observability coverage Reduce operational toil through automation Support incident response and continuous improvement Contribute to performance and scaling of services Production operations Participate … Incident response & improvement Support incident triage and resolution Participate in post-incident reviews and implement remediation actions Maintain and improve runbooks and operational documentation Observability Implement and maintain monitoring using: + Splunk + CloudWatch + Grafana Improve: + Logging quality + Metrics coverage + Alerting accuracy Contribute to linking system ...

Senior Platform Engineer - Developer Experience

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
paved roads that other engineers use. You will work across the software development lifecycle, from creating a new service through to testing, deployment, observability and operating it in production. You will join an established Platform team and work alongside our existing Developer Experience Engineer. You will speak directly with engineers … reliability and usability of our CI/CD systems. Developing reusable platform capabilities that product engineers can consume through self-service. Helping engineers use observability effectively, with good defaults for logs, metrics, traces and service‐level indicators. Working directly with engineers to understand friction, test ideas and support adoption. Using ...

Network Automation Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Jinja2 Integrating network automation into CI/CD pipelines for reliable, repeatable deployments Creating APIs and self‐service tooling for engineering teams Implementing observability and telemetry solutions for performance and reliability Partnering with network, platform and security teams to deliver resilient, scalable systems Contributing to incident response and production reliability … Ansible, Terraform and Jinja2; also must have experience leveraging AI tools, such as Claude Code Familiarity with Docker and Kubernetes Exposure to monitoring, observability or telemetry in distributed systems Pragmatic problem solver who can operate in ambiguity and take ownership Comfortable working in collaborative, fast‐paced engineering teams Deep understanding ...

Lead Software Engineer - Java / Python - Equity Derivatives - Front Office Quant Developer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
test strategy, defect triage, evidence collection, and sign-offs Maintain strong release discipline, including regression assessment, rollback/fallback planning, post-deployment verification, and observability improvements Gather and synthesize data/telemetry to develop reporting and metrics that improve stability, quality, and delivery predictability Identify hidden failure patterns in production … Agile environment, managing multiple priorities/projects, and producing clear documentation Operational excellence mindset, including incident leadership, postmortems, and continuous improvement in observability and production reliability Knowledge of equity derivatives products and workflows, including options pricing concepts and understanding of how system issues can impact quoting, booking, and hedging Preferred ...

Vice President Software Engineering

Hiring Organisation
Jobleads-UK
Location
City of Edinburgh, Scotland, United Kingdom
where 80–90% of code is AI‐generated, with a roadmap to 95%+. Embed modern engineering excellence (CI/CD, trunk‐based development, observability, and automated testing). Partner cross‐functionally across Product, DevOps, Security, and Platform teams. Build a high‐performance culture grounded in accountability, innovation, and continuous … where software is shipped to production frequently or daily. Expertise in modern practices including CI/CD pipelines, trunk‐based development, automated testing strategies, observability and system reliability. Proven ability to use engineering metrics to drive performance and continuous improvement. Organisational Design & Methodologies Experience designing and evolving engineering organisations using ...

Software Engineering Manager - Tooling and Optimisations

Hiring Organisation
Jobleads-UK
Location
Windsor, England, United Kingdom
practice, reduce duplication, and support maintainable, secure and high-performing systems. Improve delivery capability through platform reliability and DevOps maturity Continuously strengthen deployment pipelines, observability, alerting, incident response, recovery procedures and operational readiness across Field Ops engineering teams. Manage stakeholders and maintain clear communication Build trusted relationships across product, operations … data modelling and data quality controls. Ability to produce both high‐level and detailed design specifications. Experience leading DevOps practices, including CI/CD, observability, monitoring and incident management. Demonstrated capability leading multi‐squad engineering delivery in a product‐led organisation. Mindset & Ways of Working Comfortable working in iterative, outcome ...

Backend Engineer - Platform - Stacks | UK | Remote

Hiring Organisation
Jobleads-UK
Location
United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand ...

Team Lead, Integrations

Hiring Organisation
Genesis10
Location
Irving, Texas, United States
Employment Type
Permanent
Salary
USD Annual
Code & DevOps: IaC with Terraform; able to review and approve infrastructure changes Azure DevOps pipeline design, branch policies, and mandatory-review gate configuration Observability & Monitoring: Application Monitoring using Application Insights Log Management and Analytics Platform Monitoring using Azure Monitor Query & Analysis using KQL (Kusto Query Language) Alerting, dashboards, and validating … solution observability Application Development: Minimum 8 years of solid .NET coding experience (C#, ASP.NET Core, background/worker services) Expert knowledge of common design and patterns including .NET patterns, libraries and Azure services Documentation & Communication: Strong technical writing skills Visual Modeling with Lucid chart/Azure architecture diagramming Comfortable delivering ...

Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
tenant environments for enterprise customers. Developer Velocity: Architect CI/CD pipelines (GitHub Actions, Docker) that allow our team to ship safely and quickly. Observability & Reliability: Instrument the stack with OpenTelemetry and Datadog to ensure we detect issues before our users do. AI Performance Tuning: Tune autoscaling and network routing … Haves: Experience with ECS, container orchestration, and distributed task queues (Celery/SQS). Strong Python skills. Familiarity with the Datadog/Grafana observability stack. A background in offensive security, CTFs, or AI/ML infrastructure. What We Offer Competitive Salary + Significant Equity: We want you to have true ...

Software Engineer / AI Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
within a defined problem, building and testing tool use, retrieval pipelines and agent workflows, integrating AI capabilities into enterprise systems, and contributing to evaluation, observability and guardrails. You will hold a high bar on code quality, flag risks and blockers early, and work alongside host‐function stakeholders to make sure … agentic AI solutions to production standard within a defined technical approach. Implement and test tool use, retrieval pipelines, and agent workflows. Contribute to evaluation, observability and guardrails for agentic systems. Integrate AI capabilities into existing enterprise workflows and systems. Maintain high code quality and documentation so patterns can be reused. ...

Lead Ai Engineer (AI/ML/R&D)

Hiring Organisation
Jobleads-UK
Location
Manchester, England, United Kingdom
powered applications. Extensive experience building and operating cloud-native applications on AWS, with strong knowledge of CI/CD, infrastructure as code, observability and modern engineering practices. Experience delivering high-performance, real-time systems that operate reliably at scale with demanding latency requirements. Demonstrated experience leading the successful transition … powered applications. Extensive experience building and operating cloud-native applications on AWS, with strong knowledge of CI/CD, infrastructure as code, observability and modern engineering practices. Experience delivering high-performance, real-time systems that operate reliably at scale with demanding latency requirements. Demonstrated experience leading the successful transition ...

Engineering Manager (Remote - UK)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
evolving AWS-native and cloud-based platforms, alongside legacy systems Set and uphold strong engineering standards across code quality, testing, CI/CD, observability and documentation Stay close to technical decisions through design reviews, architecture discussions and hands‐on coaching Balance new feature delivery with technical debt, reliability, security … least one modern programming language Strong grasp of system design and software engineering fundamentals Experience with Infrastructure as Code, CI/CD, observability and secure production systems Able to communicate technical ideas clearly to both technical and non‐technical stakeholders What Altus Group Offers Rewarding performance: competitive compensation, incentive ...

Senior Machine Learning Engineer - News

Hiring Organisation
Disney Entertainment and ESPN Product & Technology
Location
New York, United States
Employment Type
Permanent
Salary
USD Annual
sensitive outcomes, while proactively identifying, communicating, and mitigating risks to ensure successful execution Champion engineering best practices across code quality, testing, CI/CD, observability, and incident response Mentor and coach engineers, fostering a culture of ownership, collaboration, and continuous improvement Contribute to technical documentation and promote knowledge sharing across … Databricks, Kinesis, Kafka Proven leadership, coaching, and mentoring skills, with the ability to inspire and empower a team towards achieving business goals Experience with observability tools for metrics, logging, and monitoring such as Datadog Experience working in Agile/Scrum development environments Excellent communication skills and a commitment to collaboration ...

Senior Machine Learning Engineer - News

Hiring Organisation
Disney Entertainment and ESPN Product & Technology
Location
Glendale, California, United States
Employment Type
Permanent
Salary
USD Annual
sensitive outcomes, while proactively identifying, communicating, and mitigating risks to ensure successful execution Champion engineering best practices across code quality, testing, CI/CD, observability, and incident response Mentor and coach engineers, fostering a culture of ownership, collaboration, and continuous improvement Contribute to technical documentation and promote knowledge sharing across … Databricks, Kinesis, Kafka Proven leadership, coaching, and mentoring skills, with the ability to inspire and empower a team towards achieving business goals Experience with observability tools for metrics, logging, and monitoring such as Datadog Experience working in Agile/Scrum development environments Excellent communication skills and a commitment to collaboration ...

Software Engineer / AI Engineer

Hiring Organisation
Elsevier
Location
Greater London, United Kingdom
Employment Type
Full Time
within a defined problem, building and testing tool use, retrieval pipelines and agent workflows, integrating AI capabilities into enterprise systems, and contributing to evaluation, observability and guardrails. You will hold a high bar on code quality, flag risks and blockers early, and work alongside host-function stakeholders to make sure … agentic AI solutions to production standard within a defined technical approach. Implement and test tool use, retrieval pipelines, and agent workflows. Contribute to evaluation, observability and guardrails for agentic systems. Integrate AI capabilities into existing enterprise workflows and systems. Maintain high code quality and documentation so patterns can be reused. ...

Principal Software Engineer SC&L

Hiring Organisation
Jobleads-UK
Location
City of Westminster, England, United Kingdom
source technology Support recruitment, onboarding and internal and external brand outreach activities Tech Stack Java, Micronaut, PL/SQL ReactJS, Next.js Azure Cloud, Dynatrace (observability) Mule, Kafka, MQ Blue Yonder Dispatcher Essential Experience Significant track record of strategic and innovative thinking, as well as execution and implementation Specialist in clean … customers to a desired outcome, without prescribing it Authoritative skills at cloud computing (network, security, serverless, Kubernetes etc) and automation Experience with implementation of Observability and Reliability using market technologies (e.g.: Dynatrace) Advocate and experience of Continuous Integration and Continuous Delivery Advanced experience of DevOps: you build ...

Lead AI Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
solutions Enhance CI/CD pipelines for AI systems, introducing automation, traceability and controlled release processes to ensure safe and repeatable deployments. Define monitoring, observability and operational strategies, improving visibility across quality, performance, safety, cost and reliability to support effective issue diagnosis and resolution Set standards for prompt engineering, evaluation … using LLM APIs/Bedrock, including RAG architectures, vector databases and orchestration frameworks DevOps skills including CI/CD pipelines, containerisation and monitoring/observability, with experience defining release and operating practices for AI services Experience Leading Engineering teams: providing technical guidance, aligning on standards/patterns, and adapting plans ...

Software Engineer

Hiring Organisation
RWS
Location
Sheffield, United Kingdom
Employment Type
Full Time
security, reliability, and operation of the services you build, with a DevSecOps approach throughout Improving engineering practices including CI/CD pipelines, automated testing, observability, security scanning, and deployment workflows Collaborating closely with product managers, designers, domain experts, and other engineers to deliver meaningful outcomes Mentoring and supporting engineers across … DevSecOps mindset with experience owning the security and operation of services you build Understanding of modern delivery practices: CI/CD, automated testing, observability, and production ownership Ability to work across the full development lifecycle, from early design through to deployment and operations Clear communication skills and a collaborative working ...

Senior Software Engineer - Customer Engineering

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
designing and maintaining RESTful APIs, web hooks, and service-to-service integrations Comfortable operating in AWS, GCP, or Azure environments and working with modern observability tooling, logging, and monitoring platforms A track record of delivering performant, reliable and scalable applications Excellent collaboration and communication skills in cross-functional teams, including … build systems, but why architectural decisions matter You can balance scalability, reliability, maintainability, and speed of execution You think critically about security, observability, and operational excellence from day one Love the idea of blending software development, distributed systems and data-intensive applications Strong familiarity with authentication and identity technologies such ...

SRE Technical Lead

Hiring Organisation
Capgemini
Location
Surrey, United Kingdom
Employment Type
Full Time
point for major incidents and high risk releases, protecting service stability and ensuring blameless post incident reviews lead to measurable improvement. • Define and govern observability and capacity practices so reliability risks are visible, actionable, and proactively managed. • Ensure SRE practices align with service governance, security, and compliance requirements, and contribute … including: • Strong expertise in Kubernetes and OpenShift. • Experience with multi cloud and hybrid architectures, including service mesh (e.g. Istio). • Hands on experience with observability platforms such as Prometheus, Grafana, Loki, Tempo, and OpenTelemetry. • Strong Infrastructure as Code and GitOps experience (Helm, Kustomize, ArgoCD, Tekton). • Experience with CI/ ...