226 to 237 of 237 Observability Jobs in the South East

Cloud FinOps Analyst

Location
Tonbridge, England, United Kingdom
across Azure and Snowflake environments. A key focus of the role is leading the FinOps optimisation activities, embedding governance frameworks, and overseeing AKS cost observability using tooling such as Power BI, Kubecost etc. The FinOps Analyst partners closely with Engineering, Data, Cloud Operations, and Finance teams to enable a cost … optimisation, and waste elimination. Develop, maintain, and enforce cloud and data platform cost governance frameworks including tagging, budgeting, guardrails, and accountability processes. Oversee cost observability tooling (Kubecost, Snowflake dashboards, cloud cost portals) to ensure visibility of usage, forecasts, and budget performance. Manage budgeting, forecasting, cost allocation, and financial reporting ...

Senior Connectivity Engineer / Network Engineer

Location
Wallingford, England, United Kingdom
WireGuard/Tailscale or equivalent): access‐as‐code, policy patterns, posture/health automation, and resilience/disaster recovery planning. Deliver fleet‐wide connectivity observability: monitoring, alerting, reporting, and actionable signals that help teams diagnose end‐to‐end issues quickly. Improve cellular/SIM lifecycle management: provisioning automation, usage/…/PMTUD, conntrack, nftables/iptables) and diagnosing kernel‐level networking behaviour. Proficient in Go and/or Python and experienced with modern observability tooling; bonus points for containers/IoT OS, ACL‐as‐code patterns, and carrier/router API integrations. #J-18808-Ljbffr ...

Staff Site Reliability Engineer

Hiring Organisation
Genomics
Location
Oxford, Oxfordshire, UK
Employment Type
Full-time
multi-trillion-row, petabyte scale — sharding and replication, materialised views, merge and query optimisation, tenant isolation and cost/performance trade-offs. Owning SLOs, observability, capacity planning and incident response for data-intensive systems and pipelines, alongside the orchestration and job execution that power them. Shaping Data-as-a-Service … level rather than as a black box — and ideally have contributed code upstream. Reliability engineering for data platforms. You bring true SRE discipline — SLOs, observability, capacity planning and incident response — to analytical data systems and pipelines. Data-as-a-Service productisation. You think in terms of data as a product ...

Strategic Public Sector Account Exec (Observability & AI)

Location
Maidenhead, England, United Kingdom
Account Executive to lead a portfolio of 2-3 UK Public Sector customers and a select set of prospects. You will drive adoption across observability, security and AI-powered platform capabilities while managing executive-level relationships. You will orchestrate cross-functional teams, navigate complex procurement environments and partner with large ...

Performance and Monitoring Engineer

Hiring Organisation
Solus Accident Repair Centres
Location
Stansted, Essex, South East, United Kingdom
Employment Type
Permanent
Salary
£50,000
talented Performance and Monitoring Engineer to help us strengthen the stability, reliability and performance of our systems. If you're passionate about monitoring, observability and using data to proactively improve service health, this is a great opportunity to make a real impact across a large, modern technology estate. Responsibilities … improve speed, accuracy and consistency Supporting major changes, deployments and post-incident reviews with data-driven evidence Qualifications Strong experience with monitoring and observability tools (LogicMonitor, Azure Monitor, App Insights, Log Analytics, Defender for Cloud) Excellent understanding of cloud performance, IaaS/PaaS, networking fundamentals, API performance and capacity modelling ...

Clickhouse Solutions Architect

Location
Slough, England, United Kingdom
Role We are looking for a ClickHouse Solutions Architect to join our team supporting the design and implementation of a greenfield, enterprise-scale ClickHouse observability platform for a global banking client. This is a genuine greenfield build at significant scale — there is no incumbent platform to inherit or work around. … where benchmarks disprove the design Establish infrastructure-as-code, CI/CD and environment promotion for schema and configuration changes Productionisation Define and implement observability — system table monitoring, metrics, alerting thresholds, capacity headroom tracking Establish backup, restore and disaster recovery, and validate them by test Implement security and governance — RBAC ...

Platform Engineer

Hiring Organisation
Oscar Technology
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£580.00 - £610.00 per day
Platform/Observability Engineer | £610p/day (Outside IR35) | Remote | 6 months (Initially) | Grafana Cloud/Grafana Dashboard Our client is looking for an experienced Platform or Observability Engineer to focus on building and maturing their observability stack, with particular emphasis on Grafana Cloud, Alloy agent management, and dashboarding. … Working arrangement: Fully remote Key Responsibilities Set up and configure a Grafana Cloud test environment, ensuring it is production-representative and suitable for validating observability changes before rollout. Design, configure, and manage Alloy agent configurations, including planning and executing agent distribution across environments/fleets. Create, iterate on, and maintain ...

Infrastructure Engineer - Azure/M365

Hiring Organisation
The Curve Group
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
Salary negotiable
Using PowerShell to automate administration, remediation and operational tasks Investigating complex incidents, carrying out root-cause analysis and driving permanent fixes Supporting monitoring and observability across infrastructure and applications, identifying issues before they become major incidents Working with backup, disaster recovery and resilience, including understanding RPO/RTO requirements Supporting … deployment and renewal Building and managing complex Group Policy environments Using PowerShell to automate repetitive administration or remediate issues at scale Implementing monitoring/observability and using logs and telemetry to investigate incidents Supporting environments with significant numbers of users, devices, servers or applications Working within structured ITIL/change ...

Global Account Manager

Location
Maidenhead, England, United Kingdom
business and the customer in a trusted advisor/consultative approach; and establish credibility quickly with senior-level executives across the organizations You understand Observability, Security and/or technology-as-a-service space (Infrastructure-as-a-Service, Software-as-a-Service, Platform-as-a-Service) You can bring your … have knowledge of the IT ecosystem in large, global, enterprises Why you will love being a Dynatracer Dynatrace is a leader in unified observability and security. We provide a culture of excellence with competitive compensation packages designed to recognize and reward performance. Our employees work with the largest cloud providers ...

Software Engineer - Life AI Platform AI & Robotics Oxford, England, United Kingdom

Location
Oxford, England, United Kingdom
hand a ticket to at the beginning. Build and operate the technical foundations for agentic systems, including orchestration, tool interfaces, state management, evaluation, observability and failure recovery. Design evaluative frameworks to understand if the deployed model is working Work directly with scientists, observe how they operate and translate what … have dealt with agentsand their typical failure modes. You’ve designed andbuiltproduction-readysystems.You can make sound decisions about APIs, data models, testing, deployment, observability and failure recovery. You’ve built policy-aware systems with durable audit trails.Constraints are enforced by design, and every decision leaves a trustworthy, queryable record. ...

Cloud-Native Network Automation Engineer

Location
Abingdon, England, United Kingdom
Capgemini is seeking an Automation & Platform Engineer in the UK to build the execution layer for autonomous network operations. You will design, implement, and operate automation workflows, orchestration pipelines, and platform services to safely trigger ...

IT Operations Programme Lead - AIOps

Hiring Organisation
Lorien
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
Salary negotiable
enterprise IT operations, knows how incidents, monitoring, and on-call really work at scale, and can delivery complex transformation. Experience: Lead the observability & AIOps programme: research, scoping, platform evaluation, proof-of-concept, business case, and delivery into production. Design the operating model Run a product evaluation and proof-of-concept … prem infrastructure, networking, identity, end-user computing. Deep expertise in one or two; credibility in all. Real-world experience with the monitoring and observability landscape - events, metrics, logs, traces - and a clear-eyed view of what AIOps genuinely delivers versus what the marketing says. Enterprise-scale ITSM/CMDB ...

Service Architecture Lead

Hiring Organisation
Danaher
Location
Portsmouth, United Kingdom
Employment Type
Full Time
deployability, recoverability, diagnostics, and upgradeability Own core design domains including: Field Installation & upgrade strategy Backup, restore & data integrity Crash recovery & resilience Diagnostics, logging, and observability IT/OT integration (networking, domain, security) Translate field and service requirements into technical architecture requirements early in product development Lead or direct work … Broad technical knowledge across multiple service domains, including software deployment and upgrades, backup and recovery, IT/OT integration, industrial networking, historian platforms, diagnostics, observability, automation validation, or system resilience. Travel, Motor Vehicle Record & Physical/Environment Requirements: Ability to travel – up to 20% domestic and international Must have ...

Incident Command Manager

Hiring Organisation
Barclays
Location
london (city of london), south east england, united kingdom
Role : Major Incident Manager Location : London Duration : 6 months (PAYE) Role Overview Lead the Major Incident capability for critical technology, with a specific focus on payments and financial crime platforms. You will command crisis response ...