Senior HPC/linux Engineer -

Mode of working - Hybrid/office based

Hybrid

If Hybrid, how many days are required in office?

3 days

Number of positions

1

The Role

We are seeking an experienced Senior HPC Engineer to manage, secure and maintain the organisation's scientific computing environment. The role will focus on RHEL-based HPC infrastructure, Slurm workload management and scientific application support, while working closely with research scientists and technical teams to deliver reliable, secure and high-performing computing services.

Your responsibilities:

Administer, patch, secure and maintain RHEL 7, 8 and 9 environments across HPC clusters and high-end workstations.

Deploy, configure and manage Slurm, including scheduling, queues, partitions and fair-share policies.

Monitor and optimise cluster health, resource utilisation, storage, networking and job throughput.

Install and support scientific software, compilers, libraries and MPI environments.

Work directly with scientists to understand computational requirements, optimise workloads and resolve application issues.

Troubleshoot complex hardware, operating system, scheduler and application problems, including root-cause analysis.

Manage incidents, problems and service requests through ServiceNow.

Collaborate with networking, storage, security, DevOps teams and external vendors to deliver reliable HPC services.

Your Profile

Essential skills/knowledge/experience:

Minimum of 10 years' enterprise IT experience, including at least 3 to 5 years in an HPC or research-computing role.

Extensive hands-on administration and troubleshooting experience with RHEL 7, 8 and 9.

Proven experience managing HPC clusters and deploying, configuring and operating Slurm.

Experience supporting scientific or research applications in a Linux HPC environment.

Strong troubleshooting skills across hardware, operating systems, schedulers and applications.

Working knowledge of ServiceNow or an equivalent ITSM platform.

Ability to work collaboratively with research scientists and translate computational requirements into practical technical solutions.

Excellent communication and stakeholder-management skills.

Ability to work onsite for a minimum of three days per week and attend at short notice when physical-system support is required.

Desirable skills/knowledge/experience:

Experience with Docker or other container technologies in an HPC environment.

Knowledge of Ansible or similar configuration-management and automation tools.

Experience supporting GPU computing, CUDA and GPU-accelerated workloads on RHEL.

Understanding of MPI libraries, particularly OpenMPI or MPICH, and their integration with Slurm.

Familiarity with InfiniBand, high-speed Ethernet and HPC networking concepts.

Experience with web server configuration and SSL certificate management.

Red Hat certifications such as RHCSA or RHCE, or equivalent qualifications.

Job Details

Company
Infoplus Technologies UK Ltd
Location
Stevenage, Hertfordshire, United Kingdom SG1 1
Employment Type
Contract
Salary
GBP Annual
Posted