CUDA Engineer
DescriptionFuse Energy is an energy startup on a mission to make energy abundant and affordable, fast. We combine first-principles thinking with cutting-edge technology to build a radically better energy system.
We've raised over $200M from top-tier investors including Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, 20VC, Hummingbird and Collaborative Fund, alongside strategic angels including Nico Rosberg and GPs behind Meta, Revolut, Spotify and Uber.
We're building a fully integrated energy company: developing our own solar, batteries and other generation projects, building our own hardware, improving and developing grid infrastructure, trading power in real time, using AI across the business, and installing distributed energy in homes. By selling directly to consumers we cut out the middleman, lower costs and pass the savings on to our customers.
As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure at the intersection of energy and AI. We're looking for a CUDA Engineer to write and optimise the low-level GPU code that powers our inference workloads: designing custom CUDA kernels, tuning performance across memory bandwidth and compute bottlenecks, and squeezing maximum throughput out of every GPU in our fleet, working at the level of SMs, warps and memory hierarchies.
ResponsibilitiesWrite and optimise custom CUDA kernels for core transformer inference operationsProfile kernels to identify and eliminate bottlenecks in occupancy, memory throughput and warp divergenceApply kernel fusion to reduce memory round-trips and launch overhead across inference pipelinesOptimise memory access patterns and manage the memory hierarchy for maximum bandwidth utilisationImplement quantisation-aware kernels and mixed-precision arithmetic to reduce latency and memory footprintBuild and tune caching mechanisms for efficient autoregressive decodingTune kernel launch configurations for target GPU architecturesBenchmark kernels against existing baselines and drive measurable throughput and latency improvementsWrite tests for CUDA code to catch performance and correctness regressionsMaintain internal CUDA libraries and contribute to team coding standards and documentationRequirements4+ years writing production CUDA code, with a track record of shipping performance-critical kernelsDeep understanding of GPU microarchitecture: warps, occupancy, register pressure and memory hierarchyStrong CUDA C++ skills, including streams and asynchronous executionHands-on experience profiling to diagnose compute-bound vs memory-bound bottlenecksExperience with kernel fusion, memory coalescing and avoiding warp divergenceExperience writing quantised and mixed-precision kernelsSolid grasp of parallel algorithm design and numerical precision tradeoffsBonus: transformer/attention-style kernels or autoregressive decoding; building high-performance GPU libraries from scratch; HPC or latency-critical performance engineering; multi-GPU or multi-node kernel-level optimisation; comfortable reading PTX/SASS to validate kernel efficiencyBenefitsCompetitive salary and eligibility for equityBiannual bonus schemeFully expensed tech to match your needsPrivate health insuranceBreakfast and dinner allowance for office-based employeesAs we hire globally, benefits vary by location. Job SummaryID: C0F3C2B4F1Department: Type: full time