Software Engineer, GPU Infrastructure- ChatGPT Engineering
- Hiring Organisation
- Jobleads-UK
- Location
- Greater London, England, United Kingdom
reduce operational overhead. Improve observability, reliability, and operational efficiency across thousands of GPUs. Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response. Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance. Partner closely with research, platform, networking, and systems teams to continuously improve … rather than relying on manual processes. Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure. Understand infrastructure observability, monitoring, capacity planning, and incident management. Enjoy identifying cross‐team pain points and building reusable platforms that improve developer productivity. Are comfortable working across software engineering and systems operations ...