Systems Engineering Manager, Site Reliability Engineering, ML Compute
- Location
- Greater London, England, United Kingdom
. Track record of mentoring technical leads. Proven success leading and influencing multiple technical teams. About the job Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google's services—both … internally critical and our externally-visible systems—have reliability, uptime appropriate to users' needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance. Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating ...