Manager Technology Ops Engineering
Job ID: 26011749Posted: 2026-08-06Location: BURGESS HILL, WEST SUSSEX, United KingdomJob Function: Technology OperationsSchedule: Full timeShift: DayWorkplace: HybridCareer Area: TechnologyCompany: American ExpressDescriptionThe Enterprise Technology Services organization partners with every part of the American Express business to power the company’s growth and innovation with trust and efficiency, and drive competitive differentiation with speed. We support the delivery and operations of technology, digital, and data capabilities, platforms, and services globally. Specifically, our team is responsible for the company’s technology engineering, architecture, and infrastructure, providing 24x7 support to ensure an uninterrupted, high-quality experience for customers and colleagues. We also provide product management for core enterprise platforms, and lead technology risk and information security, enterprise data governance and platforms, digital product and design, and enterprise AI platforms on behalf of the company. Manager, Site Reliability Engineering leads and mentors Site Reliability Engineering (SRE) teams, fostering a culture of continuous improvement and inclusivity, while collaborating across the organization to enhance system resilience, scalability, and alignment with business objectives.ResponsibilitiesManages and leads a team of Site Reliability Engineering colleagues, enabling a culture of continuous learning, growth opportunities, and inclusivity for all individual colleagues and teamsProvides leadership, guidance, and coaching to Site Reliability Engineering teams, supporting training and development of best practices in software development, resiliency, and non-functional system requirementsRecruit and develop a high-performing team, recognizing and rewarding achievements, and creating an environment that motivates and energizes colleagues to achieve best business objectivesOversees and facilitates collaboration with Software Engineering teams to design and implement features that improve system resilience, scalability, and performance; ensuring optimal functionalityCollaborates with executives, product managers, and other stakeholders to ensure SRE principles are embedded throughout the organizationLeads comprehensive chaos engineering experiments and resiliency tests, driving the analyzation of outcomes and implementation of improvements that enhance system robustness and recovery capabilitiesPlans regular drills and strategic planning to ensure organization is prepared for and can swiftly recover from complex and unexpected disruptionsCollaborates and co-creates effectively with teams in product and the business to align technology initiatives with business objectives. Balance feature development speed and reliability with well-defined service level objectivesRecognizes opportunities to adopt innovative technologies to enable business capabilities. Explores new automation techniques to refine the agility, speed and quality of engineering initiatives and effortsQualificationsBachelor’s degree in computer science, Information Technology, Engineering, and/or comparable experience; advance degree preferredKnowledge of modern observability stack – Splunk, Elastic Search, Prometheus, GrafanaKnowledge of containerization technologies (e.g., Kubernetes, Docker) and microservices architectureKnowledge of observability tools and methodologies, including experience with logging, monitoring, tracing, and performance analysis platformsKnowledge of cloud-based Site Reliability Engineering (SRE) practices and experience with public cloud platforms such as AWS, Azure, or Google Cloud.Knowledge of Jira, confluence, rally and project management tools including MS office suit. Work Experience:Must have Domain knowledge of Cards Payments systems. Understanding of E2E workflows of Authorization Approval and clearing & reconciliation processes. Have SRE experience with knowledge of SRE functionsHave knowledge on Splunk, ELF/Kibana and Prometheus/Grafana and experience to use these tools to automate and configure alerting and build dashboards.Have experience in application support (Must have). Application support in cloud-based environment. Conceptual/support knowledge of microservices in cloud environment and deployment processIncident management system knowledge (service now)Good communication skills to run production bridges.Experience in source control using tools such as Git with DevOps and IT automation concepts.Basic UNIX knowledge and any programing language (preferably java, go lang or UI stack).Some knowledge in Redhat Open Shift 3.9/3.11 or Kubernetes 1.9/1.11 and above.Perform day today support activities to track incidents, respond timely on incidents and review and analyze issues at level 2. Create automation dashboards using Splunk, Grafana and KibanaFlexible to work shifts (only day shift, start may be little late than usual time). And ready to provide weekend support as per roster.Review current issues and work with Engineering team to get code fixed and deployed.Familiar with Agile or other rapid application development methodsExperience with design and coding across one or more platforms and languages as appropriateExperience with distributed (multi-tiered) systems, algorithms, and relational databasesA proactive approach to spotting problems, areas for improvement, and performance bottlenecks.Good Knowledge of Networking & Services like TCP/UDP, rest APIs. Able to understand and use complex data structures and associated componentsDesigns, codes, tests, maintains, and documents applicationsTakes part in reviews of own work and reviews of colleagues' workDefines test conditions based on the requirements and specifications providedDepending on factors such as business unit requirements, the nature of the position, cost and applicable laws, American Express may provide visa sponsorship for certain positions.