Job Summary
The role is accountable for providing tactical and operational support for production services across one or more technology platforms/domains, ensuring optimal service stability, availability, performance, and resilience.
Key Responsibilities
• Ensure maximum service quality and production stability through rapid and effective response to technical incidents, while driving continual service improvement through trend analysis, problem management, and proactive identification of improvement opportunities.
• Manage technical recovery and service restoration for High Severity Incidents (CERT/MIM), service outages, and medium/high severity incidents, providing end-to-end support and implementing timely resolutions within agreed SLAs.
• Lead and coordinate incident management activities, including stakeholder communications, escalation management, and technical bridge facilitation during critical incidents.
• Perform root cause analysis (RCA) for High Severity and recurring incidents, ensuring corrective and preventive actions are identified, tracked, and implemented to closure.
• Own the operational stability, availability, and performance of production systems, directing second- and third-level support teams for problem diagnosis and resolution in accordance with agreed SLAs and OLAs.
• Manage production changes, releases, deployments, and rollouts with zero or minimal impact to business services. Ensure comprehensive implementation, validation, rollback, contingency, and communication plans are in place for all production activities.
• Review and assess the impact of dependent changes across applications, infrastructure, databases, middleware, cloud platforms, and networks to minimize production risk.
• Drive proactive monitoring and event management by identifying opportunities for automation, alert optimization, early issue detection, and operational efficiency improvements.
• Support capacity management, resiliency testing, disaster recovery (DR), and business continuity planning (BCP) activities to ensure operational readiness.
• Create, maintain, and continuously improve Production Engineering documentation, operational procedures, runbooks, recovery guides, knowledge articles, and contingency plans.
• Ensure adherence to operational governance, security, compliance, audit, and risk management requirements across supported services.
• Provide inputs to the PE Manager for operational dashboards and service reviews, including incident trends, problem trends, availability metrics, service improvement plans (SIPs), RCA action tracking, and platform health indicators.
• Collaborate with application development, infrastructure, security, architecture, and business teams to improve reliability, reduce technical debt, and enhance service resilience.
• Participate in and support cross-training, knowledge transfer, mentoring, and capability-building activities within the Production Engineering organization.
• Identify and drive opportunities for automation, shift-left initiatives, operational simplification, and reduction of manual effort to improve service reliability and support efficiency.
• Act as a technical lead during major incidents, complex problem investigations, production go-lives, and critical business events, providing technical guidance and decision-making support to internal and external stakeholders.
Strategy
• To be accountable to execute the strategy devised for the business unit.
Business
• Fully accountable in incident, problem, change and risk management which relates to the production application/system.
Processes
• Create, Review and update Production Engineering documentation. Update of contingency (DR/BCP) documentation and processes
People & Talent
• Participate in cross-training and knowledge transfer activities within support teams
Risk Management
• Responsible to proactively identify the risks in the application and manage the mitigation actions. Responsible for managing, tracking and timely closure of risks and other compliance related issues in Riskwise (Information Security risks) & M7 (Operational Risks).
Governance
• Provide inputs to management for monthly dashboard that provide information on incident and problem trends along with SIP and RCA Action Items.
Regulatory & Business Conduct
• Display exemplary conduct and live by the Group’s Values and Code of Conduct.
• Take personal responsibility for embedding the highest standards of ethics, including regulatory and business conduct, across Standard Chartered Bank. This includes understanding and ensuring compliance with, in letter and spirit, all applicable laws, regulations, guidelines and the Group Code of Conduct.
• Effectively and collaboratively identify, escalate, mitigate and resolve risk, conduct and compliance matters.
Key stakeholders
• PE Manager / PE Lead
• Application Development Teams
• Business & Product Stakeholders
• CIO / Technology Leadership
• Infrastructure, DBA & Platform Teams
• Security, Risk & Compliance Teams
• Change, Release & Service Management Teams (ITCRM)
• MIM / CERT Teams
• External Vendors and Service Providers
Other Responsibilities
• Embed Here for good and Group’s brand and values in the SRE team; Perform other responsibilities assigned under Group, Country, Business or Functional policies and procedures; Multiple functions (double hats).
Skills and Experience
• AWS
• Oracle
• Linux
• Kubernetes
• API
Qualifications
• Bachelor's Degree in Computer Science, Engineering, or related discipline.
• AWS Certification (Preferred).
• SRE Foundation Certification (Preferred).
• ITIL Foundation Certification (Preferred).
• Experience supporting mission-critical Banking/Risk Analytics applications.
• Strong understanding of Incident, Problem, Change, and Release Management.
• Experience managing High Severity Incidents (CERT/MIM) and RCA execution.
• Exposure to AWS, Linux, SQL, Monitoring & Observability tools, and automation scripting.
• Knowledge of DevOps, CI/CD, resilience engineering, DR, and operational risk controls.
About Standard Chartered
We're an international bank, nimble enough to act, big enough for impact. For more than 170 years, we've worked to make a positive difference for our clients, communities, and each other. We question the status quo, love a challenge and enjoy finding new opportunities to grow and do better than before. If you're looking for a career with purpose and you want to work for a bank making a difference, we want to hear from you. You can count on us to celebrate your unique talents and we can't wait to see the talents you can bring us.
Our purpose, to drive commerce and prosperity through our unique diversity, together with our brand promise, to be here for good are achieved by how we each live our valued behaviours. When you work with us, you'll see how we value difference and advocate inclusion.
Together we:
- Do the right thing and are assertive, challenge one another, and live with integrity, while putting the client at the heart of what we do
- Never settle, continuously striving to improve and innovate, keeping things simple and learning from doing well, and not so well
- Are better together, we can be ourselves, be inclusive, see more good in others, and work collectively to build for the long term
What we offer
In line with our Fair Pay Charter, we offer a competitive salary and benefits to support your mental, physical, financial and social wellbeing.
- Core bank funding for retirement savings, medical and life insurance, with flexible and voluntary benefits available in some locations.
- Time-off including annual leave, parental/maternity (20 weeks), sabbatical (12 months maximum) and volunteering leave (3 days), along with minimum global standards for annual and public holiday, which is combined to 30 days minimum.
- Flexible working options based around home and office locations, with flexible working patterns.
- Proactive wellbeing support through Unmind, a market-leading digital wellbeing platform, development courses for resilience and other human skills, global Employee Assistance Programme, sick leave, mental health first-aiders and all sorts of self-help toolkits
- A continuous learning culture to support your growth, with opportunities to reskill and upskill and access to physical, virtual and digital learning.
- Being part of an inclusive and values driven organisation, one that embraces and celebrates our unique diversity, across our teams, business functions and geographies - everyone feels respected and can realise their full potential.