Job Details

Lead, SRE Engineer
Job Description
Requisition Number:  58871
Job Location:  Guangzhou, CHN
Global Grade:  Band 5
Work Type:  Hybrid Working
Employment Type:  Permanent
Posting Start Date:  27/07/2026
Posting End Date:  01/11/2026
Job Description: 

Job Summary

As the Lead, Markets Site Reliability Engineering, you will define and drive the reliability engineering strategy across the Markets technology estate. You will lead a global team of Site Reliability Engineers, establishing SRE standards, governance, observability practices, resilience objectives, and automation capabilities that improve service reliability, reduce operational risk, and enhance engineering productivity. The role combines technical leadership, people leadership, and operational governance, ensuring SRE practices are consistently adopted across Markets applications and platforms.

 

Accountable for the overall effectiveness, maturity and adoption of Site Reliability Engineering practices across the Markets technology estate, including service reliability, observability maturity, operational resilience, incident reduction and automation outcomes.

Key Responsibilities

SRE Strategy & Governance
•    Define and maintain the Markets SRE strategy, operating model, standards and roadmap.
•    Establish and govern SLI, SLO and error budget frameworks across Markets technology.
•    Develop reliability maturity assessments and improvement plans across critical applications and platforms.
•    Drive reduction of operational risk through systematic remediation of recurring reliability issues.
•    Present reliability metrics, risk themes and resilience improvement plans to senior technology leadership.
•    Chair or participate in reliability reviews, service health reviews and major incident governance forums.
•    Ensure consistency of observability, resilience and operational readiness standards across Markets technology teams.

 

Site Reliability Engineering
•    Drive and oversee reliability improvement programmes across critical Markets services.
•    Establish and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to drive reliability-focused engineering decisions.
•    Lead the diagnosis and resolution of production incidents, reducing Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).
•    Drive root cause analysis and permanent remediation of recurring production issues.
•    Proactively identify reliability risks and implement measures to improve resilience and fault tolerance.
•    Define standards and provide technical leadership for monitoring, observability and capacity management solutions across the Markets estate.
•    Partner with development teams to improve application operability, observability, recoverability, and production readiness.
•    Support disaster recovery, resiliency testing, failover exercises, and service continuity planning.
•    Drive operational excellence through automation, self-healing capabilities, and reduction of manual operational effort.
•    Act as the senior SRE escalation point during major incidents, coordinating technical recovery activities and ensuring effective communication to senior stakeholders.
•    Drive the identification and remediation of systemic reliability risks revealed through production incidents.

Observability & Automation
•    Develop and enhance observability solutions covering metrics, logs, traces, synthetic monitoring, and service health monitoring.
•    Build automated operational workflows, runbooks, and remediation capabilities.
•    Develop dashboards and analytics that provide actionable insights into application health, performance, capacity, and reliability trends.
•    Support the ongoing evolution of the bank’s monitoring and observability capabilities.

 

AI & Copilot Engineering
•    Define the Markets AIOps and Reliability Automation roadmap.
•    Prioritise and sponsor AI-enabled use cases that improve operational efficiency, resilience and engineering productivity. 
•    Govern the design and implementation of Microsoft Copilot-integrated operational assistants. 
•    Establish standards and controls for AI adoption within production support and reliability engineering. 
•    Identify opportunities to reduce operational toil through AI-assisted diagnosis, change risk assessment, incident management and knowledge management.


Additional Responsibilities
•    Act as the primary SRE representative within Markets Technology leadership forums. 
•    Partner with Application Development, Production Support, Architecture, Infrastructure and Cloud Engineering leaders to improve reliability outcomes. 
•    Influence engineering teams to adopt best practices relating to resilience, observability, automation and operational readiness. 
•    Provide regular updates on reliability performance, operational risk and resilience improvements to senior stakeholders.

Technical Skills

Strong experience with several of the following technologies:
•    Prometheus and AlertManager
•    Grafana
•    OpenTelemetry (Metrics, Logs, Tracing)
•    Elastic Stack or equivalent observability platforms
•    Application Performance Monitoring (APM) tools
•    Synthetic monitoring solutions
•    Kafka / Confluent Kafka
•    ServiceNow ITOM and Event Management
•    Azure, AWS, AKS, EKS, or Kubernetes-based platforms
•    Terraform, Ansible, Chef, Puppet, or similar automation tools
•    Experience with Shell scripting, Java, Python or Ruby
•    Experience with Web Technologies (Apache, HTML, JavaScript, HTTP, XML)

•     Experience in define the Markets AIOps and Reliability Automation roadmap is a big plus

 

Programming and scripting experience in one or more of:
•    Python
•    Java
•    Go
•    Shell scripting
•    PowerShell

Reliability Engineering Competencies
•    Advanced troubleshooting skills across applications, middleware, cloud infrastructure, databases, networking, and distributed systems.
•    Production incident management and major incident support experience.
•    Root cause analysis and problem management expertise.
•    Capacity planning, performance engineering, and resiliency testing experience.
•    Strong understanding of operational risk and production stability principles.

 

AI & Automation Experience (Preferred) Experience in one or more of the following areas would be highly advantageous:
•    Microsoft Copilot Studio and Copilot extensibility.
•    AI agent development and orchestration.
•    Generative AI platforms and Large Language Models (LLMs).
•    AIOps platforms and intelligent event management.
•    Retrieval-Augmented Generation (RAG) architectures.
•    Operational knowledge management and search platforms.
•    AI-enabled observability and incident management solution

Strategy
•    Awareness and understanding of the T&O 30 business strategy and model appropriate to the role. Support and the enablement of the Central Monitoring & Observability strategy, goals and objectives by developing prioritized features aligned to the Catalyst and Tech Simplification programmes.

 

Business
•    The Markets Site Reliability Engineering (SRE) team is a global function focused on ensuring the reliability, performance, resilience, and operational stability of the bank’s critical Markets applications and platforms. Working closely with development, infrastructure, and cloud engineering teams, the SRE team drives observability, automation, incident reduction, and operational excellence across the Markets technology estate.

•    The ideal candidate will have strong experience supporting large-scale production environments and deep expertise in observability technologies such as Elastic, Grafana, Open Telemetry, or Geneos, together with supporting technologies including Kafka, cloud platforms, and automation frameworks. They will apply SRE principles to improve service reliability, accelerate incident resolution, reduce operational toil, and enhance platform resilience.

•    In addition, the candidate will contribute to the development of AI-powered operational tooling and Microsoft Copilot-integrated agents, helping automate troubleshooting, knowledge retrieval, incident analysis, and operational workflows. This role requires a blend of reliability engineering, software development, and automation skills, with a focus on improving both platform stability and engineering productivity across the Markets domain.

 

Processes
•    As the SRE lead, you will play a crucial role in ensuring the stability, reliability, and performance of our Markets applications and platforms, thereby enabling our organization to deliver exceptional services to our internal stakeholders by adhering to the Enterprise SDLC (eSDLC) framework and guidelines.

People & Talent
•    Lead, coach and develop a global team of Site Reliability Engineers located across China and Poland.
•    Establish capability development plans, technical mentoring programmes and career development frameworks for the SRE team.
•    Drive recruitment, succession planning, performance management and workforce planning activities.
•    Foster a culture of engineering excellence, continuous improvement, accountability and operational ownership.
•    Promote adoption of modern SRE, observability, automation and AI engineering practices across the team.
•    Actively engaging in stakeholders’ conversations, providing timely, clear and actionable feedback to deliver solution within timeline. 


Risk Management
•    Identify and escalate systemic technology risks impacting service reliability, resilience and operational stability, ensuring appropriate mitigation plans are established and executed.
•    The ability to interpret the Group’s technical and security (ICS) control requirements and information to identify potential risks and key issues based on this information and put in place appropriate controls and measures to mitigate or minimize risk to the central monitoring & observability platform delivery.

 

Governance
•    Awareness and understanding of the eSDLC framework, in which the T&O software delivery operates, and the requirements and expectations relevant to the role. 
•    Responsible for adhering to the effectiveness of the central monitoring and observability platform deliver governance, based on oversight and controls of the eSDLC framework.

Regulatory & Business Conduct
•    Display exemplary conduct and live by the Group’s Values and Code of Conduct. 
•    Take personal responsibility for embedding the highest standards of ethics, including regulatory and business conduct, across Standard Chartered Bank. This includes understanding and ensuring compliance with, in letter and spirit, all applicable laws, regulations, guidelines and the Group Code of Conduct.
•    Effectively and collaboratively identify, escalate, mitigate and resolve risk, conduct and compliance matters.

 

Key stakeholders
•    Global Head, Markets Production Management  
•    Markets PSS Leads 
•    Markets PSS Managers
•    Head, Observability 

 

Other Responsibilities
•    Embed Here for good and Group’s brand and values in the Observability Platform Team; Perform other responsibilities assigned under Group, Country, Business or Functional policies and procedures; Multiple functions (double hats); [List all responsibilities associated with the role]

Skills and Experience

•    Reliability Engineering 
•    Observability Architecture
•    Incident & Crisis management 
•    Software Engineering
•    Software Quality Assurance
•    Cloud Computing

Qualifications

•    Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent practical experience.
•    Minimum 10+ years of IT experience, including at least 5 years in SRE, Production Engineering, Platform Engineering, DevOps, or a similar reliability-focused role.
•    Experience leading technical teams within financial markets or mission-critical environments.
•    Proven experience driving enterprise-wide reliability, observability and automation initiatives.
•    Strong understanding of distributed systems, cloud-native architectures, and large-scale production operations.
•    Experience with software engineering principles, source control, CI/CD pipelines, and deployment automation.
•    Training: Agile Delivery, SRE
•    Any Professional certifications in SRE, monitoring and observability platforms such as ElasticSearch, Grafana, or ITRS Geneos:
•    Certified Kubernetes Administrator (CKA)
•    Kubernetes and Cloud Native Associate (KCNA)
•    Certified Administrator for Apache Kafka
•    Red Hat Certified Specialist in Event-Driven Development with Kafka
•    AWS Certified SysOps Administrator – Associate
•    Languages: English

About Standard Chartered

We're an international bank, nimble enough to act, big enough for impact. For more than 170 years, we've worked to make a positive difference for our clients, communities, and each other. We question the status quo, love a challenge and enjoy finding new opportunities to grow and do better than before. If you're looking for a career with purpose and you want to work for a bank making a difference, we want to hear from you. You can count on us to celebrate your unique talents and we can't wait to see the talents you can bring us.

Our purpose, to drive commerce and prosperity through our unique diversity, together with our brand promise, to be here for good are achieved by how we each live our valued behaviours. When you work with us, you'll see how we value difference and advocate inclusion.

Together we:

  • Do the right thing and are assertive, challenge one another, and live with integrity, while putting the client at the heart of what we do
  • Never settle, continuously striving to improve and innovate, keeping things simple and learning from doing well, and not so well
  • Are better together, we can be ourselves, be inclusive, see more good in others, and work collectively to build for the long term

What we offer

In line with our Fair Pay Charter, we offer a competitive salary and benefits to support your mental, physical, financial and social wellbeing.

  • Core bank funding for retirement savings, medical and life insurance, with flexible and voluntary benefits available in some locations.
  • Time-off including annual leave, parental/maternity (20 weeks), sabbatical (12 months maximum) and volunteering leave (3 days), along with minimum global standards for annual and public holiday, which is combined to 30 days minimum.
  • Flexible working options based around home and office locations, with flexible working patterns.
  • Proactive wellbeing support through Unmind, a market-leading digital wellbeing platform, development courses for resilience and other human skills, global Employee Assistance Programme, sick leave, mental health first-aiders and all sorts of self-help toolkits
  • A continuous learning culture to support your growth, with opportunities to reskill and upskill and access to physical, virtual and digital learning.
  • Being part of an inclusive and values driven organisation, one that embraces and celebrates our unique diversity, across our teams, business functions and geographies - everyone feels respected and can realise their full potential.
Information at a Glance