Summary
Overview
Work History
Education
Skills
Timeline
Hi, I’m

Soma Sundaram

Information Technology
Chennai,TN
Soma Sundaram

Summary

Head of SRE leading reliability for retail banking systems across a 10+ team member SRE team. Defines error budget strategy, drives blameless RCA for 8+ incidents per month for recurring incidents, and automates 10+ manual production tasks to reduce toil across production services. Aligns infrastructure and IT Ops ownership to improve availability across a multi-country banking footprint.

Overview

14
years of professional experience

Work History

Equity Bank Group

Head of SRE
05.2025 - Current

Job overview

  • Lead a team of around 15 SRE team members supporting the retail banking division at head office.
  • Define SRE strategy for production reliability and availability across Equity Group operations in 6 countries.
  • Drive RCA for weekly and monthly major incidents, tracking blameless action items to strengthen service stability.
  • Coordinate shared ownership of infrastructure and IT Ops applications, reducing manual toil through targeted automation and runbook improvements.
  • Support monitoring, observability, and on-call rotation management to improve response readiness across production systems.

NatWest Group

Site Reliability Engineer
10.2022 - 04.2025

Job overview

  • Resolve 5-10 incidents per sprint across microservices and deployment workflows while keeping support queues moving.
  • Supported microservices on Tomcat by handling deployments and incident management across production support activities.
  • Handled deployment issues in Artifactory, GitLab, TeamCity, Kafka, and the bank deployment tool during release activities.
  • Managed the SQL database lifecycle and supported application maintenance in Oracle and Sybase environments.
  • Worked in Scrum and SRE ways of working to support migration from data center to cloud technologies.
  • Onboarded new servers and services in Splunk, managed SSL certificates in AWS CloudFront, and used Postman for API support.

HCL Technologies

Technical Specialist
09.2021 - 09.2022

Job overview

  • Reduced incident resolution time by 5% through faster triage and escalation of weekly support issues.
  • Maintained application uptime of 9995% while supporting trader-facing fixed income systems.
  • Troubleshot bond and securities trading issues, guiding users through buy and sell activity in the fixed income platform.
  • Performed identity and access management activities and maintained data records for front-office traders.
  • Analyzed bond and equity rules, traced application dependencies, and delivered solutions when issues stalled trading workflows.
  • Supported application enhancement roadmaps and standards while updating Confluence with the latest changes for the team.
  • Coordinated weekly incidents and changes, partnered with the Disaster Recovery team, and managed backup schedules and data cleanup tasks.

Royal Bank of Scotland

Technical Specialist
09.2013 - 08.2021

Job overview

  • Managed a team of 8 members supporting multiple client onboarding applications.
  • Provided application support for client onboarding systems across production and weekend release cycles.
  • Managed change and incident management activities to stabilize applications and coordinate timely recovery.
  • Led annual disaster recovery testing to support compliance and validate application readiness.
  • Drove continuous improvement activities to reduce recurring issues and prevent outages.
  • Trained new joiners and facilitated sessions for business and internal teams on application processes.
  • Coordinated with QA, downstream teams, and release stakeholders on IAM roadmap and standards work.

Tata consultancy services

Technical Manager
05.2012 - 09.2013

Job overview

  • Maintained 9995% application uptime through monitoring, alert review, and timely production support actions.
  • Resolved 30 incidents per shift and restored application stability for Credit Suisse support queues.
  • Managed application support for Credit Suisse, triaging production issues and coordinating resolution across technical teams.
  • Performed root cause analysis and documented recurring failure patterns to reduce repeat incidents.
  • Supported release and change management activities, validating production readiness before deployment windows.
  • Used SQL and Splunk to investigate errors, trace logs, and accelerate issue diagnosis.
  • Maintained runbooks and coordinated on-call handoffs to improve operational consistency across support shifts.

Education

University of Madras
Chennai

from Computer Science
10-2003

Skills

IT operations

Application support

Application maintenance

Release management

Change management

Incident management

Monitoring & observability

Splunk

SQL

Cloud operations

Capacity planning

Runbook automation

Service level objectives

Service reliability management

Delivery management

Project management

Scrum master

Stakeholder management

Root cause analysis

On-call rotation management

Timeline

Head of SRE

Equity Bank Group
05.2025 - Current

Site Reliability Engineer

NatWest Group
10.2022 - 04.2025

Technical Specialist

HCL Technologies
09.2021 - 09.2022

Technical Specialist

Royal Bank of Scotland
09.2013 - 08.2021

Technical Manager

Tata consultancy services
05.2012 - 09.2013

University of Madras

from Computer Science
Soma SundaramInformation Technology